Intelligent campus safety early warning method and system based on edge computing and big data

By processing video data in real time by edge computing nodes and combining the multi-stage training model of the center server, the delay and resource waste of the existing campus behavior monitoring system is solved, real-time and accurate cross-scene behavior monitoring is achieved, and the adaptability and sustainability of the system is improved.

CN120339957AActive Publication Date: 2025-07-18GUANGDONG SANZHU TECH CO LTD

Patent Information

Application Number
CN202510542402.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Due to the centralized processing architecture, the existing campus behavior monitoring system has high delay in video data transmission, which is difficult to meet the real-time monitoring needs. The fixed model cannot adapt to changes in behavior patterns in multiple scenarios, and the resource configuration pattern cannot be dynamically adjusted, resulting in reduced detection accuracy and waste of resources.

Method used

Edge computing nodes are used to collect video data streams in real time, perform frame sequence segmentation and spatial and temporal feature extraction, combine the multi-stage training model of the central server to generate behavioral semantic description sequences, and adjust the acquisition parameters through dynamic priority strategies to realize cross-scene behavioral mode migration and adaptive configuration of resources.

Benefits of technology

It reduces the delay in video data processing, improves the real-time and accuracy of behavior monitoring, enhances the sustainability of the system and scene adaptability, avoids resource waste, and improves the accuracy of abnormal behavior detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339957A_ABST
    Figure CN120339957A_ABST
Patent Text Reader

Abstract

The invention provides a smart campus safety early warning method and system based on edge computing and big data. The method comprises the following steps: collecting multiple paths of video data streams in real time through an edge computing node, performing frame sequence segmentation and spatial-temporal feature extraction, generating an initial behavior feature set, performing multi-dimensional correlation analysis on the initial behavior feature set based on a preset behavior semantic tag, extracting a spatial-temporal behavior feature vector corresponding to a target monitoring scene, and obtaining a spatial-temporal behavior feature vector; and transmitting the time-space behavior feature vector to a central server, inputting the time-space behavior feature vector into a target behavior recognition model, generating a behavior semantic description sequence corresponding to the video data stream, determining a real-time behavior monitoring result according to a matching result of the behavior semantic description sequence and a preset abnormal behavior rule base, receiving the real-time behavior monitoring result through an edge computing node, and sending the real-time behavior monitoring result to the central server. And adaptive adjustment is carried out on acquisition parameters of the video data stream based on a dynamic priority strategy. According to the invention, the real-time performance, the accuracy and the system sustainability of campus behavior monitoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and more particularly, to a smart campus security early warning method and system based on edge computing and big data. Background Art

[0002] With the increasing demand for campus security management, behavior monitoring technology based on data analysis has become an important means to ensure campus security. Existing behavior monitoring systems usually adopt a centralized video processing architecture, where the cloud server performs unified behavior recognition and anomaly detection on surveillance videos, and relies on fixed feature extraction models and preset rule libraries to achieve behavior classification. However, such solutions have significant drawbacks: the centralized processing architecture leads to high video data transmission latency, making it difficult to meet the requirements of real-time monitoring; the fixed model cannot adapt to the dynamic changes of behavior patterns in multiple campus scenarios (such as classrooms, playgrounds, corridors), resulting in a decrease in cross-scenario detection accuracy; behavior recognition based on low-level visual features lacks semantic relevance and is vulnerable to environmental interference, leading to false detections and missed detections; at the same time, the static resource allocation mode cannot dynamically adjust the acquisition parameters according to the real-time risk level, resulting in coexistence of resource contention in high-load scenarios and resource idleness in low-load scenarios. The above problems seriously restrict the practicability and scalability of campus behavior monitoring systems, and there is an urgent need for a new monitoring solution that can integrate the real-time nature of edge computing and the overall nature of big data analysis, and has both scene adaptation ability and in-depth semantic understanding. Summary of the Invention

[0003] The present invention provides a smart campus security early warning method and system based on edge computing and big data.

[0004] In a first aspect, an embodiment of the present invention provides a smart campus security warning method based on edge computing and big data. The method includes: collecting multiple video data streams in real time through edge computing nodes deployed in campus monitoring areas, segmenting the video data streams into frame sequences and extracting spatio-temporal features to generate an initial behavior feature set; performing multi-dimensional correlation analysis on the initial behavior feature set based on preset behavior semantic tags, extracting spatio-temporal behavior feature vectors corresponding to the target monitoring scenarios, and transmitting the spatio-temporal behavior feature vectors to the central server; in the central server, training an initial behavior recognition model in multiple stages according to a historical behavior data set to obtain a target behavior recognition model, where the multi-stage training includes feature fusion based on spatio-temporal correlation and cross-scenario behavior pattern migration; inputting the spatio-temporal behavior feature vectors into the target behavior recognition model to generate a behavior semantic description sequence corresponding to the video data stream, and determining a real-time behavior monitoring result according to the matching result between the behavior semantic description sequence and a preset abnormal behavior rule library; receiving the real-time behavior monitoring result through the edge computing nodes, and adaptively adjusting the acquisition parameters of the video data streams based on a dynamic priority strategy.

[0005] In a second aspect, an embodiment of the present invention provides a smart campus security warning system, including an edge computing node and a central server that are communicatively connected to each other. Computer programs are respectively stored in the memories of the edge computing node and the central server; when the processors of the edge computing node and the central server respectively load and execute the corresponding computer programs, the above-mentioned smart campus security warning method based on edge computing and big data is implemented.

[0006] The smart campus security warning method provided by the present invention collects and extracts spatio-temporal behavior feature vectors of video data streams in real time through edge computing nodes, generates a behavior semantic description sequence in combination with a multi-stage training model of the central server, and dynamically feeds back and optimizes the acquisition parameters, realizing a collaborative improvement of multi-dimensional behavior monitoring capabilities. Through a hierarchical processing architecture of edge computing and central analysis, this method not only effectively reduces the video data processing delay and optimizes the computing resource allocation, but also enhances the adaptability of the model to complex campus scenarios through spatio-temporal feature fusion and cross-scenario behavior pattern migration. At the same time, it uses a semantic rule matching mechanism to improve the accuracy and interpretability of abnormal behavior detection; moreover, the adaptive parameter adjustment mechanism based on the dynamic priority strategy can intelligently balance the detection accuracy and resource consumption according to the real-time risk level, avoiding resource waste or missed detection problems under traditional fixed parameter configurations; in addition, the closed-loop feedback generation of the semantic description sequence and the optimization of the acquisition parameters form a continuously iterative learning system, gradually improving the generalization ability and scenario coverage of behavior monitoring without manual warning, thereby comprehensively improving the real-time performance, accuracy and system sustainability of campus behavior monitoring. Description of the Drawings

[0007] Figure 1 is a flowchart of a smart campus security early warning method based on edge computing and big data provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the composition of a smart campus security early warning system provided by an embodiment of the present invention. Detailed Implementation Manner

[0008] Please refer to Figure 1 , Figure 1 which is a flowchart of a smart campus security early warning method based on edge computing and big data provided by an embodiment of the present invention, and includes the following steps: Step S100: Real-time collect multiple video data streams through edge computing nodes deployed in campus monitoring areas, and perform frame sequence segmentation and spatio-temporal feature extraction on the video data streams to generate an initial behavior feature set.

[0009] In an embodiment of the present invention, edge computing nodes are devices with computing and data processing capabilities distributed in campus monitoring areas. They are close to data sources and can perform preliminary processing on the collected data at the edge side to reduce data transmission pressure. Video data streams are sequences of video signals continuously collected and transmitted by monitoring cameras. Frame sequence segmentation is to divide continuous video data streams into individual video frames in chronological order for subsequent individual processing of each frame. Spatio-temporal feature extraction is to extract feature information related to time and space from the segmented video frames, such as the movement trajectories and posture changes of objects. The initial behavior feature set is a set obtained by integrating various spatio-temporal features extracted, containing the basic feature information of various behaviors presented in the video.

[0010] An exemplary implementation process is as follows: Edge computing nodes use built-in video acquisition modules to obtain multiple video data streams from cameras in various campus monitoring areas. These video data streams exist in the form of continuous video signals and contain information about various scenes and activities on campus. Then, perform frame sequence segmentation on the obtained video data streams, and use a frame-by-frame algorithm to divide the continuous video signal into a series of discrete video frames. Exemplarily, it can be segmented at a fixed time interval according to the frame rate of the video, and the video signal per second is divided into several independent video frames. Next, perform spatio-temporal feature extraction on the segmented video frames. This process can be achieved through computer vision technology, such as using a convolutional neural network (CNN) to process the video frames and extract spatio-temporal features such as the movement trajectories and posture changes of objects. Finally, integrate the various spatio-temporal features extracted to generate an initial behavior feature set. This set can be represented by a multi-dimensional vector, and each dimension corresponds to a specific spatio-temporal feature.

[0011] As an implementation manner, in step S100, the video data stream is subjected to frame sequence segmentation and spatio-temporal feature extraction to generate an initial behavior feature set, which may specifically include: Step S110: The video data stream is subjected to frame splitting processing to obtain a continuous video frame sequence, and the video frame sequence is filtered for noise based on the pixel change rate between adjacent video frames.

[0012] Frame splitting processing is the process of splitting a continuous video data stream into individual video frames in chronological order. Through frame splitting processing, the video data can be converted into a discrete data form suitable for computer processing. A continuous video frame sequence is a series of video frames arranged in chronological order obtained after frame splitting processing. The pixel change rate between adjacent video frames is the degree of change in attributes such as color and brightness of corresponding pixel points in two adjacent video frames. Noise filtering is to remove noise interference generated for various reasons (such as camera jitter, light change, etc.) in the video frame sequence to make the video frames clearer and more accurate.

[0013] In specific implementation, a video decoding algorithm can be used to perform frame splitting processing on the video data stream, and the video signal per second is split into a number of independent video frames according to the frame rate. Exemplarily, for a video with a frame rate of 25 frames per second, the video signal per second will be split into 25 independent video frames. Then, the pixel change rate between adjacent video frames is calculated. By comparing the color values or brightness values of corresponding pixel points in adjacent video frames, the difference between them can be calculated, and the average value of the differences is used as the pixel change rate. For pixel points with a pixel change rate less than a preset threshold, they can be considered as noise points, and their values are set to the average value of adjacent pixel points or other appropriate values, thereby realizing noise filtering.

[0014] Step S120: Perform dynamic region detection on the filtered video frame sequence, determine the local image region containing the moving target, and perform multi-scale spatio-temporal convolution processing on the local image region.

[0015] Dynamic region detection is to detect the region containing the moving target in the video frame sequence. A moving target is an object whose position or posture changes in the video, such as a pedestrian, a vehicle, etc. A local image region is a specific region containing the moving target, which is a part intercepted from the entire video frame. Multi-scale spatio-temporal convolution processing is to perform convolution operations of different scales on the local image region to extract spatio-temporal features at different scales.

[0016] When performing dynamic region detection, a background subtraction algorithm can be adopted to compare the current video frame with a pre-established background model to determine regions with significant differences from the background model, and these regions are the dynamic regions containing moving objects. Exemplarily, by calculating the difference between the corresponding pixel points in the current video frame and the background model, the pixel points with a difference greater than a preset threshold are marked as the moving object regions. After determining the local image region, multi-scale spatio-temporal convolution processing is performed on it. A convolutional neural network (CNN) is used to perform convolution operations on the local image region at different scales. Exemplarily, convolutional kernels of different sizes, such as 3*3, 5*5, 7*7, etc., can be used to perform convolution on the local image region to extract spatio-temporal features at different scales.

[0017] Step S130: Extract the motion trajectory features, posture change features, and environmental interaction features of the local image region after multi-scale spatio-temporal convolution processing, and align the motion trajectory features, posture change features, and environmental interaction features along the time axis.

[0018] The motion trajectory features are features such as the motion path and speed of the moving object in the video frame sequence. The posture change features are the changes in the body posture of the moving object in the video frame sequence, such as the standing, walking, bending postures of a human body. The environmental interaction features are the interaction relationships between the moving object and the objects in the surrounding environment, such as the contact and collision between a person and an object. Time axis alignment is to arrange the extracted motion trajectory features, posture change features, and environmental interaction features in chronological order to make them consistent in time.

[0019] Specifically, when implementing, feature extraction is performed on the local image region after multi-scale spatio-temporal convolution processing. For the motion trajectory features, a target tracking algorithm can be adopted to track the position of the moving object in the video frame sequence and record information such as its motion path and speed. For the posture change features, a human pose estimation algorithm can be used to detect the skeletal joint points of the moving object and analyze its posture changes. For the environmental interaction features, through object detection and relationship analysis algorithms, the interaction relationships between the moving object and the objects in the surrounding environment are identified. Then, the extracted motion trajectory features, posture change features, and environmental interaction features are aligned along the time axis. These features can be arranged in chronological order according to the timestamps of the video frames to make them consistent in time.

[0020] As an implementation manner, the above step S130 may specifically include the following steps: Step S131: Divide the local image region after multi-scale spatio-temporal convolution processing into time segments to generate a set of time segments containing start timestamps and end timestamps.

[0021] Time segment division is to divide the local image regions after multi-scale spatio-temporal convolution processing into several time periods in chronological order. The start timestamp is the start time of each time segment, and the end timestamp is the end time of each time segment. The time segment set is a set composed of all the divided time segments.

[0022] When performing time segment division, it can be divided according to the timestamps of video frames at a fixed time interval. Exemplarily, the video frame sequence is divided in the way of one time segment per second, and each time segment contains a certain number of video frames. For each time segment, its start timestamp and end timestamp are recorded to generate a time segment set.

[0023] Step S132: Perform trajectory tracking processing on each time segment in the time segment set, obtain the continuous displacement coordinate sequence of the moving target in the local image region, and generate motion trajectory features according to the vector direction change rate of adjacent coordinates in the displacement coordinate sequence.

[0024] Trajectory tracking processing is to track the position of the moving target within each time segment and record its continuous displacement coordinates. The continuous displacement coordinate sequence is a sequence composed of the continuous displacement coordinates of the moving target within each time segment. The vector direction change rate is the degree of change of the vector direction between adjacent displacement coordinates. The motion trajectory feature is the feature information generated according to the vector direction change rate of adjacent coordinates in the displacement coordinate sequence, which reflects the motion trajectory and speed change of the moving target.

[0025] In specific implementation, a target tracking algorithm is used to perform trajectory tracking processing on each time segment in the time segment set. Exemplarily, algorithms such as the Kalman filter or particle filter can be used to predict and update the position of the moving target to obtain its continuous displacement coordinate sequence. Then, calculate the vector direction change rate of adjacent coordinates in the displacement coordinate sequence. The vector direction change rate can be determined by calculating the change in the angle between adjacent displacement coordinates. Generate motion trajectory features according to the vector direction change rate. For example, the average value, standard deviation, etc. of the vector direction change rate can be used as part of the motion trajectory features.

[0026] Step S133: Perform key point detection processing on each time segment in the time segment set, extract the spatial position sequence of the moving target at the skeletal joint points, and generate pose change features based on the relative displacement amount between adjacent time segments of the spatial position sequence.

[0027] Key point detection processing is to detect the skeletal joint points of a moving target within each time segment and determine their spatial positions. Skeletal joint points are the joint parts of the human body or other objects, such as the shoulders, elbows, knees of the human body, etc. The spatial position sequence is a sequence composed of the spatial positions of the skeletal joint points of the moving target within each time segment. The relative displacement amount is the change amount of the spatial positions of the skeletal joint points between adjacent time segments. The pose change feature is the feature information generated according to the relative displacement amount between adjacent time segments in the spatial position sequence, which reflects the pose change situation of the moving target.

[0028] When performing key point detection processing, a human pose estimation algorithm such as OpenPose can be used to process each time segment in the time segment set and extract the spatial position sequence of the moving target on the skeletal joint points. Then, calculate the relative displacement amount between adjacent time segments in the spatial position sequence. The relative displacement amount can be determined by calculating the spatial distance between the corresponding skeletal joint points between adjacent time segments. Generate pose change features according to the relative displacement amount. For example, the average value, standard deviation, etc. of the relative displacement amount can be used as part of the pose change features.

[0029] Step S134: Perform target relationship parsing processing on each time segment in the time segment set, identify the spatial distance sequence and contact state change sequence between the moving target and surrounding static objects, and generate environmental interaction features based on the coupling relationship between the spatial distance sequence and the contact state change sequence.

[0030] The target relationship analysis process analyzes the relationships between moving targets and surrounding static objects within each time segment. The spatial distance sequence is a sequence composed of the spatial distances between a moving target and surrounding static objects within each time segment. The contact state change sequence is a sequence composed of the changes in the contact states (such as contact or non-contact) between a moving target and surrounding static objects within each time segment. The coupling relationship is the mutual relationship between the spatial distance sequence and the contact state change sequence. The environmental interaction feature is the feature information generated based on the coupling relationship between the spatial distance sequence and the contact state change sequence, which reflects the interaction between the moving target and the surrounding environment. In specific implementation, target detection and relationship analysis algorithms can be used to perform target relationship analysis processing on each time segment in the time segment set. Exemplarily, target detection algorithms such as YOLO can be used to detect the positions of moving targets and surrounding static objects, and then calculate the spatial distances between them to generate the spatial distance sequence. At the same time, by analyzing the pixel overlap situation between the moving target and surrounding static objects, etc., their contact states are determined to generate the contact state change sequence. Then, the coupling relationship between the spatial distance sequence and the contact state change sequence is analyzed. For example, when the spatial distance is small, the contact state may be contact; when the spatial distance is large, the contact state may be non-contact. Environmental interaction features are generated based on this coupling relationship. For example, the correlation coefficient between the spatial distance sequence and the contact state change sequence can be used as part of the environmental interaction features.

[0031] Step S135: According to the start timestamp and end timestamp corresponding to each time segment in the time segment set, perform timestamp synchronization processing on the motion trajectory feature, posture change feature, and environmental interaction feature to generate a time-aligned version of the motion trajectory feature, posture change feature, and environmental interaction feature with a unified time reference.

[0032] Timestamp synchronization processing aligns the motion trajectory feature, posture change feature, and environmental interaction feature according to the start timestamp and end timestamp corresponding to each time segment in the time segment set, so that they have a unified time reference. The time-aligned version is the version of the motion trajectory feature, posture change feature, and environmental interaction feature obtained after timestamp synchronization processing, and they are consistent in time.

[0033] In specific implementation, according to the start timestamp and end timestamp corresponding to each time segment in the time segment set, the motion trajectory features, posture change features, and environment interaction features are arranged in chronological order. For each time segment, the corresponding motion trajectory features, posture change features, and environment interaction features are matched to make them corresponding in time. Exemplarily, if the start timestamp of a certain time segment is t1 and the end timestamp is t2, then the motion trajectory features, posture change features, and environment interaction features extracted within this time range are combined to generate a time-aligned version with a unified time reference.

[0034] Step S136: Based on the time continuity of the motion trajectory features, posture change features, and environment interaction features in the time-aligned version, perform adaptive trajectory interpolation compensation on the feature data corresponding to the missing timestamps to generate the motion trajectory features, posture change features, and environment interaction features after alignment on the complete time axis.

[0035] Time continuity refers to the continuous change of the motion trajectory features, posture change features, and environment interaction features over time. The feature data corresponding to the missing timestamps is the feature data corresponding to some timestamps that are missing in the time-aligned version due to various reasons (such as data loss, detection failure, etc.). Adaptive trajectory interpolation compensation is to estimate and supplement the feature data corresponding to the missing timestamps according to the time continuity of the existing feature data in the time-aligned version. The motion trajectory features, posture change features, and environment interaction features after alignment on the complete time axis are the feature data obtained after adaptive trajectory interpolation compensation, and they are complete on the time axis without missing values. In specific implementation, the time continuity of the motion trajectory features, posture change features, and environment interaction features in the time-aligned version can be analyzed. It can be judged whether their changes are continuous by calculating the difference or change rate between the feature data corresponding to adjacent timestamps. For the feature data corresponding to the missing timestamps, an adaptive trajectory interpolation algorithm is used for compensation. Exemplarily, a linear interpolation algorithm can be used to estimate the value of the feature data corresponding to the missing timestamp according to the values of the feature data corresponding to adjacent timestamps. More complex interpolation algorithms, such as spline interpolation algorithms, can also be used to improve the accuracy of interpolation. Through adaptive trajectory interpolation compensation, the motion trajectory features, posture change features, and environment interaction features after alignment on the complete time axis are generated.

[0036] Step S140: Perform weighted fusion on the aligned features according to the preset spatio-temporal weight distribution matrix to generate an initial behavior feature set.

[0037] The preset spatio-temporal weight distribution matrix is a predefined matrix that contains the weight information of different spatio-temporal features. Weighted fusion is to perform weighted summation on the aligned motion trajectory features, pose change features, and environmental interaction features according to the preset spatio-temporal weight distribution matrix, and fuse them into a unified feature vector. The initial behavior feature set is a set of feature vectors obtained after weighted fusion, which contains the basic feature information of various behaviors presented in the video.

[0038] In specific implementation, weighted summation can be performed on the aligned motion trajectory features, pose change features, and environmental interaction features according to the preset spatio-temporal weight distribution matrix. Suppose the preset spatio-temporal weight distribution matrix is W, the aligned motion trajectory feature is T, the pose change feature is P, and the environmental interaction feature is E. Then the initial behavior feature set F can be expressed as: F = W1*T + W2*P + W3*E, where W1, W2, and W3 are the weight vectors corresponding to the motion trajectory feature, pose change feature, and environmental interaction feature respectively. Through weighted fusion, different spatio-temporal features are fused into a unified feature vector to generate the initial behavior feature set. It can be understood that before weighted fusion after feature alignment, a process of feature normalization or standardization can be included to unify the dimensions of each feature. For example, perform normalization processing (such as Min-Max normalization or Z-Score standardization) on T, P, and E respectively, and map them to the same numerical range (such as [0,1] or standard normal distribution). In other content of the embodiments of the present invention, if it involves similar technical content of fusing features with different dimensions, it should be understood that those skilled in the art can unify the dimensions based on their own well-known technical knowledge for subsequent further feature operations.

[0039] Step S200: Perform multi-dimensional correlation analysis on the initial behavior feature set based on the preset behavior semantic labels, extract the spatio-temporal behavior feature vector corresponding to the target monitoring scenario, and transmit the spatio-temporal behavior feature vector to the central server.

[0040] The preset behavior semantic labels are a series of predefined labels used to describe behaviors, such as "walking", "running", "falling", etc. Multi-dimensional correlation analysis is to analyze the initial behavior feature set from multiple perspectives, determine the features related to the preset behavior semantic labels among them, and establish the correlation relationship between them. The target monitoring scenario is a specific scenario that needs to be monitored, such as the playground or classroom of a campus. The spatio-temporal behavior feature vector is the feature vector corresponding to the target monitoring scenario extracted from the initial behavior feature set, which contains the spatio-temporal feature information of the behavior in the target monitoring scenario. The central server is a server used for centralized processing and storage of data. It receives the spatio-temporal behavior feature vectors from the edge computing nodes and performs further analysis and processing.

[0041] For example, the specific implementation process can be as follows: First, extract a set of semantic keywords that match the target monitoring scenario from the preset behavior semantic tags, and assign a scenario association weight coefficient to each semantic keyword. Exemplarily, for the monitoring scenario of a school campus playground, the semantic keywords may include "running", "kicking a ball", etc. According to actual needs, a weight coefficient is artificially assigned to each keyword to represent its association degree with the monitoring scenario. Then, perform semantic similarity matching between each behavior feature in the initial behavior feature set and the set of semantic keywords, and filter out a subset of candidate behavior features whose similarity exceeds a preset threshold. Natural language processing techniques, such as word vector models, can be used to calculate the similarity between the behavior feature and the semantic keyword. Next, perform cross-correlation analysis on the behavior features in the subset of candidate behavior features in the time dimension and the space dimension to generate a joint feature map of a time continuity distribution map and a space density heat map. By analyzing the distribution of behavior features in time and space, the association relationship between them is determined. Dynamically weight the feature regions in the joint feature map based on the scenario association weight coefficient, enhance the weight value of the feature regions strongly related to the target monitoring scenario, and suppress the weight value of the irrelevant feature regions. Input the weighted joint feature map into a spatio-temporal convolution kernel for local feature aggregation to extract a sequence of feature segments with time dependence and space aggregation. According to the cumulative distribution of the feature segment sequence on the time axis and the overlapping ratio of the space regions, generate a spatio-temporal behavior feature vector that fuses multi-dimensional association relationships. Finally, transmit the spatio-temporal behavior feature vector to the central server through the network.

[0042] As an implementation manner, step S200 includes the following steps: Step S210: Extract a set of semantic keywords that match the target monitoring scenario from the preset behavior semantic tags, and assign a scenario association weight coefficient to each semantic keyword.

[0043] The preset behavior semantic tags are a predefined set of tags used to describe behaviors. The set of semantic keywords is a set composed of keywords related to the target monitoring scenario filtered from the preset behavior semantic tags. The scenario association weight coefficient is a numerical value assigned to each semantic keyword to represent the association degree of the keyword with the target monitoring scenario. When extracting the set of semantic keywords, first analyze the preset behavior semantic tags to determine the tags related to the target monitoring scenario. Then, assign a scenario association weight coefficient to each semantic keyword, and the size of the weight coefficient can be determined according to experience or statistical data.

[0044] Step S220: Perform semantic similarity matching between each behavior feature in the initial behavior feature set and the set of semantic keywords, and filter out a subset of candidate behavior features whose similarity exceeds a preset threshold.

[0045] Semantic similarity matching calculates the similarity between each behavioral feature in the initial behavioral feature set and each keyword in the semantic keyword set. The candidate behavioral feature subset is a subset composed of behavioral features screened from the initial behavioral feature set whose similarity with the semantic keyword set exceeds a preset threshold.

[0046] In specific implementation, a word vector model in natural language processing technology, such as Word2Vec or GloVe, can be used to convert behavioral features and semantic keywords into vector representations. After mapping them to the same vector space, the cosine similarity between them is calculated to measure their semantic similarity. For behavioral features whose similarity exceeds the preset threshold, they are added to the candidate behavioral feature subset.

[0047] Step S230: Conduct cross-correlation analysis of the behavioral features in the candidate behavioral feature subset in the time dimension and the space dimension to generate a joint feature map of the time continuity distribution map and the space density heat map.

[0048] The cross-correlation analysis of the time dimension and the space dimension simultaneously considers the distribution of behavioral features in time and space to determine the correlation relationship between them. The time continuity distribution map is a graph used to represent the continuous change of behavioral features in time. The space density heat map is a graph used to represent the distribution density of behavioral features in space. The joint feature map is a feature map obtained by fusing the time continuity distribution map and the space density heat map, which synthesizes the information of behavioral features in time and space.

[0049] In specific implementation, first, analyze the behavioral features in the candidate behavioral feature subset in the time dimension, count the occurrence frequency of each behavioral feature at different time points, and generate a time continuity distribution map. Then, analyze the behavioral features in the space dimension, count the occurrence frequency of each behavioral feature at different spatial positions, and generate a space density heat map. Finally, fuse the time continuity distribution map and the space density heat map to generate a joint feature map. An image fusion algorithm, such as the weighted average method, can be used to fuse the two graphs.

[0050] Step S240: Dynamically weight the feature regions in the joint feature map based on the scene association weight coefficient, enhance the weight values of the feature regions strongly related to the target monitoring scene, and suppress the weight values of the irrelevant feature regions.

[0051] The scene correlation weight coefficient is a numerical value assigned to each semantic keyword to represent its degree of association with the target monitoring scene. The feature region is different regions in the joint feature map, and each region corresponds to a specific behavior feature or a combination of behavior features. Dynamic weighting is to adjust the feature regions in the joint feature map in real time according to the scene correlation weight coefficient, enhance the weight value of the feature regions strongly related to the target monitoring scene, and suppress the weight value of the irrelevant feature regions. When specifically implemented, a weight value can be assigned to each feature region in the joint feature map according to the scene correlation weight coefficient. For the feature regions strongly related to the target monitoring scene, a higher weight value is assigned; for the irrelevant feature regions, a lower weight value is assigned. Then, the numerical value of each feature region in the joint feature map is multiplied by the corresponding weight value to obtain the weighted joint feature map.

[0052] Step S250: Input the weighted joint feature map into the spatio-temporal convolution kernel for local feature aggregation to extract a sequence of feature segments with time dependence and spatial aggregation.

[0053] The spatio-temporal convolution kernel is a convolution kernel used to process spatio-temporal data, which can consider the information of data in both time and space. Local feature aggregation is to merge and extract the local features in the weighted joint feature map to obtain more representative features. The sequence of feature segments is a sequence composed of feature segments with time dependence and spatial aggregation extracted from the weighted joint feature map. When specifically implemented, the weighted joint feature map is input into the spatio-temporal convolution kernel for convolution operation. The spatio-temporal convolution kernel slides on the joint feature map, performs convolution calculations on local regions, and extracts local features. Then, the extracted local features are aggregated to obtain a sequence of feature segments with time dependence and spatial aggregation. Pooling operations, such as max pooling or average pooling, can be used to aggregate the local features.

[0054] Step S260: Generate a spatio-temporal behavior feature vector that fuses multi-dimensional association relationships according to the cumulative distribution of the sequence of feature segments on the time axis and the overlapping ratio of the spatial regions.

[0055] The cumulative distribution of the feature segment sequence on the time axis is the occurrence frequency and distribution of the feature segment sequence on the time axis. The overlapping ratio of the spatial regions is the overlapping degree of the feature segment sequence in different spatial regions. The spatio-temporal behavior feature vector that fuses multi-dimensional correlation relationships is a feature vector obtained by fusing the cumulative distribution of the feature segment sequence on the time axis and the overlapping ratio of the spatial regions, which comprehensively reflects the multi-dimensional correlation relationships of the behavior features in time and space. In specific implementation, first, the cumulative distribution of the feature segment sequence on the time axis is statistically analyzed, and the occurrence frequency of the feature segment at each time point is calculated. Then, the overlapping ratio of the feature segment sequence in different spatial regions is calculated. The overlapping ratio can be determined by calculating the ratio of the intersection area to the union area of the feature segments in different spatial regions. Finally, the cumulative distribution of the feature segment sequence on the time axis and the overlapping ratio of the spatial regions are fused to generate a spatio-temporal behavior feature vector that fuses multi-dimensional correlation relationships, such as by using vector concatenation or weighted summation for fusion.

[0056] Step S300: In the central server, the initial behavior recognition model is trained in multiple stages based on the historical behavior data set to obtain the target behavior recognition model, where the multiple-stage training includes feature fusion based on spatio-temporal correlation and cross-scene behavior pattern transfer.

[0057] The historical behavior data set is the behavior data in the past period stored in the central server, which contains the spatio-temporal features of various behaviors and the corresponding behavior type labels. The initial behavior recognition model is a pre-constructed model for recognizing behaviors, which can be a deep learning model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN). The multiple-stage training is to train the initial behavior recognition model in multiple stages to gradually improve the performance of the model. The feature fusion based on spatio-temporal correlation is to fuse different spatio-temporal features to extract more representative features. The cross-scene behavior pattern transfer is to transfer the behavior patterns learned in one scene to another scene to improve the generalization ability of the model. The target behavior recognition model is the model obtained after multiple-stage training, which has a high behavior recognition accuracy and generalization ability.

[0058] As an implementation manner, the above step S300 can be specifically implemented as the following steps S310 to S370: Step S310: Extract the first training set with spatio-temporal annotations and the second training set of unannotated cross-scene data from the historical behavior data set, where each sample in the first training set contains the annotated spatio-temporal behavior feature vector and the corresponding normal behavior type label, and each sample in the second training set contains unannotated video segments under different monitoring scenes.

[0059] The historical behavior data set is a large amount of historical behavior data stored in the central server, and this data contains behavior information in different scenarios. The first training set with spatio-temporal annotation is extracted from the historical behavior data set, and each sample contains the annotated spatio-temporal behavior feature vector and the corresponding normal behavior type label. The spatio-temporal behavior feature vector is a vector that describes the features of behavior in time and space, and the normal behavior type label is the annotation of the behavior type, such as "walking", "standing", etc. The second training set of unannotated cross-scenario data is extracted from the historical behavior data set, and each sample contains unannotated video segments in different monitoring scenarios. The unannotated video segment is a video segment without behavior type annotation, and different monitoring scenarios are different areas on campus, such as playgrounds, classrooms, etc.

[0060] When specifically implemented, the historical behavior data set is screened and sorted, and the data with spatio-temporal annotation is extracted to form the first training set. At the same time, the unannotated cross-scenario data is extracted to form the second training set. The screening can be carried out according to the annotation information and scenario information of the data to ensure the accuracy and representativeness of the first training set and the second training set.

[0061] Step S320: Input the first training set into the initial behavior recognition model, and perform cross-frame feature alignment on the spatio-temporal behavior feature vector through the spatio-temporal correlation constraint module to generate the first intermediate feature representation with time continuity.

[0062] The initial behavior recognition model is a pre-constructed model for recognizing behaviors. The spatio-temporal correlation constraint module is a module in the model, which is used to perform cross-frame feature alignment on the spatio-temporal behavior feature vector. For example, it can adopt a bidirectional long short-term memory network (Bi-LSTM) based on the attention mechanism. Cross-frame feature alignment is to align the spatio-temporal behavior feature vectors of different frames so that they have continuity in time. The first intermediate feature representation is the feature representation obtained after cross-frame feature alignment, which has time continuity and can better reflect the dynamic changes of behaviors.

[0063] In specific implementation, the first training set is input into the spatio-temporal correlation constraint module of the initial action recognition model. The module extracts the feature subsequences corresponding to consecutive time segments from the spatio-temporal action feature vectors, and calculates the dynamic time warping distance between the feature subsequences of adjacent time segments. The dynamic time warping distance is a method for measuring the similarity between two time series, which can handle the stretching and deformation of time series. Then, according to the dynamic time warping distance, non-linear interpolation processing is performed on the feature subsequences to generate the fourth intermediate feature representation aligned with the time axis. Next, time dimension slicing is performed on the fourth intermediate feature representation to generate local feature blocks within multiple overlapping time windows. Learnable spatio-temporal attention weights are assigned to each local feature block, and feature enhancement is performed on the local feature blocks based on the spatio-temporal attention weights to generate the fifth intermediate feature representation. Finally, the fifth intermediate feature representation is input into the bidirectional temporal convolutional network, which captures future context dependencies through forward propagation and captures historical context dependencies through backward propagation to generate the sixth intermediate feature representation with bidirectional temporal correlation. The sixth intermediate feature representation is connected with the original spatio-temporal action feature vector through residual connection to generate the first intermediate feature representation.

[0064] As an implementation manner, step S320 may specifically include the following steps S321 to S326: Step S321: Extract the feature subsequences corresponding to consecutive time segments from the spatio-temporal action feature vectors, and calculate the dynamic time warping distance between the feature subsequences of adjacent time segments.

[0065] The spatio-temporal action feature vector is a vector that describes the features of an action in time and space. The feature subsequences corresponding to consecutive time segments are extracted from the spatio-temporal action feature vectors, and each subsequence corresponds to a consecutive time segment. The dynamic time warping distance is a method for measuring the similarity between two time series, which can handle the stretching and deformation of time series.

[0066] In specific implementation, according to the timestamp information, the feature subsequences corresponding to consecutive time segments are extracted from the spatio-temporal action feature vectors. Exemplarily, the spatio-temporal action feature vectors can be segmented into multiple consecutive time segments at a fixed time interval, and each time segment corresponds to a feature subsequence. Then, the dynamic time warping algorithm is used to calculate the distance between the feature subsequences of adjacent time segments. The dynamic time warping algorithm calculates the distance between two time series by finding the optimal matching path between them.

[0067] Step S322: Perform non-linear interpolation processing on the feature subsequences according to the dynamic time warping distance to generate the fourth intermediate feature representation aligned with the time axis.

[0068] Nonlinear interpolation processing performs an interpolation operation on the feature subsequences based on the dynamic time warping distance to achieve time-axis alignment. Time-axis alignment aligns the feature subsequences of different time segments in time so that they have the same time reference. The fourth intermediate feature representation is the feature representation obtained after nonlinear interpolation processing, and it is aligned on the time axis. In specific implementation, the mapping relationship between the feature subsequences is determined according to the dynamic time warping distance. Then, a nonlinear interpolation algorithm (such as the spline interpolation algorithm) is used to perform an interpolation operation on the feature subsequences to generate the fourth intermediate feature representation that is aligned on the time axis. Through nonlinear interpolation processing, the feature subsequences of different time segments can be made more continuous and consistent in time.

[0069] Step S323: Perform time-dimensional slicing on the fourth intermediate feature representation to generate local feature blocks within multiple overlapping time windows.

[0070] Time-dimensional slicing divides the fourth intermediate feature representation according to the time dimension to generate local feature blocks within multiple overlapping time windows. An overlapping time window is one where there is a certain overlapping part between adjacent time windows, which can better capture the dynamic changes of the behavior. A local feature block is a feature block extracted within each time window, which contains the behavior feature information within that time window. In specific implementation, the fourth intermediate feature representation is sliced in the time dimension according to a fixed time window size and overlapping ratio. Exemplarily, if the time window size is 10 frames and the overlapping ratio is 50%, then slicing is performed every 5 frames to generate local feature blocks within multiple overlapping time windows. Each local feature block contains the spatio-temporal behavior feature information within that time window.

[0071] Step S324: Assign learnable spatio-temporal attention weights to each local feature block, and perform feature enhancement on the local feature blocks based on the spatio-temporal attention weights to generate the fifth intermediate feature representation.

[0072] Learnable spatio-temporal attention weights are the weights assigned to each local feature block, and these weights can be learned and adjusted through the training process of the model. The spatio-temporal attention weights are used to represent the importance of each local feature block in the entire behavior feature. Feature enhancement is to perform weighted summation on the local feature blocks according to the spatio-temporal attention weights to enhance the influence of important features. The fifth intermediate feature representation is the feature representation obtained after feature enhancement, and it more prominently highlights the important behavior features.

[0073] In specific implementation, a fully connected layer or an attention mechanism is used to assign learnable spatio-temporal attention weights to each local feature block. Then, each local feature block is multiplied by the corresponding spatio-temporal attention weight and weighted summation is performed to obtain the enhanced feature representation. Through feature enhancement, the model can pay more attention to important behavior features and improve the accuracy of behavior recognition.

[0074] Step S325: Input the fifth intermediate feature representation into a bidirectional temporal convolutional network, capture future context dependencies through forward propagation, and capture historical context dependencies through backward propagation to generate a sixth intermediate feature representation with bidirectional temporal correlation.

[0075] A bidirectional temporal convolutional network is a convolutional network that can process both forward and backward information of a time series simultaneously. Forward propagation calculates from the starting point to the ending point of the time series to capture future context dependencies. Backward propagation calculates from the ending point to the starting point of the time series to capture historical context dependencies. The sixth intermediate feature representation is the feature representation obtained after being processed by the bidirectional temporal convolutional network. It has bidirectional temporal correlation and can better reflect the dynamic changes of behaviors.

[0076] In specific implementation, input the fifth intermediate feature representation into the bidirectional temporal convolutional network. The network calculates future context dependencies and historical context dependencies through forward propagation and backward propagation respectively. Then, merge the results of forward propagation and backward propagation to generate a sixth intermediate feature representation with bidirectional temporal correlation. The bidirectional temporal convolutional network can effectively capture the time series information of behaviors and improve the accuracy of behavior recognition.

[0077] Step S326: Perform a residual connection between the sixth intermediate feature representation and the original spatio-temporal behavior feature vector to generate a first intermediate feature representation.

[0078] Residual connection is to add the input and output of the network to alleviate the vanishing gradient problem and improve the training efficiency of the network. The original spatio-temporal behavior feature vector is the spatio-temporal behavior feature vector input into the spatio-temporal correlation constraint module. The first intermediate feature representation is the feature representation obtained after the residual connection. It combines the information of the original spatio-temporal behavior feature vector and the processed sixth intermediate feature representation.

[0079] In specific implementation, perform element-wise addition between the sixth intermediate feature representation and the original spatio-temporal behavior feature vector to obtain the first intermediate feature representation. Through the residual connection, the information of the original spatio-temporal behavior feature vector can be retained, and at the same time, the information of the processed sixth intermediate feature representation can be fused to improve the quality of the feature representation.

[0080] Step S330: Input the first intermediate feature representation into the spatio-temporal feature fusion module, perform dynamic weight allocation and feature superposition on the features within adjacent time windows to generate a fused second intermediate feature representation.

[0081] The spatio-temporal feature fusion module is a module in the model, which is used to fuse the features within adjacent time windows. For example, it can be implemented as a gated recurrent unit (GRU). Dynamic weight assignment assigns different weights to the features within adjacent time windows according to the importance of the features. Feature superposition superimposes the features within adjacent time windows according to the assigned weights to obtain the fused feature representation. The second intermediate feature representation is the feature representation obtained after being processed by the spatio-temporal feature fusion module, which fuses the feature information within adjacent time windows. In specific implementation, the first intermediate feature representation is input into the spatio-temporal feature fusion module. This module first assigns dynamic weights to the features within adjacent time windows according to the importance of the features. A fully connected layer or an attention mechanism can be used to learn these weights. Then, the features within adjacent time windows are superimposed according to the assigned weights to obtain the fused second intermediate feature representation. Through dynamic weight assignment and feature superposition, the feature information within adjacent time windows can be better fused, improving the accuracy of behavior recognition.

[0082] Step S340: Based on the similarity loss between the second intermediate feature representation and the normal behavior type label, perform supervised training on the initial behavior recognition model to obtain the first intermediate model.

[0083] The similarity loss is a loss function used to measure the degree of difference between the second intermediate feature representation and the normal behavior type label. Supervised training adjusts the parameters of the model by minimizing the similarity loss in the case of labeled data. The first intermediate model is the model obtained after supervised training, which can accurately identify the normal behavior type to a certain extent.

[0084] In specific implementation, a suitable similarity loss function, such as the cross-entropy loss function, is used to calculate the loss between the second intermediate feature representation and the normal behavior type label. Then, an optimization algorithm, such as the stochastic gradient descent algorithm, is used to update the parameters of the initial behavior recognition model to minimize the similarity loss. After multiple iterations of training, the first intermediate model is obtained. Through supervised training, the model can learn the features of the normal behavior type, improving the accuracy of behavior recognition.

[0085] Step S350: Input the second training set into the first intermediate model, and extract scene-invariant features from the unlabeled video segments through the cross-scene behavior pattern migration module to generate a third intermediate feature representation that is independent of the monitoring scene.

[0086] The cross-scenario behavior pattern migration module is a module in the model that is used to extract scene-invariant features from unannotated video segments. Scene-invariant features are features that are not affected by the monitored scene and can reflect the essential features of behaviors. The third intermediate feature representation is the feature representation obtained after being processed by the cross-scenario behavior pattern migration module. It is independent of the monitored scene and has good generalization ability.

[0087] In specific implementation, the second training set is input into the cross-scenario behavior pattern migration module of the first intermediate model. This module uses some unsupervised learning algorithms, such as autoencoders or variational autoencoders, to extract features from unannotated video segments. By learning the latent features in the video segments, scene-invariant features are extracted. Then, these scene-invariant features are combined into the third intermediate feature representation. Through cross-scenario behavior pattern migration, the model can learn the behavior patterns in different scenarios and improve the generalization ability of the model.

[0088] Step S360: Compress and reconstruct the third intermediate feature representation in the time dimension to generate a cross-scenario behavior pattern embedding vector, and perform self-supervised training on the first intermediate model based on the reconstruction error to obtain the second intermediate model.

[0089] Time dimension compression is to compress the third intermediate feature representation in the time dimension to reduce the dimension of the features. Feature reconstruction is to reconstruct the compressed features to restore them to the original feature dimension. The cross-scenario behavior pattern embedding vector is the vector obtained after time dimension compression and feature reconstruction, which represents the cross-scenario behavior pattern. The reconstruction error is the degree of difference between the reconstructed features and the original features. Self-supervised training is to adjust the parameters of the model by minimizing the reconstruction error in the case of unlabeled data. The second intermediate model is the model obtained after self-supervised training, which has better performance in cross-scenario behavior recognition.

[0090] In specific implementation, a time series compression module, such as a long short-term memory network (LSTM) or a gated recurrent unit (GRU), can be used to compress the third intermediate feature representation in the time dimension. Then, a reconstruction module, such as a fully connected layer or a convolutional layer, is used to reconstruct the compressed features. Calculate the reconstruction error between the reconstructed features and the original features, and use an optimization algorithm, such as the stochastic gradient descent algorithm, to update the parameters of the first intermediate model to minimize the reconstruction error. After multiple iterative trainings, the second intermediate model is obtained. Through self-supervised training, the model can learn the cross-scenario behavior patterns and improve the generalization ability of the model.

[0091] Step S370: Alternately and weightedly fuse the parameter of the feature extraction layer of the first intermediate model and the parameter of the feature reconstruction layer of the second intermediate model to generate a target behavior recognition model with cross-scenario adaptability.

[0092] The feature extraction layer parameters are the parameters of the layer in the first intermediate model used for feature extraction. The feature reconstruction layer parameters are the parameters of the layer in the second intermediate model used for feature reconstruction. Alternating weighted fusion is to alternately weight the feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model to generate new parameters. The target behavior recognition model is the model obtained after alternating weighted fusion, which has cross-scene adaptability and can accurately recognize behaviors in different monitoring scenarios.

[0093] In specific implementation, the feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model can be alternately weighted according to a certain weight ratio. For example, a linear combination method such as α * parameter 1+(1 - α) * parameter 2 can be used, where α is the weight coefficient. Through alternating weighted fusion, the parameters of the target behavior recognition model with cross-scene adaptability are generated. Applying these parameters to the model, the final target behavior recognition model is obtained. Through alternating weighted fusion, the advantages of the first intermediate model and the second intermediate model can be comprehensively utilized to improve the cross-scene adaptability and behavior recognition accuracy of the model.

[0094] Step S400: Input the spatio-temporal behavior feature vector into the target behavior recognition model to generate a behavior semantic description sequence corresponding to the video data stream, and determine the real-time behavior monitoring result according to the matching result between the behavior semantic description sequence and the preset abnormal behavior rule library.

[0095] The spatio-temporal behavior feature vector is extracted from the video data stream and contains the feature information of the behavior in time and space. The target behavior recognition model is a model trained for behavior recognition. The behavior semantic description sequence is a sequence that describes the behavior in the video data stream in natural language, which contains information such as the action, interaction, and intention of the behavior. The preset abnormal behavior rule library is a set of predefined abnormal behavior rules, and each rule corresponds to an abnormal behavior type. The real-time behavior monitoring result is determined according to the matching result between the behavior semantic description sequence and the preset abnormal behavior rule library, which indicates whether there is an abnormal behavior in the current video data stream and the type of the abnormal behavior.

[0096] The specific implementation process is as follows: First, perform a time-axis sliding window segmentation on the spatio-temporal behavior feature vector to generate a sequence of feature segments corresponding to consecutive time segments. Then, input the sequence of feature segments into the basic action parsing layer in the target behavior recognition model to extract the limb displacement rate change pattern and direction consistency features of the moving target within the time segment, and generate an atomic action label sequence. Next, input the atomic action label sequence into the interaction relationship reasoning layer in the target behavior recognition model, and based on the time overlap and spatial proximity between atomic action labels, identify the collaborative action pattern between multiple moving targets and generate a set of interaction relationship descriptions. Input the set of interaction relationship descriptions into the intention inference layer in the target behavior recognition model, combine the scene context features corresponding to the time segment, analyze the logical matching degree between the collaborative action pattern and the preset behavior intention template, and generate a behavior intention probability distribution. According to the atomic action label sequence, the set of interaction relationship descriptions, and the behavior intention probability distribution, generate a structured behavior semantic description tree including action level, interaction level, and intention level. Perform time-axis compression and semantic aggregation on the structured behavior semantic description tree, remove redundant description nodes and merge the same type of action branches to generate a behavior semantic description sequence. Finally, perform semantic parsing on the behavior semantic description sequence, extract the key behavior action nodes and the temporal dependency relationships between the nodes. Construct a behavior state transition graph according to the temporal dependency relationships, and calculate the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule library. When the topological similarity exceeds the preset threshold, generate an abnormal behavior type identifier that matches the abnormal behavior rule library. According to the risk level corresponding to the abnormal behavior type identifier, generate a real-time behavior monitoring result including a behavior warning strategy.

[0097] As an implementation manner, step S400 may specifically include the following steps: Step S410: Perform a time-axis sliding window segmentation on the spatio-temporal behavior feature vector to generate a sequence of feature segments corresponding to consecutive time segments.

[0098] The time-axis sliding window segmentation is to use a window with a fixed size to slide on the time axis of the spatio-temporal behavior feature vector, and divide the spatio-temporal behavior feature vector into a sequence of feature segments corresponding to multiple consecutive time segments. Each sequence of feature segments contains the spatio-temporal behavior feature information within a time segment. Specifically, when implementing, according to the preset window size and sliding step, perform a sliding window segmentation on the time axis of the spatio-temporal behavior feature vector. Exemplarily, the window size is 10 frames and the sliding step is 5 frames, then a segmentation is performed every 5 frames to generate a sequence of feature segments corresponding to multiple consecutive time segments. Each sequence of feature segments contains the spatio-temporal behavior feature information within this time segment.

[0099] Step S420: Input the feature segment sequence into the basic action parsing layer in the target behavior recognition model, extract the change pattern of the limb displacement rate and the direction consistency feature of the moving target within the time segment, and generate an atomic action label sequence.

[0100] The basic action parsing layer is a layer in the target behavior recognition model, which is used to parse the basic actions of the moving target within the time segment. The change pattern of the limb displacement rate is the change of the displacement rate of the limb of the moving target within the time segment. The direction consistency feature is the degree of consistency of the movement direction of the limb of the moving target within the time segment. The atomic action label sequence is the sequence obtained by labeling the basic actions within the time segment, and each label corresponds to an atomic action, such as "walking", "turning around", etc.

[0101] In specific implementation, input the feature segment sequence into the basic action parsing layer in the target behavior recognition model. This layer uses some computer vision algorithms, such as the optical flow method or the key point detection algorithm, to extract the change pattern of the limb displacement rate and the direction consistency feature of the moving target within the time segment. Then, according to these features, use a classifier, such as a support vector machine or a neural network, to classify the basic actions within the time segment and generate an atomic action label sequence.

[0102] Step S430: Input the atomic action label sequence into the interaction relationship reasoning layer in the target behavior recognition model, and based on the time overlap degree and spatial proximity between the atomic action labels, identify the collaborative action pattern among multiple moving targets and generate an interaction relationship description set.

[0103] The interaction relationship inference layer is a network layer in the target behavior recognition model. It is used to infer the interaction relationships between multiple moving targets. For example, it is the inference layer in a graph neural network (GNN). The temporal overlap degree is the degree of overlap of atomic action labels in time, and the spatial proximity is the degree of proximity of moving targets in space. The collaborative action pattern is the way of collaborative actions between multiple moving targets, such as "walking together", "colliding with each other", etc. The interaction relationship description set is a set obtained by describing the interaction relationships between multiple moving targets, and each description corresponds to a collaborative action pattern. In specific implementation, the atomic action label sequence is input into the interaction relationship inference layer in the target behavior recognition model. This layer first calculates the temporal overlap degree and spatial proximity between atomic action labels. Then, based on the temporal overlap degree and spatial proximity, a graph neural network or a rule inference system is used to identify the collaborative action patterns between multiple moving targets and generate an interaction relationship description set. Exemplarily, an atomic action label sequence is obtained, and each label corresponds to the basic action of a moving target in a certain time period, such as student A "running" and student B "passing the ball", and the action time range and target position are recorded. Then the temporal overlap degree and spatial proximity are calculated. The temporal overlap degree of the overlapping part of the time ranges of different atomic action labels is determined. For example, the time period of student A "running" is 8:00 - 8:10, and the time period of student B "passing the ball" is 8:05 - 8:12, and the overlap degree is 5 minutes; the spatial proximity is judged by calculating the spatial distance through the target positions, and if the distance is less than the threshold, it is considered proximate. When the temporal overlap degree and spatial proximity of two atomic action labels are high, they are recognized as a collaborative action pattern, such as "passing - receiving collaboration", and finally these patterns are sorted out and described to form an interaction relationship description set such as "there is a passing - receiving collaboration between student A and student B".

[0104] Step S440: Input the interaction relationship description set into the intention inference layer in the target behavior recognition model, and combine the scene context features corresponding to the time segment to analyze the logical matching degree between the collaborative action pattern and the preset behavior intention template, and generate a behavior intention probability distribution.

[0105] The intention inference layer is a layer in the target behavior recognition model. It is used to infer the behavior intention of a moving target and can include a knowledge graph and the network layer of LSTM. The scene context feature is the relevant feature of the scene corresponding to the time segment, such as the type of the scene, environmental information, etc. The preset behavior intention template is a series of predefined templates of behavior intentions, and each template corresponds to a behavior intention, such as "playing", "fighting", etc. The logical matching degree is the logical matching degree between the collaborative action pattern and the preset behavior intention template. The behavior intention probability distribution is the probability distribution obtained by matching the collaborative action pattern with the preset behavior intention template, and each probability corresponds to the possibility of a behavior intention. Specifically in implementation, the set of interaction relationship descriptions is input into the intention inference layer in the target behavior recognition model. This layer combines the scene context features corresponding to the time segment and uses a deep learning model or a knowledge graph to analyze the logical matching degree between the collaborative action pattern and the preset behavior intention template. Then, according to the logical matching degree, a behavior intention probability distribution is generated. Exemplarily, for a certain collaborative action pattern, calculate its matching score with each preset behavior intention template, and normalize the scores to obtain the behavior intention probability distribution.

[0106] Step S450: Generate a structured behavior semantic description tree including an action level, an interaction level, and an intention level according to the atomic action label sequence, the set of interaction relationship descriptions, and the behavior intention probability distribution.

[0107] The structured behavior semantic description tree is a tree structure that contains information about the action level, the interaction level, and the intention level. The action level represents the basic actions of the moving target, the interaction level represents the interaction relationships between multiple moving targets, and the intention level represents the behavior intention of the moving target. For example, each label in the atomic action label sequence can be mapped to a leaf node of the structured behavior semantic description tree, and the start timestamp and duration corresponding to each leaf node are recorded. Then, according to the collaborative action pattern in the set of interaction relationship descriptions, horizontal association edges are established between the leaf nodes. The horizontal association edges contain interaction type identifiers and the number information of the participating moving targets. Based on the intention categories with probability values exceeding the preset threshold in the behavior intention probability distribution, intention branch nodes are created below the root node of the structured behavior semantic description tree, and the leaf nodes corresponding to the horizontal association edges are clustered into the matching intention branch nodes. The leaf nodes with timestamp overlap exceeding the preset ratio in the structured behavior semantic description tree are merged to generate composite action nodes, and vertical association edges are established between the composite action nodes and the corresponding intention branch nodes. According to the time coverage range of the vertical association edges and the logical consistency of the intention branch nodes, semantic generalization processing is performed on the composite action nodes and replaced with high-level behavior description phrases. Through these steps, a structured behavior semantic description tree including an action level, an interaction level, and an intention level is generated.

[0108] Step S460: Perform timeline compression and semantic aggregation on the structured behavioral semantic description tree, remove redundant description nodes, and merge same-type action branches to generate a behavioral semantic description sequence.

[0109] Timeline compression is to compress the structured behavioral semantic description tree on the timeline and remove redundant information in terms of time. Semantic aggregation is to merge nodes with the same semantics to reduce the complexity of the description. Redundant description nodes are nodes in the structured behavioral semantic description tree that do not substantially contribute to the description of behavioral semantics. Same-type action branches are branches with the same action type. The behavioral semantic description sequence is the sequence obtained after timeline compression and semantic aggregation, which more concisely describes the semantic information of the behavior.

[0110] In specific implementation, first perform timeline compression on the structured behavioral semantic description tree. Redundant information in terms of time can be removed by merging nodes with overlapping timestamps. Then, perform semantic aggregation to merge nodes with the same semantics. Exemplarily, multiple nodes representing "walking" are merged into one node. Remove redundant description nodes, such as some intermediate nodes without practical meaning. Finally, convert the processed structured behavioral semantic description tree into a sequence form to generate a behavioral semantic description sequence.

[0111] As an implementation manner, the process of generating the structured behavioral semantic description tree and performing timeline compression and semantic aggregation may specifically include: Step S451: Map each label in the atomic action label sequence to a leaf node of the structured behavioral semantic description tree, and record the start timestamp and duration corresponding to each leaf node.

[0112] The structured behavioral semantic description tree is a tree structure used to represent behavioral semantics. Leaf nodes are the bottommost nodes of the tree, and each leaf node corresponds to an atomic action label. The start timestamp is the time when the atomic action starts, and the duration is the length of time that the atomic action lasts.

[0113] In specific implementation, traverse the atomic action label sequence, map each label to a leaf node of the structured behavioral semantic description tree. At the same time, according to the information of the time segment, record the start timestamp and duration corresponding to each leaf node. Exemplarily, if the time segment corresponding to a certain atomic action label is from frame 10 to frame 20, then the start timestamp of this leaf node is frame 10, and the duration is 10 frames.

[0114] Step S452: Establish horizontal association edges between leaf nodes according to the collaborative action patterns in the interaction relationship description set. The horizontal association edges include interaction type identifiers and the number information of participating moving targets.

[0115] The set of interaction relationship descriptions contains the information of the collaborative action patterns among multiple moving objects. The horizontal association edge is the edge established between the leaf nodes of the structured behavior semantic description tree, which represents the interaction relationship between two atomic actions. The interaction type identifier is a symbol used to identify the interaction type, such as "together", "collision", etc. The number information of the participating moving objects is the number of moving objects participating in the collaborative action. In the specific implementation, traverse the set of interaction relationship descriptions, and establish horizontal association edges between the corresponding leaf nodes according to the information in the collaborative action pattern. Exemplarily, if a certain collaborative action pattern indicates that two moving objects are walking together, then establish a horizontal association edge between the leaf nodes corresponding to these two moving objects, and mark the interaction type identifier "together" and the number of participating moving objects 2 on the edge.

[0116] Step S453: Based on the intention categories with probability values exceeding the preset threshold in the behavior intention probability distribution, create intention branch nodes below the root node of the structured behavior semantic description tree, and cluster the leaf nodes corresponding to the horizontal association edges to the matching intention branch nodes.

[0117] The behavior intention probability distribution is the possibility distribution of each behavior intention. The preset threshold is a pre-set probability value used to screen out the behavior intentions with relatively high possibilities. The intention branch node is a branch node below the root node in the structured behavior semantic description tree, and each intention branch node corresponds to a behavior intention. Clustering is to group the relevant leaf nodes under the same intention branch node. In the specific implementation, traverse the behavior intention probability distribution to determine the intention categories with probability values exceeding the preset threshold. Create intention branch nodes corresponding to these intention categories below the root node of the structured behavior semantic description tree. Then, according to the information of the horizontal association edges, cluster the corresponding leaf nodes to the matching intention branch nodes. Exemplarily, if a certain intention category is "playing", then cluster the leaf nodes corresponding to the horizontal association edges related to "playing" under the "playing" intention branch node.

[0118] Step S454: Merge the leaf nodes in the structured behavior semantic description tree with time stamp overlap exceeding the preset ratio to generate composite action nodes, and establish vertical association edges between the composite action nodes and the corresponding intention branch nodes.

[0119] The time stamp overlap means that there is an overlapping part in the time segments corresponding to two leaf nodes. The preset ratio is a pre-set ratio value used to judge whether the time stamp overlap reaches the condition for merging. The composite action node is a node formed by merging multiple leaf nodes, which represents a composite action. The vertical association edge is the edge from the composite action node to the intention branch node in the structured behavior semantic description tree, which represents the relationship between the composite action and the behavior intention.

[0120] During specific implementation, traverse the structured behavior semantic description tree to determine the leaf nodes where the timestamp overlap exceeds a preset ratio. Merge these leaf nodes into a composite action node. Then, based on the intention corresponding to the composite action node, establish a vertical association edge with the corresponding intention branch node. Exemplarily, if the intention corresponding to a certain composite action node is "play", then establish a vertical association edge between this composite action node and the "play" intention branch node.

[0121] Step S455: According to the time coverage range of the vertical association edge and the logical consistency with the intention branch node, perform semantic generalization processing on the composite action node and replace it with a high-level behavior description phrase.

[0122] The time coverage range of the vertical association edge is the time range corresponding to the composite action node and the intention branch node connected by the vertical association edge. Logical consistency is the degree of logical correspondence between the composite action node and the intention branch node. Semantic generalization processing is to abstract and generalize the description of the composite action node and replace it with a higher-level behavior description phrase.

[0123] During specific implementation, analyze the time coverage range of the vertical association edge and the logical consistency with the intention branch node. For the composite action nodes that meet the logic, perform semantic generalization processing. Exemplarily, if a composite action node includes atomic actions such as "running" and "jumping", and the corresponding intention branch node is "play", then this composite action node can be replaced with a high-level behavior description phrase such as "lively play".

[0124] Step S456: Generate a compressed behavior semantic description sequence based on the chronological order of the high-level behavior description phrases.

[0125] The high-level behavior description phrase is the behavior description phrase obtained after semantic generalization processing. Chronological order arrangement is to arrange according to the chronological order corresponding to the high-level behavior description phrases. The compressed behavior semantic description sequence is the sequence obtained by arranging the high-level behavior description phrases in chronological order, which removes redundant information and is more concise and clear.

[0126] During specific implementation, sort the high-level behavior description phrases according to the start timestamp corresponding to the high-level behavior description phrases. Then, connect the sorted high-level behavior description phrases in sequence to generate a compressed behavior semantic description sequence. Exemplarily, if there are three high-level behavior description phrases "read books quietly", "play lively", and "queue up orderly", and their start timestamps are t1, t2, t3 in sequence, and t1 < t2 < t3, then the compressed behavior semantic description sequence is "read books quietly -> play lively -> queue up orderly".

[0127] As an implementation manner, in step S400, according to the matching result between the behavior semantic description sequence and the preset abnormal behavior rule library, the real-time behavior monitoring result is determined, which may specifically include: step S470: perform semantic parsing on the behavior semantic description sequence, and extract the key behavior action nodes and the temporal dependence relationship between the nodes.

[0128] Semantic parsing is to analyze the behavior semantic description sequence and understand its semantic information. The key behavior action nodes are the nodes that have an important impact on the behavior semantics in the behavior semantic description sequence, and they represent the main actions of the behavior. The temporal dependence relationship between the nodes is the sequence or concurrency relationship between the key behavior action nodes, which reflects the temporal order and logical relationship of the behavior.

[0129] Specifically in implementation, the behavior semantic description sequence can be first segmented into multiple consecutive basic semantic units according to a preset time granularity, and redundant filtering is performed on the basic semantic units based on a semantic coherence threshold to generate a set of redundant-free semantic units. Action keyword matching is performed on each semantic unit in the redundant-free semantic unit set to identify candidate semantic units containing the target action vocabulary in the preset behavior action dictionary, and the candidate semantic units are mapped to initial behavior action nodes. Based on the distribution density of the initial behavior action nodes on the time axis, the initial behavior action nodes within adjacent time windows are clustered and merged to generate key behavior action nodes with time aggregation. Dependence analysis is performed on the time intervals between the key behavior action nodes. If the time interval between two key behavior action nodes is less than the preset action association threshold, a temporal dependence edge is established between the two key behavior action nodes. According to the directionality of the temporal dependence edge, the sequence or concurrency relationship between the key behavior action nodes is determined to generate a set of temporal dependence relationships with temporal tags. Context consistency verification is performed on the temporal dependence edges in the set of temporal dependence relationships. If there are dependence edges that conflict with the verified temporal logic in the historical behavior pattern, the conflicting dependence edges are removed and the temporal dependence relationship is recalculated to generate optimized key behavior action nodes and the temporal dependence relationship between the nodes.

[0130] When segmenting the behavior semantic description sequence according to the preset time granularity, the preset time granularity can be flexibly set according to the actual monitoring scenario and requirements, such as in seconds, minutes, etc. For the semantic coherence threshold, it is used to measure the semantic association degree between the basic semantic units. When the semantic coherence of a basic semantic unit with other units is lower than this threshold, it can be considered redundant and should be filtered. When performing action keyword matching, the preset behavior action dictionary is a pre-constructed set containing various behavior action vocabularies. By matching the semantic units in the redundant-free semantic unit set with the target action vocabulary in this dictionary, the semantic units with key behavior actions can be accurately identified.

[0131] Clustering and merging based on the distribution density of initial behavior action nodes on the timeline can merge initial behavior action nodes that are close in time into a more representative key behavior action node. Exemplarily, in a campus surveillance scenario, if the initial behavior action node of "running" appears multiple times within a short period and they are densely distributed on the timeline, then these "running" nodes can be merged into a key behavior action node representing "continuous running".

[0132] When performing dependency analysis on the time intervals between key behavior action nodes, the preset action association threshold is the basis for determining whether there is an association between two key behavior action nodes. If the time interval between two nodes is less than this threshold, it is considered that there is a temporal dependency relationship between them, and this relationship is represented by establishing a temporal dependency edge. When determining the sequential triggering relationship or concurrent relationship, it is judged according to the directionality of the temporal dependency edge. Exemplarily, if the temporal dependency edge points from node A to node B, it means that node A triggers before node B; if there is a two-way temporal dependency edge, it means that the two nodes may be in a concurrent relationship.

[0133] When performing context consistency verification, the historical behavior pattern is a behavior logic pattern obtained by analyzing and learning a large amount of historical data. When it is found that a certain temporal dependency edge in the set of temporal dependency relationships conflicts with the verified temporal logic in the historical behavior pattern, it indicates that this dependency edge may be incorrect, and this conflicting dependency edge needs to be removed and the temporal dependency relationship needs to be recalculated to ensure the accuracy and reliability of the key behavior action nodes and the temporal dependency relationships between the nodes.

[0134] Step S480: Construct a behavior state transition graph based on the temporal dependency relationship and calculate the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule library.

[0135] The behavior state transition graph is a graph structure used to represent the changes and transition relationships of behavior states, consisting of state nodes and state transition edges. The state nodes represent key behavior action nodes, and the state transition edges represent the temporal dependency relationships between the nodes. The reference state transition graph is a graph structure predefined in the abnormal behavior rule library for representing abnormal behavior state transitions. The topological similarity is an index used to measure the similarity degree of the topological structures of two graph structures. By comparing the topological similarity between the behavior state transition graph and the reference state transition graph, it can be judged whether the current behavior matches the abnormal behavior rules.

[0136] In specific implementation, key behavioral action nodes can be first mapped to state nodes in the state transition graph, and the temporal dependency edges in the temporal dependency relationship can be mapped to state transition edges to generate an initial behavioral state transition graph. Time constraint annotation is performed on each state transition edge in the initial behavioral state transition graph. The time interval range corresponding to the temporal dependency edge in the temporal dependency relationship is extracted, and the time interval range is used as the time constraint condition for the state transition edge. According to the time constraint condition of the state transition edge and the state transition frequency in the historical behavior pattern, a transition probability weight is assigned to each state transition edge to generate a behavioral state transition graph with temporal constraints and probability attributes. A reference state transition graph matching the current monitoring scenario is extracted from the abnormal behavior rule library. The reference state transition graph includes a predefined set of abnormal state nodes and a set of abnormal state transition edges. The state nodes in the behavioral state transition graph are semantically aligned with the abnormal state nodes in the reference state transition graph to identify pairs of state nodes with the same behavioral semantic description. Based on the semantic alignment result of the state node pairs, subgraph matching is performed on the behavioral state transition graph and the reference state transition graph to extract the structural difference degree in the direction of the state transition edge, the time constraint condition, and the transition probability weight. The topological coverage degree between the behavioral state transition graph and the reference state transition graph is calculated according to the structural difference degree, and the topological similarity is generated based on the comparison result between the topological coverage degree and the preset abnormal threshold.

[0137] When mapping key behavioral action nodes to state nodes in the state transition graph, each key behavioral action node corresponds to a state node. The attributes of the state node can include information such as the name of the behavioral action and the occurrence time. When mapping the temporal dependency edges in the temporal dependency relationship to state transition edges, the attributes of the state transition edge can include time constraint conditions and transition probability weights, etc. When performing time constraint annotation on the state transition edge, by analyzing the time interval range corresponding to the temporal dependency edge in the temporal dependency relationship, the time constraint condition of the state transition edge is determined. For example, it is stipulated that a certain state transition edge must occur within a specific time range.

[0138] According to the time constraint condition of the state transition edge and the state transition frequency in the historical behavior pattern, a transition probability weight is assigned to the state transition edge. The state transition frequency in the historical behavior pattern can be obtained through statistical analysis of a large amount of historical data. Exemplarily, if the transition frequency from state A to state B is high in the historical data, a higher transition probability weight is assigned to this state transition edge.

[0139] When performing semantic alignment, the state node pairs with the same semantics are determined by comparing the behavior semantic descriptions of the state nodes in the behavior state transition graph with the abnormal state nodes in the reference state transition graph. When performing subgraph matching, the differences between the behavior state transition graph and the reference state transition graph in terms of state transition edge direction, time constraints, and transition probability weights are analyzed to calculate the structural difference. The topological coverage can be obtained by calculating the proportion of matching parts in the two graph structures. When the topological coverage exceeds the preset abnormal threshold, it is considered that the behavior state transition graph and the reference state transition graph have a high topological similarity, and the current behavior may be abnormal.

[0140] Step S490: When the topology similarity exceeds a preset threshold, an abnormal behavior type identifier matching the abnormal behavior rule base is generated.

[0141] The preset threshold is a pre-set value used to determine whether the topological similarity meets the abnormal standard. The abnormal behavior type identifier is a symbol or code used to identify the abnormal behavior type, which corresponds to a specific abnormal behavior type in the abnormal behavior rule library. When the topological similarity between the behavior state transition diagram and the reference state transition diagram exceeds the preset threshold, it means that the current behavior is highly matched with an abnormal behavior pattern in the abnormal behavior rule library. At this time, the corresponding abnormal behavior type identifier needs to be generated for subsequent processing.

[0142] In specific implementation, when the calculated topological similarity exceeds a preset threshold, the abnormal behavior rule matching the current behavior state transition diagram is searched in the abnormal behavior rule library. Each abnormal behavior rule corresponds to a specific abnormal behavior type identifier, which is extracted as the abnormal behavior type identifier of the current behavior. Exemplarily, the abnormal behavior rule library may define abnormal behavior types such as "fighting" and "vandalism", each type has a corresponding identifier, and when the current behavior matches the abnormal behavior rule of "fighting", an abnormal behavior type identifier corresponding to "fighting" is generated.

[0143] Step S4100: Generate a real-time behavior monitoring result including a behavior warning strategy according to the risk level corresponding to the abnormal behavior type identifier.

[0144] The risk level corresponding to the abnormal behavior type identifier is the risk level pre-set for each abnormal behavior type, which can be divided into different levels such as high, medium, and low. The behavior warning strategy is a response measure formulated for different abnormal behavior types and risk levels. For example, for high-risk abnormal behaviors, it may be necessary to immediately notify security personnel for on-site intervention; for low-risk abnormal behaviors, it may only be necessary to handle them through broadcast reminders and other methods. The real-time behavior monitoring result is a comprehensive result that includes information such as abnormal behavior type, risk level, and behavior warning strategy, which provides a basis for subsequent behavior management and intervention.

[0145] In specific implementation, the corresponding risk level is searched in the abnormal behavior rule library according to the abnormal behavior type identifier. The risk level for each abnormal behavior type is clearly specified in the abnormal behavior rule library. For example, "fighting" may be set as a high risk level, and "making loud noise" may be set as a low risk level. According to the risk level, the corresponding behavior warning strategy is selected from the preset behavior warning strategy library. The behavior warning strategy library contains specific intervention measures for different risk levels and abnormal behavior types. For example, for high-risk abnormal behaviors, there may be a warning strategy of "immediately notify the security personnel to rush to the scene to stop and handle"; for low-risk abnormal behaviors, there may be a warning strategy of "remind through the campus broadcast and require relevant personnel to abide by the order".

[0146] Integrate information such as the abnormal behavior type, risk level, and behavior warning strategy to generate real-time behavior monitoring results. The real-time behavior monitoring results can be presented in the form of a report for convenient viewing and processing by relevant personnel. Exemplarily, the real-time behavior monitoring results may show "Abnormal behavior type: fighting, Risk level: high, Behavior warning strategy: immediately notify the security personnel to rush to the scene to stop and handle".

[0147] Step S500: Receive the real-time behavior monitoring results through the edge computing node and adaptively adjust the acquisition parameters of the video data stream based on the dynamic priority strategy.

[0148] The edge computing node is a device with computing and data processing capabilities distributed in the campus monitoring area, responsible for receiving the real-time behavior monitoring results sent by the central server. The dynamic priority strategy is a strategy that dynamically adjusts the priority of the acquisition parameters of the video data stream according to information such as the risk level in the real-time behavior monitoring results. The acquisition parameters of the video data stream include the acquisition frame rate, resolution, and transmission bandwidth, etc. The adaptive adjustment automatically adjusts these acquisition parameters according to the real-time situation to meet different monitoring requirements.

[0149] In specific implementation, the edge computing node receives the real-time behavior monitoring results sent by the central server through the network. According to the abnormal behavior type identifier in the real-time behavior monitoring results, determine the risk level coefficient of the current monitoring area. The risk level coefficient is a quantitative indicator used to represent the risk degree of the current monitoring area, and different abnormal behavior types correspond to different risk level coefficients. Based on the risk level coefficient and the resource load status of the edge computing node, calculate the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream.

[0150] Schedule the video acquisition tasks of multiple edge computing nodes according to the priority weights, so that the edge computing nodes at the preset priority weights are allocated the preset computing resources. Synchronize the adjusted acquisition parameters to all edge computing nodes, and update the acquisition task queue of the video data stream.

[0151] As an implementation manner, in step S500, receive the real-time behavior monitoring result through the edge computing node, and adaptively adjust the acquisition parameters of the video data stream based on the dynamic priority policy, which may specifically include: step S510: Determine the risk level coefficient of the current monitoring area according to the abnormal behavior type identifier in the real-time behavior monitoring result.

[0152] The abnormal behavior type identifier is a symbol or code used to identify the type of abnormal behavior, and different types of abnormal behaviors have different risk levels. The risk level coefficient is a quantitative indicator used to represent the risk level of the current monitoring area, and it can be found in the abnormal behavior rule library according to the abnormal behavior type identifier.

[0153] Specifically, when the edge computing node receives the real-time behavior monitoring result, extract the abnormal behavior type identifier from it. Look up the risk level coefficient corresponding to the identifier in the abnormal behavior rule library. Exemplarily, the risk level coefficient of "fighting" is specified as 0.8 and the risk level coefficient of "loud noise" is 0.2 in the abnormal behavior rule library. When the abnormal behavior type identifier in the real-time behavior monitoring result is "fighting", determine that the risk level coefficient of the current monitoring area is 0.8.

[0154] Step S520: Calculate the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream based on the risk level coefficient and the resource load status of the edge computing node.

[0155] The risk level coefficient represents the risk level of the current monitoring area, and the resource load status of the edge computing node reflects the current computing and processing capabilities of the node. The priority weight is a value used to determine the importance degree of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream. By calculating the priority weights, resources can be reasonably allocated to meet different monitoring requirements.

[0156] In specific implementation, a multi-objective optimization function is constructed with resource utilization rate, behavior detection accuracy, and response delay as constraint conditions. The resource utilization rate refers to the usage of computing resources, storage resources, etc. of the edge computing node. The behavior detection accuracy refers to the accuracy of detecting and identifying behaviors. The response delay refers to the time delay from data collection to the output of the processing result. The multi-objective optimization function is solved through the gradient descent algorithm to obtain the Pareto optimal solution set of the acquisition frame rate, resolution, and transmission bandwidth. The Pareto optimal solution set is a set of solutions that, under the condition of meeting the constraint conditions, cannot improve a certain objective without reducing other objectives. According to the behavior detection accuracy and resource consumption ratio corresponding to each solution in the Pareto optimal solution set, a solution that meets the preset balance condition is selected as the priority weight. The preset balance condition can be set according to actual requirements. For example, it is required that the behavior detection accuracy reaches a certain level while the resource consumption is within an acceptable range.

[0157] As an implementation manner, step S520 may specifically include: step S521: Construct a multi-objective optimization function with resource utilization rate, behavior detection accuracy, and response delay as constraint conditions.

[0158] The resource utilization rate is the usage ratio of computing resources, storage resources, etc. of the edge computing node. The behavior detection accuracy is the ability to accurately detect and identify behaviors. The response delay is the time taken from data collection to the output of the processing result. The multi-objective optimization function is a function that comprehensively considers the resource utilization rate, behavior detection accuracy, and response delay, and its purpose is to find the optimal video data stream acquisition parameters under the premise of meeting these constraint conditions.

[0159] In specific construction, let the resource utilization rate be U, the behavior detection accuracy be A, and the response delay be D. The multi-objective optimization function F can be expressed as , where are weight coefficients, respectively representing the importance of the resource utilization rate, behavior detection accuracy, and response delay in the function, and can be adjusted according to actual requirements. are sub-functions regarding the resource utilization rate, behavior detection accuracy, and response delay respectively. For example, can be a function inversely proportional to the resource utilization rate to ensure that the resource utilization rate is within a reasonable range; can be a function directly proportional to the behavior detection accuracy to improve the accuracy of behavior detection; can be a function inversely proportional to the response delay to reduce the response delay. At the same time, it is necessary to set the constraint conditions for the resource utilization rate, behavior detection accuracy, and response delay. For example, it is required that the resource utilization rate does not exceed a certain threshold U max , the behavior detection accuracy is not lower than a certain threshold A min , and the response delay does not exceed a certain threshold D max .

[0160] Step S522: Solve the multi-objective optimization function through the gradient descent algorithm to obtain the Pareto optimal solution set of the acquisition frame rate, resolution, and transmission bandwidth.

[0161] The gradient descent algorithm is an optimization algorithm used to find the minimum value of a function. The Pareto optimal solution set is the set of solutions in a multi-objective optimization problem where it is impossible to improve one objective without degrading other objectives. By solving the multi-objective optimization function using the gradient descent algorithm, the Pareto optimal solution set that satisfies the constraint conditions can be found, thereby determining the optimal values of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream.

[0162] In specific implementation, the values of the acquisition frame rate, resolution, and transmission bandwidth can be initialized as the initial solution of the algorithm. Then, calculate the gradient of the multi-objective optimization function at the current solution. The gradient represents the rate of change of the function at this point. According to the direction of the gradient, update the values of the acquisition frame rate, resolution, and transmission bandwidth to gradually decrease the value of the multi-objective optimization function. Repeat this process until the convergence condition is met, such as the norm of the gradient being less than a preset threshold. During the solution process, it is necessary to ensure that the updated solution satisfies the constraint conditions of resource utilization rate, behavior detection accuracy, and response delay. The final set of solutions obtained is the Pareto optimal solution set of the acquisition frame rate, resolution, and transmission bandwidth.

[0163] Step S523: Select the solution that meets the preset balance condition as the priority weight according to the ratio of behavior detection accuracy to resource consumption corresponding to each solution in the Pareto optimal solution set.

[0164] The ratio of behavior detection accuracy to resource consumption is the ratio of behavior detection accuracy to resource consumption under a certain solution, which reflects the behavior detection accuracy that can be achieved when consuming a certain amount of resources. The preset balance condition is a condition set in advance for screening out solutions that satisfy the balance between behavior detection accuracy and resource consumption. The priority weight is determined according to the selected solution and is used to adjust the acquisition frame rate, resolution, and transmission bandwidth of the video data stream. In specific implementation, traverse each solution in the Pareto optimal solution set and calculate the corresponding ratio of behavior detection accuracy to resource consumption. The preset balance condition can be that the ratio of behavior detection accuracy to resource consumption reaches a certain threshold, or that the behavior detection accuracy and resource consumption are kept balanced within a certain range. Select the solution that meets the preset balance condition as the priority weight. Exemplarily, if the preset balance condition is that the ratio of behavior detection accuracy to resource consumption is not less than 0.8, then select the solution in the Pareto optimal solution set whose ratio of behavior detection accuracy to resource consumption is not less than 0.8, and use the corresponding values of the acquisition frame rate, resolution, and transmission bandwidth as the priority weight.

[0165] Step S530: Schedule the video capture tasks of multiple edge computing nodes according to the priority weights, and allocate preset computing resources to the edge computing nodes with preset priority weights.

[0166] The priority weights represent the importance levels of the capture frame rate, resolution, and transmission bandwidth of the video data stream. By scheduling the video capture tasks of multiple edge computing nodes according to the priority weights, computing resources can be reasonably allocated to ensure clearer and more accurate video data can be obtained in high-risk areas. The preset priority weights are the weight values preset for distinguishing the priorities of different edge computing nodes, and the preset computing resources are the amounts of computing resources allocated to the edge computing nodes according to the priority weights.

[0167] In specific implementation, sort the multiple edge computing nodes according to the priority weights, and the edge computing nodes with higher priorities are ranked in the front. For the edge computing nodes within the preset priority weight range, allocate preset computing resources to them. Exemplarily, if the preset priority weights stipulate that the top 30% of the edge computing nodes are high-priority nodes, then more computing resources will be allocated to these 30% of the edge computing nodes, such as higher CPU usage rates, larger memory spaces, etc., to ensure that they can perform video data capture at higher capture frame rates, resolutions, and transmission bandwidths.

[0168] Step S540: Synchronize the adjusted capture parameters to all edge computing nodes and update the capture task queue of the video data stream.

[0169] The adjusted capture parameters are the parameters such as the capture frame rate, resolution, and transmission bandwidth of the video data stream determined according to the priority weights and resource allocation situations. Synchronization is to send the adjusted capture parameters to all edge computing nodes so that they can perform video data capture according to the new parameters. Updating the capture task queue of the video data stream is to rearrange the video capture tasks of the edge computing nodes according to the adjusted capture parameters to ensure the smooth progress of the capture tasks. In specific implementation, the adjusted capture parameters can be sent to all edge computing nodes through the network. After receiving the adjusted capture parameters, the edge computing nodes update their own capture configurations and perform video data capture according to the new capture frame rate, resolution, and transmission bandwidth. At the same time, update the capture task queue of the video data stream and rearrange the execution order and time of the capture tasks according to the new capture parameters. Exemplarily, if the capture frame rate of an edge computing node increases, then the time interval of its capture task needs to be adjusted accordingly to ensure that the capture can be performed according to the new frame rate.

[0170] As an implementation manner, the method provided by the present invention may further include the following derivative steps: Step S600: Construct a behavior pattern evolution graph in the central server, where the behavior pattern evolution graph includes historical behavior pattern nodes, real-time behavior pattern nodes, and pattern transition edges.

[0171] The behavior pattern evolution graph is a graph structure used to represent the evolution of behavior patterns over time, which can help analyze the change trends and rules of behavior patterns. The historical behavior pattern nodes are the nodes of the behavior patterns that occurred in the past period of time, and each node represents a specific behavior pattern. The real-time behavior pattern nodes are the nodes of the behavior patterns that occur at the current moment. The pattern transition edges are the edges connecting the historical behavior pattern nodes and the real-time behavior pattern nodes, which represent the transition relationships between the behavior patterns.

[0172] Specifically, when constructing, different behavior patterns can be extracted from the historical behavior data and used as historical behavior pattern nodes. Each historical behavior pattern node can be represented by a feature vector, which contains information such as the spatio-temporal features and behavior semantics of the behavior pattern. For the real-time behavior pattern nodes, the current behavior pattern is extracted by real-time monitoring and analysis of the current behavior data and used as the real-time behavior pattern node. Then, the relationship between the historical behavior pattern nodes and the real-time behavior pattern nodes is analyzed to determine the pattern transition edges between them. The pattern transition edges can have weights, and the weights represent the probability or frequency of the behavior pattern transition. Exemplarily, if the frequency of transferring from the historical behavior pattern A to the real-time behavior pattern B is relatively high, the weight of the pattern transition edge A->B can be set to a relatively large value.

[0173] Step S700: Update the real-time behavior pattern nodes according to the real-time behavior monitoring results, and calculate the pattern similarity between the real-time behavior pattern nodes and the historical behavior pattern nodes.

[0174] The real-time behavior monitoring results contain the behavior information at the current moment, and the features of the real-time behavior pattern nodes can be updated according to this information. The pattern similarity is an index used to measure the similarity degree between the real-time behavior pattern nodes and the historical behavior pattern nodes. By calculating the pattern similarity, the differences and connections between the current behavior pattern and the historical behavior pattern can be understood.

[0175] Specifically, when implementing, the feature vector of the real-time behavior pattern nodes is updated according to the information such as the behavior semantic description sequence and key behavior action nodes in the real-time behavior monitoring results. Multiple methods can be used to calculate the pattern similarity between the real-time behavior pattern nodes and the historical behavior pattern nodes, such as cosine similarity, Euclidean distance, etc. Taking cosine similarity as an example, let the feature vector of the real-time behavior pattern node be X and the feature vector of the historical behavior pattern node be Y, then the cosine similarity S between them can be expressed as , where represents the dot product of vectors X and Y, and represent the magnitudes of vectors X and Y respectively. By calculating the pattern similarity between the real-time behavior pattern node and all historical behavior pattern nodes, a similarity list can be obtained.

[0176] Step S800: When it is detected that the pattern similarity is continuously lower than the preset threshold, trigger the incremental learning process of the target behavior recognition model.

[0177] The preset threshold is a numerically value set in advance for judging whether the pattern similarity is too low. When the pattern similarity between the real-time behavior pattern node and the historical behavior pattern node is continuously lower than the preset threshold, it indicates that there are significant differences between the current behavior pattern and the historical behavior patterns, and the target behavior recognition model may not be able to accurately recognize these new behavior patterns. At this time, it is necessary to trigger the incremental learning process to update and optimize the target behavior recognition model.

[0178] Among them, the incremental learning process can specifically include: Step S810: Collect the video data stream segment corresponding to the current real-time behavior pattern node.

[0179] The video data stream segment corresponding to the real-time behavior pattern node is a video data segment containing the current real-time behavior pattern. These segments record the detailed information of the newly emerged behavior patterns and are important data sources for incremental learning.

[0180] In specific implementation, according to the timestamp information of the real-time behavior pattern node, extract the corresponding video data segment from the video data stream collected by the edge computing node. The time period containing the real-time behavior pattern can be located by querying the timestamp index of the video data, and then the video data stream segment within this time period is intercepted. Exemplarily, if the time range of the real-time behavior pattern node is from the 100th frame to the 200th frame, then the segment from the 100th frame to the 200th frame in the video data stream is intercepted as the video data stream segment corresponding to the current real-time behavior pattern node.

[0181] Step S820: Re-extract the features of the video data stream segment to generate an incremental training sample set.

[0182] Feature re-extraction is to re-extract the features of the collected video data stream segment to obtain more accurate and representative features. The incremental training sample set is a set composed of samples obtained after feature re-extraction, and these samples will be used for incremental learning of the target behavior recognition model.

[0183] In specific implementation, the same method as in step S100 is used to extract features from the video data stream segment. That is, the video data stream segment is segmented into a frame sequence and spatio-temporal features are extracted to generate an initial set of behavior features. Then, according to the method of step S200, multi-dimensional correlation analysis is performed on the initial set of behavior features to extract spatio-temporal behavior feature vectors corresponding to the target monitoring scenario. The extracted spatio-temporal behavior feature vectors are used as samples in the incremental training sample set. Each sample can be associated with a corresponding behavior label, and the behavior label can be determined according to the sequence of behavior semantic descriptions in the real-time behavior monitoring results.

[0184] Step S830: Fine-tune the parameters of the target behavior recognition model based on the incremental training sample set, and update the weights of the pattern transition edges in the behavior pattern evolution graph.

[0185] Parameter fine-tuning is to make minor adjustments to the parameters of the target behavior recognition model based on the incremental training sample set so that it can better recognize newly emerging behavior patterns. Updating the weights of the pattern transition edges in the behavior pattern evolution graph is to adjust the weights of the pattern transition edges according to the relationship between the newly emerging behavior pattern and the historical behavior patterns to reflect the evolution of the behavior patterns.

[0186] In specific implementation, the incremental training sample set is input into the target behavior recognition model, and an optimization algorithm, such as the stochastic gradient descent algorithm, is used to fine-tune the parameters of the model. During the fine-tuning process, the loss function of the model is calculated, and the parameters of the model are updated according to the gradient of the loss function to make the output of the model closer to the true label of the sample. For the weights of the pattern transition edges in the behavior pattern evolution graph, the weights of the pattern transition edges are adjusted according to the pattern similarity and transition frequency between the newly emerging real-time behavior pattern node and the historical behavior pattern nodes. Exemplarily, if the similarity between the newly emerging real-time behavior pattern and a certain historical behavior pattern is high and the transition frequency increases, the weight of the pattern transition edge between these two nodes is increased accordingly. Through parameter fine-tuning, the target behavior recognition model can better adapt to newly emerging behavior patterns and improve the accuracy of behavior recognition; by updating the weights of the pattern transition edges, the behavior pattern evolution graph can more accurately reflect the evolution trend of the behavior patterns.

[0187] Please refer to Figure 2 , Figure 2It is a schematic diagram of the composition of a smart campus security warning system provided by an embodiment of the present invention. The system includes an edge computing node 100 and a central server 300 that are communicatively connected through a network 200 to each other. Among them, the number of edge computing nodes 100 is one or more. Computer programs are respectively stored in the memories of the edge computing node 100 and the central server 300; when the processors of the edge computing node 100 and the central server 300 respectively load and execute the corresponding computer programs, the smart campus security warning method based on edge computing and big data provided by the embodiment of the present invention is implemented.

Claims

1. A smart campus security warning method based on edge computing and big data, characterized in that, The method includes: collecting multiple video data streams in real time through edge computing nodes deployed in the campus monitoring area, performing frame sequence segmentation and spatio-temporal feature extraction on the video data streams to generate an initial behavior feature set; performing multi-dimensional correlation analysis on the initial behavior feature set based on preset behavior semantic tags, extracting spatio-temporal behavior feature vectors corresponding to the target monitoring scenario, and transmitting the spatio-temporal behavior feature vectors to the central server; in the central server, performing multi-stage training on the initial behavior recognition model according to the historical behavior data set to obtain a target behavior recognition model, where the multi-stage training includes feature fusion based on spatio-temporal correlation and cross-scene behavior pattern migration; inputting the spatio-temporal behavior feature vectors into the target behavior recognition model to generate a behavior semantic description sequence corresponding to the video data stream, and determining the real-time behavior monitoring result according to the matching result between the behavior semantic description sequence and the preset abnormal behavior rule library; receiving the real-time behavior monitoring result through the edge computing node, and adaptively adjusting the acquisition parameters of the video data stream based on the dynamic priority policy.

2. The method according to claim 1, wherein The performing frame sequence segmentation and spatio-temporal feature extraction on the video data stream to generate an initial behavior feature set includes: performing frame splitting on the video data stream to obtain a continuous video frame sequence, and performing noise filtering on the video frame sequence based on the pixel change rate between adjacent video frames; performing dynamic area detection on the filtered video frame sequence to determine a local image area containing a moving target, and performing multi-scale spatio-temporal convolution processing on the local image area; extracting the motion trajectory feature, pose change feature and environment interaction feature of the local image area after the multi-scale spatio-temporal convolution processing, and aligning the motion trajectory feature, pose change feature and environment interaction feature along the time axis; performing weighted fusion on the aligned features according to the preset spatio-temporal weight distribution matrix to generate the initial behavior feature set.

3. The method according to claim 2, characterized in that, Extract the motion trajectory features, pose change features, and environmental interaction features of the local image region after the multi-scale spatio-temporal convolution process, and align the motion trajectory features, pose change features, and environmental interaction features along the time axis, including: dividing the local image region after the multi-scale spatio-temporal convolution process into time segments to generate a set of time segments containing start timestamps and end timestamps; performing trajectory tracking processing on each time segment in the set of time segments to obtain a continuous displacement coordinate sequence of the moving target within the local image region, and generating the motion trajectory features based on the vector direction change rate of adjacent coordinates in the displacement coordinate sequence; performing key point detection processing on each time segment in the set of time segments to extract the spatial position sequence of the moving target at the skeletal joint points, and generating the pose change features based on the relative displacement amount between adjacent time segments of the spatial position sequence; performing target relationship analysis processing on each time segment in the set of time segments to identify the spatial distance sequence and contact state change sequence between the moving target and surrounding static objects, and generating the environmental interaction features based on the coupling relationship between the spatial distance sequence and the contact state change sequence; synchronizing the timestamps of the motion trajectory features, pose change features, and environmental interaction features according to the start timestamps and end timestamps corresponding to each time segment in the set of time segments to generate a time-aligned version of the motion trajectory features, pose change features, and environmental interaction features with a unified time reference; based on the time continuity of the motion trajectory features, pose change features, and environmental interaction features in the time-aligned version, performing adaptive trajectory interpolation compensation on the feature data corresponding to the missing timestamps to generate the motion trajectory features, pose change features, and environmental interaction features after complete time-axis alignment.

4. The method according to claim 1, characterized in that Performing multi-stage training on the initial behavior recognition model according to the historical behavior data set to obtain the target behavior recognition model, including: extracting a first training set with spatio-temporal annotations and a second training set of unannotated cross-scene data from the historical behavior data set, wherein each sample in the first training set contains an annotated spatio-temporal behavior feature vector and a corresponding normal behavior type label, and each sample in the second training set contains an unannotated video segment under different monitoring scenarios; inputting the first training set into the initial behavior recognition model, and performing cross-frame feature alignment on the spatio-temporal behavior feature vector through a spatio-temporal correlation constraint module to generate a first intermediate feature representation with time continuity; inputting the first intermediate feature representation into a spatio-temporal feature fusion module, performing dynamic weight assignment and feature superposition on the features within adjacent time windows to generate a fused second intermediate feature representation; performing supervised training on the initial behavior recognition model based on the similarity loss between the second intermediate feature representation and the normal behavior type label to obtain a first intermediate model; inputting the second training set into the first intermediate model, and extracting scene-invariant features from the unannotated video segment through a cross-scene behavior pattern migration module to generate a third intermediate feature representation independent of the monitoring scenario; performing temporal dimension compression and feature reconstruction on the third intermediate feature representation to generate a cross-scene behavior pattern embedding vector, and performing self-supervised training on the first intermediate model based on the reconstruction error to obtain a second intermediate model; alternately weighted fusing the feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model to generate a target behavior recognition model with cross-scene adaptability.

5. The method according to claim 4, wherein Performing cross-frame feature alignment on the spatio-temporal behavior feature vector through a spatio-temporal correlation constraint module to generate a first intermediate feature representation with time continuity, including: extracting a feature subsequence corresponding to a continuous time segment from the spatio-temporal behavior feature vector, and calculating the dynamic time warping distance between adjacent time segment feature subsequences; performing non-linear interpolation processing on the feature subsequence according to the dynamic time warping distance to generate a fourth intermediate feature representation with time axis alignment; performing time dimension slicing on the fourth intermediate feature representation to generate local feature blocks within multiple overlapping time windows; assigning learnable spatio-temporal attention weights to each local feature block, and enhancing the features of the local feature block based on the spatio-temporal attention weights to generate a fifth intermediate feature representation; inputting the fifth intermediate feature representation into a bidirectional temporal convolutional network, capturing future context dependencies through forward propagation, and capturing historical context dependencies through backward propagation to generate a sixth intermediate feature representation with bidirectional time correlation; performing residual connection on the sixth intermediate feature representation and the original spatio-temporal behavior feature vector to generate the first intermediate feature representation.

6. The method according to claim 1, wherein Determining the real-time behavior monitoring result according to the matching result between the behavior semantic description sequence and the preset abnormal behavior rule library includes: performing semantic parsing on the behavior semantic description sequence to extract key behavior action nodes and the temporal dependency relationship between the nodes; constructing a behavior state transition graph according to the temporal dependency relationship, and calculating the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule library; when the topological similarity exceeds a preset threshold, generating an abnormal behavior type identifier matching the abnormal behavior rule library; generating the real-time behavior monitoring result including a behavior warning strategy according to the risk level corresponding to the abnormal behavior type identifier.

7. The method according to claim 6, wherein The performing semantic parsing on the behavior semantic description sequence to extract key behavior action nodes and the temporal dependency relationship between the nodes includes: dividing the behavior semantic description sequence into multiple consecutive basic semantic units according to a preset time granularity, and performing redundant filtering on the basic semantic units based on a semantic coherence threshold to generate a set of redundant-free semantic units; performing action keyword matching on each semantic unit in the set of redundant-free semantic units, identifying candidate semantic units containing target action words in a preset behavior action dictionary, and mapping the candidate semantic units to initial behavior action nodes; clustering and merging the initial behavior action nodes in adjacent time windows based on the distribution density of the initial behavior action nodes on the time axis to generate key behavior action nodes with time aggregation; performing dependency analysis on the time intervals between the key behavior action nodes, and if the time interval between two key behavior action nodes is less than a preset action association threshold, establishing a temporal dependency edge between the two key behavior action nodes; determining the sequential triggering relationship or concurrent relationship between the key behavior action nodes according to the directionality of the temporal dependency edge to generate a set of temporal dependency relationships with temporal tags; performing context consistency verification on the temporal dependency edges in the set of temporal dependency relationships, and if there are dependency edges conflicting with the verified temporal logic in the historical behavior pattern, removing the conflicting dependency edges and recalculating the temporal dependency relationship to generate optimized key behavior action nodes and the temporal dependency relationship between the nodes.

8. The method according to claim 6, wherein Constructing a behavioral state transition graph according to the temporal dependency relationship and calculating the topological similarity between the behavioral state transition graph and the reference state transition graph in the abnormal behavior rule base includes: mapping the key behavioral action nodes to state nodes in the state transition graph, and mapping the temporal dependency edges in the temporal dependency relationship to state transition edges to generate an initial behavioral state transition graph; performing time constraint annotation on each state transition edge in the initial behavioral state transition graph, extracting the time interval range corresponding to the temporal dependency edge in the temporal dependency relationship, and using the time interval range as the time constraint condition for the state transition edge; assigning a transition probability weight to each state transition edge according to the time constraint condition of the state transition edge and the state transition frequency in the historical behavior pattern to generate a behavioral state transition graph with temporal constraints and probability attributes; extracting a reference state transition graph matching the current monitoring scenario from the abnormal behavior rule base, where the reference state transition graph includes a predefined set of abnormal state nodes and a set of abnormal state transition edges; semantically aligning the state nodes in the behavioral state transition graph with the abnormal state nodes in the reference state transition graph to identify pairs of state nodes with the same behavioral semantic description; based on the semantic alignment result of the state node pairs, performing subgraph matching on the behavioral state transition graph and the reference state transition graph, and extracting the structural difference degree between them in terms of the direction of the state transition edge, the time constraint condition, and the transition probability weight; calculating the topological coverage degree between the behavioral state transition graph and the reference state transition graph according to the structural difference degree, and generating the topological similarity based on the comparison result between the topological coverage degree and a preset abnormal threshold.

9. The method according to claim 1, wherein Receiving the real-time behavior monitoring result by the edge computing node and adaptively adjusting the acquisition parameters of the video data stream based on a dynamic priority policy includes: determining the risk level coefficient of the current monitoring area according to the abnormal behavior type identifier in the real-time behavior monitoring result; calculating the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream based on the risk level coefficient and the resource load status of the edge computing node; scheduling the video acquisition tasks of multiple edge computing nodes according to the priority weights, so that the edge computing nodes at the preset priority weights are allocated preset computing resources; synchronizing the adjusted acquisition parameters to all edge computing nodes and updating the acquisition task queue of the video data stream; where calculating the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream includes: constructing a multi-objective optimization function with resource utilization rate, behavior detection accuracy, and response delay as constraint conditions; solving the multi-objective optimization function by a gradient descent algorithm to obtain the Pareto optimal solution set of the acquisition frame rate, resolution, and transmission bandwidth; and selecting the solution that meets the preset balance condition as the priority weight according to the behavior detection accuracy and resource consumption ratio corresponding to each solution in the Pareto optimal solution set.

10. A smart campus security warning system, characterized in that, It includes an edge computing node and a central server that are communicatively connected to each other, and computer programs are respectively stored in the memories of the edge computing node and the central server; when the processors of the edge computing node and the central server respectively load and execute the corresponding computer programs, the intelligent campus security early warning method based on edge computing and big data as described in any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Campus abnormal behavior analysis system and method

    CN115100572A

  • Method for training video label recommendation model, and method for determining video label

    WO2023273769A1

Cited By

  • Multi-source pressure equipment safety monitoring system based on image fusion analysis

    CN120561827A

  • Video monitoring method for intelligent medical patient nursing

    CN120881240A

  • A video monitoring method for smart medical patient care

    CN120881240B

  • Video monitoring intelligent analysis method and system based on edge calculation

    CN120913156A

  • Smart park energy data dynamic management and control system

    CN121169334A