Intelligent campus safety early warning method and system based on edge computing and big data

By combining edge computing and big data, a smart campus security early warning method is used to collect and extract spatiotemporal behavioral features of video data streams in real time. Combined with a multi-stage training model on a central server, a sequence of behavioral semantic descriptions is generated, and the collection parameters are dynamically adjusted. This solves the problems of real-time performance, adaptability, and resource management in existing technologies, and achieves efficient behavior monitoring.

CN120339957BActive Publication Date: 2026-04-14GUANGDONG SANZHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG SANZHU TECH CO LTD
Filing Date
2025-04-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing campus behavior monitoring systems suffer from high latency in video data transmission due to their centralized processing architecture, making it difficult to meet real-time monitoring needs; fixed models cannot adapt to changes in behavior patterns across multiple scenarios, leading to decreased detection accuracy; behavior recognition based on low-level visual features lacks semantic relevance and is susceptible to environmental interference, resulting in false positives and false negatives; and static resource configuration modes cannot be dynamically adjusted according to real-time risk levels, leading to resource waste or idleness.

Method used

Video data streams are acquired in real time via edge computing nodes, and frame sequence segmentation and spatiotemporal feature extraction are performed to generate an initial set of behavioral features. Based on preset behavioral semantic labels, multi-dimensional correlation analysis is conducted to extract spatiotemporal behavioral feature vectors. Cross-scene behavioral pattern transfer is performed by combining a multi-stage training model on a central server to generate behavioral semantic description sequences, and the acquisition parameters of the video data stream are adjusted through a dynamic priority strategy.

Benefits of technology

It achieves a synergistic improvement in multi-dimensional behavior monitoring capabilities, reduces video data processing latency, optimizes computing resource allocation, enhances the model's adaptability to complex campus scenarios, improves the accuracy and interpretability of abnormal behavior detection, dynamically adjusts resource consumption to avoid resource waste or missed detections, and forms a continuously iterative learning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339957B_ABST
    Figure CN120339957B_ABST
Patent Text Reader

Abstract

The application provides a kind of based on edge computing and big data's wisdom campus safety early warning method and system. Through edge computing node, real-time collection is carried out to multiple video data streams, frame sequence segmentation and space-time feature extraction are carried out, initial behavior feature set is generated, initial behavior feature set is carried out multidimensional correlation analysis based on preset behavior semantic label, space-time behavior feature vector corresponding to target monitoring scene is extracted, and is transmitted to central server, space-time behavior feature vector is input into target behavior identification model, behavior semantic description sequence corresponding to video data stream is generated, and according to its matching result with preset abnormal behavior rule library, real-time behavior monitoring result is determined, real-time behavior monitoring result is received by edge computing node, and based on dynamic priority strategy, the collection parameters of video data stream are adaptively adjusted.The application can improve the real-time performance, accuracy and system sustainability of campus behavior monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically, to a smart campus security early warning method and system based on edge computing and big data. Background Technology

[0002] With the increasing demand for campus security management, data-driven behavior monitoring technology has become an important means of ensuring campus safety. Existing behavior monitoring systems typically employ a centralized video processing architecture, using cloud servers to perform unified behavior recognition and anomaly detection on surveillance videos, relying on fixed feature extraction models and preset rule bases for behavior classification. However, this approach has significant drawbacks: centralized processing architecture leads to high video data transmission latency, making it difficult to meet real-time monitoring needs; fixed models cannot adapt to the dynamic changes in behavior patterns across multiple campus scenarios (such as classrooms, playgrounds, and corridors), resulting in decreased detection accuracy across different scenarios; behavior recognition based on low-level visual features lacks semantic relevance and is susceptible to environmental interference, leading to false positives and false negatives; simultaneously, static resource configuration cannot dynamically adjust acquisition parameters according to real-time risk levels, resulting in resource contention during high-load scenarios and resource idleness during low-load scenarios. These problems severely restrict the practicality and scalability of campus behavior monitoring systems, necessitating a new monitoring solution that integrates the real-time performance of edge computing with the global scope of big data analysis, possessing both scene adaptability and deep semantic understanding. Summary of the Invention

[0003] This invention provides a smart campus security early warning method and system based on edge computing and big data.

[0004] In a first aspect, embodiments of the present invention provide a smart campus security early warning method based on edge computing and big data. The method includes: real-time acquisition of multiple video data streams by edge computing nodes deployed in the campus monitoring area, and frame sequence segmentation and spatiotemporal feature extraction of the video data streams to generate an initial behavioral feature set; multi-dimensional correlation analysis of the initial behavioral feature set based on preset behavioral semantic tags to extract spatiotemporal behavioral feature vectors corresponding to the target monitoring scene, and transmission of the spatiotemporal behavioral feature vectors to a central server; in the central server, multi-stage training of the initial behavioral recognition model based on historical behavioral data sets to obtain a target behavioral recognition model, wherein the multi-stage training includes feature fusion based on spatiotemporal correlation and cross-scene behavioral pattern transfer; inputting the spatiotemporal behavioral feature vectors into the target behavioral recognition model to generate a behavioral semantic description sequence corresponding to the video data streams, and determining real-time behavioral monitoring results based on the matching results of the behavioral semantic description sequence and a preset abnormal behavior rule base; receiving the real-time behavioral monitoring results through the edge computing nodes, and adaptively adjusting the acquisition parameters of the video data streams based on a dynamic priority strategy.

[0005] Secondly, embodiments of the present invention provide a smart campus security early warning system, including edge computing nodes and a central server that are interconnected. The memory of the edge computing nodes and the central server respectively stores computer programs. When the processors of the edge computing nodes and the central server respectively load and execute the corresponding computer programs, the smart campus security early warning method based on edge computing and big data as described above is realized.

[0006] This invention provides a smart campus security early warning method based on edge computing and big data. It collects and extracts spatiotemporal behavioral feature vectors from video data streams in real time through edge computing nodes, combines this with a multi-stage training model on a central server to generate behavioral semantic description sequences, and dynamically optimizes the collection parameters, achieving a synergistic improvement in multi-dimensional behavior monitoring capabilities. This method, through a hierarchical processing architecture of edge computing and central analysis, not only effectively reduces video data processing latency and optimizes computing resource allocation, but also enhances the model's adaptability to complex campus scenarios through spatiotemporal feature fusion and cross-scenario behavior pattern transfer. Simultaneously, it utilizes a semantic rule matching mechanism to improve the accuracy and interpretability of abnormal behavior detection. Furthermore, an adaptive parameter adjustment mechanism based on a dynamic priority strategy can intelligently balance detection accuracy and resource consumption according to real-time risk levels, avoiding resource waste or missed detection problems under traditional fixed parameter configurations. In addition, the closed-loop feedback of semantic description sequence generation and collection parameter optimization forms a continuously iterative learning system, gradually improving the generalization ability and scenario coverage of behavior monitoring without the need for manual early warning, thereby comprehensively improving the real-time performance, accuracy, and system sustainability of campus behavior monitoring. Attached Figure Description

[0007] Figure 1 This is a flowchart of a smart campus security early warning method based on edge computing and big data provided in an embodiment of the present invention;

[0008] Figure 2 This is a schematic diagram of the composition of a smart campus security early warning system provided in an embodiment of the present invention. Detailed Implementation

[0009] Please see Figure 1 , Figure 1 The flowchart of a smart campus security early warning method based on edge computing and big data provided in the embodiment of the present invention includes the following steps: Step S100: Real-time acquisition of multiple video data streams by edge computing nodes deployed in the campus monitoring area, and frame sequence segmentation and spatiotemporal feature extraction of the video data streams to generate an initial behavioral feature set.

[0010] In this embodiment of the invention, edge computing nodes are devices with computing and data processing capabilities distributed throughout the campus monitoring area. Located close to the data source, they can perform preliminary processing of the collected data at the edge, reducing data transmission pressure. The video data stream is a sequence of video signals continuously collected and transmitted by surveillance cameras. Frame sequence segmentation divides the continuous video data stream into independent video frames according to time sequence, allowing for individual processing of each frame later. Spatiotemporal feature extraction extracts time- and space-related feature information from the segmented video frames, such as the trajectory of objects and changes in posture. The initial behavioral feature set is a collection obtained by integrating the extracted spatiotemporal features, containing basic feature information of various behaviors presented in the video.

[0011] An exemplary implementation process is as follows: Edge computing nodes utilize their built-in video acquisition modules to acquire multiple video data streams from cameras in various monitoring areas of the campus. These video data streams exist in the form of continuous video signals, containing information about various scenes and activities within the campus. Then, the acquired video data streams are segmented into frame sequences, using a frame-segmentation algorithm to divide the continuous video signal into a series of discrete video frames. For example, segmentation can be performed according to the video's frame rate at fixed time intervals, dividing the video signal per second into several independent video frames. Next, spatiotemporal feature extraction is performed on the segmented video frames. This process can be implemented using computer vision techniques, such as using a convolutional neural network (CNN) to process the video frames and extract spatiotemporal features such as the object's motion trajectory and posture changes. Finally, the extracted spatiotemporal features are integrated to generate an initial behavioral feature set. This set can be represented by a multi-dimensional vector, with each dimension corresponding to a specific spatiotemporal feature.

[0012] In one implementation, step S100 involves segmenting the video data stream into frames and extracting spatiotemporal features to generate an initial set of behavioral features. Specifically, step S110 involves performing frame segmentation on the video data stream to obtain a continuous video frame sequence, and filtering noise from the video frame sequence based on the pixel change rate between adjacent video frames.

[0013] Frame segmentation is the process of dividing a continuous video data stream into individual video frames in chronological order. Through frame segmentation, video data can be converted into a discrete data format suitable for computer processing. A continuous video frame sequence is a series of video frames arranged in chronological order after frame segmentation. The pixel change rate between adjacent video frames refers to the degree of change in attributes such as color and brightness of corresponding pixels in two adjacent video frames. Noise filtering removes noise interference from the video frame sequence caused by various factors (such as camera shake, lighting changes, etc.), making the video frames clearer and more accurate.

[0014] In practical implementation, a video decoding algorithm can be used to segment the video data stream into frames, dividing the video signal per second into several independent video frames according to the frame rate. For example, for a video with a frame rate of 25 frames per second, the video signal per second will be divided into 25 independent video frames. Then, the pixel change rate between adjacent video frames is calculated. This can be done by comparing the color or brightness values ​​of corresponding pixels in adjacent video frames, calculating the difference between them, and using the average of these differences as the pixel change rate. Pixels with a pixel change rate less than a preset threshold can be considered noise points, and their values ​​are set to the average of adjacent pixels or other suitable values, thereby achieving noise filtering.

[0015] Step S120: Perform dynamic region detection on the filtered video frame sequence to determine the local image regions containing moving targets, and perform multi-scale spatiotemporal convolution processing on the local image regions.

[0016] Dynamic region detection is the process of identifying regions containing moving targets within a video frame sequence. Moving targets are objects that change position or pose in the video, such as pedestrians or vehicles. A local image region is a specific area containing a moving target; it is a portion extracted from the entire video frame. Multi-scale spatiotemporal convolution processing involves performing convolution operations at different scales on the local image region to extract spatiotemporal features at different scales.

[0017] When performing dynamic region detection, a background subtraction algorithm can be used to compare the current video frame with a pre-established background model, identifying regions with significant differences from the background model. These regions are considered dynamic regions containing moving targets. For example, by calculating the difference between corresponding pixels in the current video frame and the background model, pixels with differences greater than a preset threshold are marked as moving target regions. After identifying local image regions, multi-scale spatiotemporal convolution processing is performed. A convolutional neural network (CNN) is used to perform convolution operations at different scales on the local image regions. For example, different sizes of convolution kernels, such as 3*3, 5*5, and 7*7, can be used to convolve the local image regions to extract spatiotemporal features at different scales.

[0018] Step S130: Extract motion trajectory features, pose change features, and environmental interaction features of the local image region after multi-scale spatiotemporal convolution processing, and align the motion trajectory features, pose change features, and environmental interaction features along the time axis.

[0019] Motion trajectory features are the motion path and velocity characteristics of a moving target within a video frame sequence. Posture change features are the changes in the moving target's body posture within the video frame sequence, such as standing, walking, and bending. Environmental interaction features are the interactions between the moving target and objects in its surrounding environment, such as contact and collisions between a person and objects. Timeline alignment arranges the extracted motion trajectory features, posture change features, and environmental interaction features in chronological order, ensuring temporal consistency.

[0020] In practical implementation, feature extraction is performed on local image regions after multi-scale spatiotemporal convolution processing. For motion trajectory features, target tracking algorithms can be used to track the position of moving targets in the video frame sequence, recording information such as their motion path and speed. For pose change features, human pose estimation algorithms can be used to detect the skeletal joints of moving targets and analyze their pose changes. For environmental interaction features, target detection and relationship analysis algorithms can be used to identify the interaction relationships between moving targets and objects in the surrounding environment. Then, the extracted motion trajectory features, pose change features, and environmental interaction features are time-axis aligned. These features can be arranged in chronological order according to the timestamps of the video frames to ensure temporal consistency.

[0021] As one implementation, step S130 may specifically include the following steps: Step S131: Divide the local image region after multi-scale spatiotemporal convolution processing into time segments to generate a time segment set containing a start timestamp and an end timestamp.

[0022] Temporal segmentation involves dividing a local image region, after multi-scale spatiotemporal convolution processing, into several time segments in chronological order. The start timestamp is the beginning time of each time segment, and the end timestamp is the end time of each time segment. The time segment set is a collection of all the resulting time segments.

[0023] When dividing a video frame into time segments, the video can be divided at fixed time intervals based on the timestamps of the video frames. For example, the video frame sequence can be divided into segments of one per second, with each segment containing a certain number of video frames. For each time segment, its start and end timestamps are recorded to generate a set of time segments.

[0024] Step S132: Perform trajectory tracking processing on each time segment in the time segment set to obtain a continuous displacement coordinate sequence of the moving target in the local image region, and generate motion trajectory features based on the rate of change of the vector direction of adjacent coordinates in the displacement coordinate sequence.

[0025] Trajectory tracking processing involves tracking the position of a moving target within each time segment and recording its continuous displacement coordinates. A continuous displacement coordinate sequence is a sequence of the moving target's continuous displacement coordinates within each time segment. The rate of change of vector direction is the degree of change in the vector direction between adjacent displacement coordinates. Motion trajectory features are feature information generated based on the rate of change of vector direction of adjacent coordinates in the displacement coordinate sequence; they reflect the moving target's trajectory and velocity changes.

[0026] In practical implementation, a target tracking algorithm is used to track the trajectory of each time segment in the time segment set. For example, algorithms such as Kalman filters or particle filters can be used to predict and update the position of the moving target, obtaining its continuous displacement coordinate sequence. Then, the rate of change of the vector direction of adjacent coordinates in the displacement coordinate sequence is calculated. The rate of change of the vector direction can be determined by calculating the change in the angle between adjacent displacement coordinates. Motion trajectory features are generated based on the rate of change of the vector direction; for example, the average value and standard deviation of the rate of change of the vector direction can be used as part of the motion trajectory features.

[0027] Step S133: Perform key point detection processing on each time segment in the time segment set, extract the spatial position sequence of the moving target on the skeletal joints, and generate posture change features based on the relative displacement of the spatial position sequence between adjacent time segments.

[0028] Keypoint detection processing involves detecting the skeletal joints of a moving target within each time segment to determine their spatial positions. Skeletal joints are joints in the human body or other objects, such as the shoulders, elbows, and knees. The spatial position sequence is a sequence of the spatial positions of the skeletal joints of the moving target within each time segment. Relative displacement is the change in the spatial position of skeletal joints between adjacent time segments. Posture change features are feature information generated based on the relative displacement of the spatial position sequence between adjacent time segments; they reflect the posture changes of the moving target.

[0029] When performing keypoint detection, human pose estimation algorithms, such as OpenPose, can be used to process each time segment in the time-segment set and extract the spatial position sequence of the moving target on the skeletal joints. Then, the relative displacement of the spatial position sequence between adjacent time segments is calculated. The relative displacement can be determined by calculating the spatial distance between corresponding skeletal joints between adjacent time segments. Pose change features are then generated based on the relative displacement; for example, the average value and standard deviation of the relative displacement can be included as part of the pose change features.

[0030] Step S134: Perform target relationship parsing processing on each time segment in the time segment set, identify the spatial distance sequence and contact state change sequence between the moving target and the surrounding static objects, and generate environmental interaction features based on the coupling relationship between the spatial distance sequence and the contact state change sequence.

[0031] Target relationship parsing involves analyzing the relationship between a moving target and surrounding static objects within each time segment. The spatial distance sequence is a sequence of spatial distances between the moving target and surrounding static objects within each time segment. The contact state change sequence is a sequence of changes in the contact state (e.g., contact, non-contact) between the moving target and surrounding static objects within each time segment. The coupling relationship is the mutual relationship between the spatial distance sequence and the contact state change sequence. Environmental interaction features are feature information generated based on the coupling relationship between the spatial distance sequence and the contact state change sequence, reflecting the interaction between the moving target and its surrounding environment. In practice, target detection and relationship analysis algorithms can be used to perform target relationship parsing for each time segment in the time segment set. For example, target detection algorithms such as YOLO can be used to detect the positions of the moving target and surrounding static objects, and then the spatial distance between them can be calculated to generate a spatial distance sequence. Simultaneously, by analyzing the pixel overlap between the moving target and surrounding static objects, their contact state is determined, generating a contact state change sequence. Then, the coupling relationship between the spatial distance sequence and the contact state change sequence is analyzed; for example, when the spatial distance is small, the contact state may be contact; when the spatial distance is large, the contact state may be non-contact. Based on this coupling relationship, environmental interaction features can be generated. For example, the correlation coefficient between spatial distance sequences and contact state change sequences can be used as part of the environmental interaction features.

[0032] Step S135: Based on the start and end timestamps corresponding to each time segment in the time segment set, perform timestamp synchronization processing on the motion trajectory features, posture change features, and environmental interaction features to generate time-aligned versions of motion trajectory features, posture change features, and environmental interaction features with a unified time reference.

[0033] Timestamp synchronization aligns motion trajectory features, posture change features, and environmental interaction features according to the start and end timestamps of each time segment in the time segment set, giving them a unified time reference. The time-aligned version is the version of motion trajectory features, posture change features, and environmental interaction features obtained after timestamp synchronization, and they are consistent in time.

[0034] In practice, based on the start and end timestamps of each time segment in the time segment set, the motion trajectory features, posture change features, and environmental interaction features are arranged in chronological order. For each time segment, the corresponding motion trajectory features, posture change features, and environmental interaction features are matched to ensure temporal correspondence. For example, if the start timestamp of a time segment is t1 and the end timestamp is t2, then the motion trajectory features, posture change features, and environmental interaction features extracted within this time range are combined to generate a time-aligned version with a unified time reference.

[0035] Step S136: Based on the temporal continuity of motion trajectory features, posture change features and environmental interaction features in the time-aligned version, adaptive trajectory interpolation compensation is performed on the feature data corresponding to the missing timestamps to generate motion trajectory features, posture change features and environmental interaction features after complete time-axis alignment.

[0036] Temporal continuity refers to the continuous change of motion trajectory features, posture change features, and environmental interaction features over time. Missing timestamps correspond to feature data that is missing in the time-aligned version due to various reasons (such as data loss, detection failure, etc.). Adaptive trajectory interpolation compensation estimates and supplements the feature data corresponding to missing timestamps based on the temporal continuity of the existing feature data in the time-aligned version. The motion trajectory features, posture change features, and environmental interaction features after complete time-axis alignment are feature data obtained after adaptive trajectory interpolation compensation; they are complete on the time axis with no missing values. In specific implementations, the temporal continuity of motion trajectory features, posture change features, and environmental interaction features in the time-aligned version can be analyzed. The continuity of their changes can be determined by calculating the difference or rate of change between feature data corresponding to adjacent timestamps. For feature data corresponding to missing timestamps, an adaptive trajectory interpolation algorithm is used for compensation. For example, a linear interpolation algorithm can be used to estimate the value of the feature data corresponding to the missing timestamp based on the values ​​of the feature data corresponding to adjacent timestamps. More complex interpolation algorithms, such as spline interpolation algorithms, can also be used to improve the accuracy of interpolation. Through adaptive trajectory interpolation compensation, motion trajectory features, posture change features, and environmental interaction features are generated after full time axis alignment.

[0037] Step S140: The aligned features are weighted and fused according to the preset spatiotemporal weight distribution matrix to generate an initial behavioral feature set.

[0038] The preset spatiotemporal weight distribution matrix is ​​a predefined matrix that contains weight information for different spatiotemporal features. Weighted fusion, based on the preset spatiotemporal weight distribution matrix, weights and sums the aligned motion trajectory features, posture change features, and environmental interaction features, fusing them into a unified feature vector. The initial behavioral feature set is the feature vector set obtained after weighted fusion, containing basic feature information of various behaviors presented in the video.

[0039] In practical implementation, the aligned motion trajectory features, posture change features, and environmental interaction features can be weighted and summed according to a preset spatiotemporal weight distribution matrix. Assuming the preset spatiotemporal weight distribution matrix is ​​W, the aligned motion trajectory features are T, the posture change features are P, and the environmental interaction features are E, then the initial behavioral feature set F can be expressed as: F = W1*T + W2*P + W3*E, where W1, W2, and W3 are the weight vectors corresponding to the motion trajectory features, posture change features, and environmental interaction features, respectively. Through weighted fusion, different spatiotemporal features are fused into a unified feature vector to generate the initial behavioral feature set. It can be understood that after feature alignment and before weighted fusion, a feature normalization or standardization process can be included to unify the dimensions of each feature. For example, T, P, and E can be normalized (e.g., Min-Max normalization or Z-Score standardization) to map them to the same numerical range (e.g., [0,1] or standard normal distribution). In other aspects of the embodiments of the present invention, if similar technical content involving the fusion of features with different dimensions is involved, it should be understood that those skilled in the art can unify the dimensions based on their own known technical knowledge in order to perform further feature operations.

[0040] Step S200: Perform multi-dimensional correlation analysis on the initial behavioral feature set based on the preset behavioral semantic tags, extract the spatiotemporal behavioral feature vector corresponding to the target monitoring scene, and transmit the spatiotemporal behavioral feature vector to the central server.

[0041] Predefined behavioral semantic labels are a predefined series of labels used to describe behaviors, such as "walking," "running," and "falling." Multi-dimensional correlation analysis analyzes the initial set of behavioral features from multiple perspectives to identify features related to the predefined behavioral semantic labels and establish relationships between them. The target monitoring scenario is the specific scenario to be monitored, such as a school playground or classroom. The spatiotemporal behavioral feature vector is a feature vector extracted from the initial set of behavioral features that corresponds to the target monitoring scenario; it contains spatiotemporal feature information of the behavior within the target monitoring scenario. The central server is a server used for centralized processing and storage of data. It receives spatiotemporal behavioral feature vectors from edge computing nodes and performs further analysis and processing.

[0042] For example, the specific implementation process could be as follows: First, extract a set of semantic keywords matching the target monitoring scene from preset behavioral semantic tags, and assign a scene association weight coefficient to each semantic keyword. For instance, for a monitoring scene of a school playground, semantic keywords might include "running," "playing football," etc. A weight coefficient is manually assigned to each keyword according to actual needs to represent its degree of association with the monitoring scene. Then, perform semantic similarity matching between each behavioral feature in the initial behavioral feature set and the set of semantic keywords, filtering out candidate behavioral feature subsets whose similarity exceeds a preset threshold. Natural language processing techniques, such as word vector models, can be used to calculate the similarity between behavioral features and semantic keywords. Next, perform cross-correlation analysis of the behavioral features in the candidate behavioral feature subset using both temporal and spatial dimensions to generate a joint feature map of temporal continuity distribution and spatial density heatmap. By analyzing the distribution of behavioral features in time and space, the correlation between them is determined. Based on the scene association weight coefficient, dynamically weight the feature regions in the joint feature map, enhancing the weight values ​​of feature regions strongly correlated with the target monitoring scene and suppressing the weight values ​​of irrelevant feature regions. The weighted joint feature map is input into a spatiotemporal convolutional kernel for local feature aggregation, extracting feature fragment sequences with temporal dependence and spatial clustering. Based on the cumulative distribution of the feature fragment sequences along the time axis and the overlap ratio of spatial regions, a spatiotemporal behavioral feature vector incorporating multi-dimensional correlations is generated. Finally, the spatiotemporal behavioral feature vector is transmitted to the central server via the network.

[0043] As one implementation method, step S200 includes the following steps: Step S210: Extract a set of semantic keywords that match the target monitoring scene from the preset behavioral semantic tags, and assign a scene association weight coefficient to each semantic keyword.

[0044] The preset behavioral semantic tags are a predefined set of tags used to describe behavior. The semantic keyword set is a collection of keywords related to the target monitoring scenario, selected from the preset behavioral semantic tags. The scenario association weight coefficient is a numerical value assigned to each semantic keyword, representing the degree of association between the keyword and the target monitoring scenario. When extracting the semantic keyword set, the preset behavioral semantic tags are first analyzed to determine the tags related to the target monitoring scenario. Then, a scenario association weight coefficient is assigned to each semantic keyword; the magnitude of the weight coefficient can be determined based on experience or statistical data.

[0045] Step S220: Perform semantic similarity matching between each behavioral feature in the initial behavioral feature set and the semantic keyword set, and filter out candidate behavioral feature subsets whose similarity exceeds a preset threshold.

[0046] Semantic similarity matching calculates the similarity between each behavioral feature in the initial behavioral feature set and each keyword in the semantic keyword set. The candidate behavioral feature subset is a subset of behavioral features selected from the initial behavioral feature set whose similarity to the semantic keyword set exceeds a preset threshold.

[0047] In practical implementation, word vector models from natural language processing techniques, such as Word2Vec or GloVe, can be used to convert behavioral features and semantic keywords into vector representations. After mapping them to the same vector space, the cosine similarity between them is calculated to measure their semantic similarity. Behavioral features with similarity exceeding a preset threshold are added to a subset of candidate behavioral features.

[0048] Step S230: Perform cross-correlation analysis of the behavioral features in the candidate behavioral feature subset in terms of time and space dimensions to generate a joint feature mapping of the temporal continuity distribution map and the spatial density heat map.

[0049] Cross-correlation analysis of temporal and spatial dimensions considers the distribution of behavioral features in both time and space to determine the relationships between them. A temporal continuity distribution map is a graph representing the continuous changes of behavioral features over time. A spatial density heatmap is a graph representing the spatial density of behavioral features. The joint feature map is a feature map obtained by fusing the temporal continuity distribution map and the spatial density heatmap; it integrates information about behavioral features in both time and space.

[0050] In practice, the behavioral features in the candidate behavioral feature subset are first analyzed over time, and the frequency of each behavioral feature at different time points is calculated to generate a temporal continuity distribution map. Then, the behavioral features are analyzed spatially, and the frequency of each behavioral feature at different spatial locations is calculated to generate a spatial density heatmap. Finally, the temporal continuity distribution map and the spatial density heatmap are fused to generate a joint feature map. Image fusion algorithms, such as the weighted average method, can be used to fuse the two images.

[0051] Step S240: Dynamically weight the feature regions in the joint feature mapping based on the scene association weight coefficient, enhance the weight values ​​of feature regions that are strongly related to the target monitoring scene, and suppress the weight values ​​of unrelated feature regions.

[0052] Scene association weight coefficients are numerical values ​​assigned to each semantic keyword to represent its degree of association with the target monitoring scene. Feature regions are different regions in the joint feature map, each corresponding to a specific behavioral feature or combination of behavioral features. Dynamic weighting adjusts the feature regions in the joint feature map in real time based on the scene association weight coefficients, enhancing the weight values ​​of feature regions strongly correlated with the target monitoring scene and suppressing the weight values ​​of irrelevant feature regions. In practice, a weight value can be assigned to each feature region in the joint feature map based on the scene association weight coefficients. Higher weight values ​​are assigned to feature regions strongly correlated with the target monitoring scene, and lower weight values ​​are assigned to irrelevant feature regions. Then, the value of each feature region in the joint feature map is multiplied by its corresponding weight value to obtain the weighted joint feature map.

[0053] Step S250: Input the weighted joint feature map into the spatiotemporal convolution kernel to perform local feature aggregation and extract feature segment sequences with temporal dependence and spatial clustering.

[0054] Spatiotemporal convolution kernels are convolution kernels used to process spatiotemporal data, simultaneously considering information about the data in both time and space. Local feature aggregation merges and extracts local features from the weighted joint feature map to obtain more representative features. A feature fragment sequence is a sequence of feature fragments with temporal dependence and spatial clustering extracted from the weighted joint feature map. In practice, the weighted joint feature map is input into the spatiotemporal convolution kernel for convolution. The spatiotemporal convolution kernel slides across the joint feature map, performing convolution calculations on local regions to extract local features. Then, the extracted local features are aggregated to obtain a feature fragment sequence with temporal dependence and spatial clustering. Pooling operations, such as max pooling or average pooling, can be used to aggregate local features.

[0055] Step S260: Generate a spatiotemporal behavioral feature vector that integrates multi-dimensional correlations based on the cumulative distribution of feature fragment sequences on the time axis and the overlap ratio of spatial regions.

[0056] The cumulative distribution of feature fragment sequences along the time axis represents the frequency and distribution of these sequences over time. The overlap ratio of spatial regions represents the degree of overlap between feature fragment sequences in different spatial regions. The spatiotemporal behavioral feature vector, which integrates multi-dimensional relationships, is obtained by fusing the cumulative distribution of feature fragment sequences along the time axis with the overlap ratio of spatial regions. It integrates the multi-dimensional relationships between behavioral features in time and space. In practice, firstly, the cumulative distribution of feature fragment sequences along the time axis is statistically analyzed, and the frequency of feature fragment occurrence at each time point is calculated. Then, the overlap ratio of feature fragment sequences in different spatial regions is calculated. This overlap ratio can be determined by calculating the ratio of the intersection area to the union area of ​​feature fragments in different spatial regions. Finally, the cumulative distribution of feature fragment sequences along the time axis and the overlap ratio of spatial regions are fused to generate a spatiotemporal behavioral feature vector that integrates multi-dimensional relationships, such as through vector concatenation or weighted summation.

[0057] Step S300: In the central server, the initial behavior recognition model is trained in multiple stages based on the historical behavior data set to obtain the target behavior recognition model. The multi-stage training includes feature fusion based on spatiotemporal correlation and cross-scene behavior pattern transfer.

[0058] The historical behavior dataset consists of behavioral data stored on a central server over a past period, containing spatiotemporal features of various behaviors and corresponding behavior type labels. The initial behavior recognition model is a pre-built model for identifying behaviors; it can be a deep learning model, such as a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN). Multi-stage training involves training the initial behavior recognition model in multiple stages to progressively improve its performance. Spatiotemporal feature fusion combines features from different spatiotemporal contexts to extract more representative features. Cross-scene behavior pattern transfer transfers behavior patterns learned in one scene to another to improve the model's generalization ability. The target behavior recognition model, obtained after multi-stage training, possesses high behavior recognition accuracy and generalization ability.

[0059] As one implementation method, the above step S300 can be specifically implemented as the following steps S310~S370: Step S310: Extract a first training set with spatiotemporal annotation and a second training set of unannotated cross-scene data from the historical behavior data set, wherein each sample in the first training set contains an annotated spatiotemporal behavior feature vector and a corresponding normal behavior type label, and each sample in the second training set contains unannotated video segments under different monitoring scenarios.

[0060] The historical behavior dataset is a large collection of historical behavior data stored on a central server, containing behavioral information from different scenarios. The first training set with spatiotemporal annotations is extracted from this dataset. Each sample contains an annotated spatiotemporal behavior feature vector and a corresponding normal behavior type label. The spatiotemporal behavior feature vector describes the temporal and spatial characteristics of the behavior, while the normal behavior type label indicates the behavior type, such as "walking" or "standing." The second training set, consisting of unannotated cross-scenario data, is also extracted from the historical behavior dataset. Each sample contains unannotated video clips from different monitoring scenarios. These unannotated video clips lack behavior type labeling, and the different monitoring scenarios represent different areas within the campus, such as the playground and classrooms.

[0061] In practice, the historical behavior data set is filtered and organized, and the data with spatiotemporal annotations is extracted to form the first training set. Simultaneously, unlabeled cross-scene data is extracted to form the second training set. The data can be filtered based on annotation and scene information to ensure the accuracy and representativeness of both the first and second training sets.

[0062] Step S320: Input the first training set into the initial behavior recognition model, and perform cross-frame feature alignment on the spatiotemporal behavior feature vector through the spatiotemporal correlation constraint module to generate a first intermediate feature representation with temporal continuity.

[0063] The initial behavior recognition model is a pre-built model for recognizing behaviors. The spatiotemporal correlation constraint module is a module within the model used for cross-frame feature alignment of spatiotemporal behavior feature vectors. For example, it can employ a bidirectional long short-term memory network (Bi-LSTM) based on an attention mechanism. Cross-frame feature alignment aligns the spatiotemporal behavior feature vectors from different frames, making them continuous in time. The first intermediate feature representation is the feature representation obtained after cross-frame feature alignment; it has temporal continuity and can better reflect the dynamic changes in behavior.

[0064] In specific implementation, the first training set is input into the spatiotemporal correlation constraint module of the initial behavior recognition model. This module extracts feature subsequences corresponding to consecutive time segments from the spatiotemporal behavior feature vector and calculates the dynamic time warping distance between adjacent time segment feature subsequences. Dynamic time warping distance is a method used to measure the similarity between two time series; it can handle the scaling and deformation of time series. Then, nonlinear interpolation is performed on the feature subsequences based on the dynamic time warping distance to generate a time-axis aligned fourth intermediate feature representation. Next, the fourth intermediate feature representation is sliced ​​along the time dimension to generate multiple local feature blocks within overlapping time windows. Learnable spatiotemporal attention weights are assigned to each local feature block, and feature enhancement is performed based on these weights to generate a fifth intermediate feature representation. Finally, the fifth intermediate feature representation is input into a bidirectional temporal convolutional network. Forward propagation captures future contextual dependencies, and backpropagation captures historical contextual dependencies, generating a sixth intermediate feature representation with bidirectional temporal correlation. The sixth intermediate feature representation is then residually concatenated with the original spatiotemporal behavior feature vector to generate a first intermediate feature representation.

[0065] As one implementation, step S320 may specifically include the following steps S321 to S326: Step S321: Extract feature subsequences corresponding to continuous time segments from the spatiotemporal behavior feature vector, and calculate the dynamic time warping distance between feature subsequences of adjacent time segments.

[0066] Spatiotemporal behavioral feature vectors are vectors that describe the characteristics of behavior in time and space. Feature subsequences corresponding to consecutive time segments are extracted from the spatiotemporal behavioral feature vectors, with each subsequence corresponding to a consecutive time segment. Dynamic time warping distance is a method used to measure the similarity between two time series; it can handle the scaling and deformation of time series.

[0067] In practical implementation, feature subsequences corresponding to consecutive time segments are extracted from the spatiotemporal behavioral feature vector based on timestamp information. For example, the spatiotemporal behavioral feature vector can be divided into multiple consecutive time segments at fixed time intervals, with each time segment corresponding to a feature subsequence. Then, a dynamic time warping algorithm is used to calculate the distance between feature subsequences of adjacent time segments. The dynamic time warping algorithm calculates the distance between two time series by finding the optimal matching path between them.

[0068] Step S322: Perform nonlinear interpolation on the feature subsequence based on the dynamic time warping distance to generate a time-axis aligned fourth intermediate feature representation.

[0069] Nonlinear interpolation is a process that interpolates feature subsequences based on dynamic time warping distance to achieve time axis alignment. Time axis alignment aligns feature subsequences from different time segments in time, giving them the same time reference. The fourth intermediate feature representation is the feature representation obtained after nonlinear interpolation, and it is aligned on the time axis. Specifically, the mapping relationship between feature subsequences is determined based on the dynamic time warping distance. Then, a nonlinear interpolation algorithm (such as spline interpolation) is used to interpolate the feature subsequences, generating the time-aligned fourth intermediate feature representation. Through nonlinear interpolation, feature subsequences from different time segments can be made more continuous and consistent in time.

[0070] Step S323: Slice the fourth intermediate feature representation in the time dimension to generate local feature blocks within multiple overlapping time windows.

[0071] Temporal slicing involves dividing the fourth intermediate feature representation along the time dimension to generate multiple local feature blocks within overlapping time windows. Overlapping time windows have a certain degree of overlap between adjacent time windows, allowing for better capture of dynamic behavioral changes. A local feature block is a feature block extracted within each time window, containing behavioral feature information within that time window. Specifically, the fourth intermediate feature representation is temporally sliced ​​according to a fixed time window size and overlap ratio. For example, if the time window size is 10 frames and the overlap ratio is 50%, slicing is performed every 5 frames, generating multiple local feature blocks within overlapping time windows. Each local feature block contains spatiotemporal behavioral feature information within that time window.

[0072] Step S324: Assign learnable spatiotemporal attention weights to each local feature block, and perform feature enhancement on the local feature blocks based on the spatiotemporal attention weights to generate the fifth intermediate feature representation.

[0073] Learnable spatiotemporal attention weights are weights assigned to each local feature block, which can be learned and adjusted during model training. These weights represent the importance of each local feature block within the overall behavioral features. Feature enhancement involves weighted summation of local feature blocks based on these spatiotemporal attention weights to amplify the influence of important features. The fifth intermediate feature representation is the feature representation obtained after feature enhancement, further highlighting important behavioral features.

[0074] In practice, a fully connected layer or attention mechanism is used to assign learnable spatiotemporal attention weights to each local feature block. Then, each local feature block is multiplied by its corresponding spatiotemporal attention weight and summed using weighted averages to obtain the enhanced feature representation. Feature enhancement allows the model to focus more on important behavioral features, improving the accuracy of behavior recognition.

[0075] Step S325: Input the fifth intermediate feature representation into a bidirectional temporal convolutional network, capture future contextual dependencies through forward propagation, and capture historical contextual dependencies through backpropagation, to generate a sixth intermediate feature representation with bidirectional temporal correlation.

[0076] Bidirectional temporal convolutional networks (TTCNNs) are convolutional networks that can simultaneously process both forward and backward information from a time series. Forward propagation computes from the start point to the end point of the time series, capturing future contextual dependencies. Backward propagation computes from the end point to the start point of the time series, capturing historical contextual dependencies. The sixth intermediate feature representation is the feature representation obtained after processing by the TTCNN; it possesses bidirectional temporal correlation and can better reflect the dynamic changes in behavior.

[0077] In its implementation, the fifth intermediate feature representation is input into a bidirectional temporal convolutional network. The network computes future context dependencies and historical context dependencies through forward and backward propagation, respectively. Then, the results of the forward and backward propagation are merged to generate a sixth intermediate feature representation with bidirectional temporal correlation. This bidirectional temporal convolutional network can effectively capture the temporal sequence information of behavior, improving the accuracy of behavior recognition.

[0078] Step S326: Perform a residual concatenation between the sixth intermediate feature representation and the original spatiotemporal behavior feature vector to generate the first intermediate feature representation.

[0079] Residual connections sum the network's input and output to mitigate the vanishing gradient problem and improve training efficiency. The original spatiotemporal behavior feature vector is the spatiotemporal behavior feature vector input to the spatiotemporal correlation constraint module. The first intermediate feature representation is the feature representation obtained after residual connections, combining information from the original spatiotemporal behavior feature vector and the processed sixth intermediate feature representation.

[0080] In practice, the sixth intermediate feature representation is added element-wise to the original spatiotemporal behavior feature vector to obtain the first intermediate feature representation. Through residual connections, information from the original spatiotemporal behavior feature vector is preserved, while information from the processed sixth intermediate feature representation is fused, thus improving the quality of the feature representation.

[0081] Step S330: Input the first intermediate feature representation into the spatiotemporal feature fusion module, dynamically assign weights and superimpose features within adjacent time windows to generate the fused second intermediate feature representation.

[0082] The spatiotemporal feature fusion module is a component of the model used to fuse features within adjacent time windows. For example, it can be implemented as a gated recurrent unit (GRU). Dynamic weight allocation assigns different weights to features within adjacent time windows based on their importance. Feature overlay combines the features from adjacent time windows according to their assigned weights to obtain the fused feature representation. The second intermediate feature representation is the feature representation obtained after processing by the spatiotemporal feature fusion module, which incorporates feature information from adjacent time windows. Specifically, the first intermediate feature representation is input into the spatiotemporal feature fusion module. This module first assigns dynamic weights to features within adjacent time windows based on their importance. These weights can be learned using a fully connected layer or an attention mechanism. Then, the features from adjacent time windows are overlaid according to their assigned weights to obtain the fused second intermediate feature representation. Through dynamic weight allocation and feature overlay, feature information from adjacent time windows can be better fused, improving the accuracy of behavior recognition.

[0083] Step S340: Based on the similarity loss between the second intermediate feature representation and the normal behavior type label, supervised training is performed on the initial behavior recognition model to obtain the first intermediate model.

[0084] Similarity loss is a loss function used to measure the degree of difference between the second intermediate feature representation and the normal behavior type label. Supervised training involves adjusting the model's parameters by minimizing the similarity loss using labeled data. The first intermediate model is the model obtained after supervised training, which can accurately identify normal behavior types to a certain extent.

[0085] In practice, a suitable similarity loss function, such as the cross-entropy loss function, is used to calculate the loss between the second intermediate feature representation and the normal behavior type label. Then, an optimization algorithm, such as stochastic gradient descent, is used to update the parameters of the initial behavior recognition model to minimize the similarity loss. After multiple iterations of training, a first intermediate model is obtained. Through supervised training, the model can learn the features of normal behavior types, improving the accuracy of behavior recognition.

[0086] Step S350: Input the second training set into the first intermediate model, and extract scene-invariant features from the unlabeled video segments through the cross-scene behavior pattern transfer module to generate a third intermediate feature representation that is unrelated to the monitoring scene.

[0087] The cross-scene behavior pattern transfer module is a component of the model used to extract scene-invariant features from unlabeled video clips. Scene-invariant features are those unaffected by the monitoring scene and reflect the essential characteristics of the behavior. The third intermediate feature representation is the feature representation obtained after processing by the cross-scene behavior pattern transfer module; it is independent of the monitoring scene and has good generalization ability.

[0088] In practice, the second training set is input into the cross-scene behavior pattern transfer module of the first intermediate model. This module uses unsupervised learning algorithms, such as autoencoders or variational autoencoders, to extract features from unlabeled video segments. By learning the latent features in the video segments, scene-invariant features are extracted. Then, these scene-invariant features are combined into a third intermediate feature representation. Through cross-scene behavior pattern transfer, the model can learn behavior patterns in different scenarios, improving the model's generalization ability.

[0089] Step S360: Perform temporal dimension compression and feature reconstruction on the third intermediate feature representation to generate cross-scene behavior pattern embedding vectors, and perform self-supervised training on the first intermediate model based on the reconstruction error to obtain the second intermediate model.

[0090] Temporal compression compresses the third intermediate feature representation along the time dimension to reduce the feature dimensionality. Feature reconstruction reconstructs the compressed features to restore them to their original dimensionality. The cross-scene behavior pattern embedding vector is the vector obtained after temporal compression and feature reconstruction, representing cross-scene behavior patterns. Reconstruction error is the degree of difference between the reconstructed features and the original features. Self-supervised training adjusts the model parameters by minimizing the reconstruction error in the presence of unlabeled data. The second intermediate model is the model obtained after self-supervised training, which exhibits better performance in cross-scene behavior recognition.

[0091] In practical implementation, a temporal compression module, such as a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU), can be used to compress the temporal dimension of the third intermediate feature representation. Then, a reconstruction module, such as a fully connected layer or a convolutional layer, is used to reconstruct the compressed features. The reconstruction error between the reconstructed features and the original features is calculated, and an optimization algorithm, such as stochastic gradient descent, is used to update the parameters of the first intermediate model to minimize the reconstruction error. After multiple iterations of training, a second intermediate model is obtained. Through self-supervised training, the model can learn behavior patterns across scenarios, improving its generalization ability.

[0092] Step S370: Alternately weight and fuse the feature extraction layer parameters of the first intermediate model with the feature reconstruction layer parameters of the second intermediate model to generate a target behavior recognition model with cross-scene adaptability.

[0093] The feature extraction layer parameters are the parameters used in the first intermediate model to extract features. The feature reconstruction layer parameters are the parameters used in the second intermediate model to reconstruct features. Alternating weighted fusion involves alternately weighting the feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model to generate new parameters. The target behavior recognition model is the model obtained after alternating weighted fusion. It has cross-scene adaptability and can accurately identify behaviors in different monitoring scenarios.

[0094] In practical implementation, the feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model can be alternately weighted according to a certain weight ratio. For example, a linear combination method can be used, such as α*parameter1 + (1-α)*parameter2, where α is the weight coefficient. Through alternating weighted fusion, parameters for a target behavior recognition model with cross-scene adaptability are generated. Applying these parameters to the model yields the final target behavior recognition model. By alternating weighted fusion, the advantages of the first and second intermediate models can be comprehensively utilized, improving the model's cross-scene adaptability and behavior recognition accuracy.

[0095] Step S400: Input the spatiotemporal behavior feature vector into the target behavior recognition model to generate a behavior semantic description sequence corresponding to the video data stream, and determine the real-time behavior monitoring result based on the matching result of the behavior semantic description sequence and the preset abnormal behavior rule library.

[0096] Spatiotemporal behavior feature vectors are extracted from the video data stream and contain temporal and spatial feature information of the behavior. The target behavior recognition model is a trained model used to identify behaviors. The behavior semantic description sequence is a sequence that describes the behaviors in the video data stream using natural language; it includes information such as the action, interaction, and intent of the behavior. The preset abnormal behavior rule base is a set of predefined abnormal behavior rules, each rule corresponding to a type of abnormal behavior. The real-time behavior monitoring result is determined based on the matching result between the behavior semantic description sequence and the preset abnormal behavior rule base; it indicates whether abnormal behavior exists in the current video data stream and the type of abnormal behavior.

[0097] The specific implementation process is as follows: First, the spatiotemporal behavior feature vector is segmented by a sliding window along the time axis to generate a sequence of feature segments corresponding to continuous time segments. Then, the feature segment sequence is input into the basic action parsing layer of the target behavior recognition model to extract the limb displacement rate change pattern and direction consistency features of the moving target within the time segment, generating an atomic action label sequence. Next, the atomic action label sequence is input into the interaction relationship inference layer of the target behavior recognition model. Based on the temporal overlap and spatial proximity between atomic action labels, the collaborative action patterns between multiple moving targets are identified, generating an interaction relationship description set. The interaction relationship description set is input into the intent inference layer of the target behavior recognition model. Combining the scene context features corresponding to the time segment, the logical matching degree between the collaborative action pattern and the preset behavior intent template is analyzed, generating a behavior intent probability distribution. Based on the atomic action label sequence, the interaction relationship description set, and the behavior intent probability distribution, a structured behavior semantic description tree containing action levels, interaction levels, and intent levels is generated. The structured behavior semantic description tree is then subjected to time axis compression and semantic aggregation, redundant description nodes are removed, and branches of the same type of action are merged to generate a behavior semantic description sequence. Finally, semantic parsing is performed on the behavioral semantic description sequence to extract key behavioral action nodes and the temporal dependencies between nodes. A behavioral state transition graph is constructed based on these temporal dependencies, and the topological similarity between the behavioral state transition graph and a reference state transition graph in the abnormal behavior rule base is calculated. When the topological similarity exceeds a preset threshold, an abnormal behavior type identifier matching the abnormal behavior rule base is generated. Based on the risk level corresponding to the abnormal behavior type identifier, real-time behavioral monitoring results containing behavioral warning strategies are generated.

[0098] As one implementation method, step S400 may specifically include the following steps: Step S410: Perform time axis sliding window segmentation on the spatiotemporal behavior feature vector to generate a feature segment sequence corresponding to continuous time segments.

[0099] Time-axis sliding window segmentation involves using a fixed-size window to slide along the time axis of the spatiotemporal behavior feature vector, dividing the vector into multiple feature segment sequences corresponding to consecutive time segments. Each feature segment sequence contains spatiotemporal behavior feature information within a specific time segment. Specifically, sliding window segmentation is performed on the time axis of the spatiotemporal behavior feature vector according to a preset window size and sliding step size. For example, if the window size is 10 frames and the sliding step size is 5 frames, segmentation is performed every 5 frames, generating multiple feature segment sequences corresponding to consecutive time segments. Each feature segment sequence contains spatiotemporal behavior feature information within that specific time segment.

[0100] Step S420: Input the feature segment sequence into the basic action parsing layer in the target behavior recognition model, extract the limb displacement rate change pattern and direction consistency features of the moving target within the time segment, and generate an atomic action label sequence.

[0101] The basic motion parsing layer is a layer in the target behavior recognition model used to parse the basic motions of a moving target within a time segment. The limb displacement rate change pattern describes the changes in the displacement rate of the moving target's limbs within a time segment. The direction consistency feature represents the degree of consistency in the movement direction of the moving target's limbs within a time segment. The atomic motion label sequence is a sequence obtained by labeling the basic motions within the time segment; each label corresponds to an atomic motion, such as "walking" or "turning around."

[0102] In practice, the feature segment sequence is input into the basic action parsing layer of the target behavior recognition model. This layer uses computer vision algorithms, such as optical flow or keypoint detection algorithms, to extract the limb displacement rate change patterns and directional consistency features of the moving target within a time segment. Then, based on these features, a classifier, such as a support vector machine or neural network, is used to classify the basic actions within the time segment, generating an atomic action label sequence.

[0103] Step S430: Input the atomic action label sequence into the interaction relationship inference layer in the target behavior recognition model. Based on the temporal overlap and spatial proximity between atomic action labels, identify the cooperative action patterns between multiple moving targets and generate an interaction relationship description set.

[0104] The interaction relationship inference layer is a network layer in the target behavior recognition model. It is used to infer the interaction relationships between multiple moving targets, such as the inference layer in a graph neural network (GNN). Temporal overlap is the degree of overlap between atomic action labels in time, and spatial proximity is the degree of proximity between moving targets in space. Cooperative action patterns are the ways in which multiple moving targets cooperate, such as "walking together" or "colliding with each other." The interaction relationship description set is a set obtained after describing the interaction relationships between multiple moving targets, with each description corresponding to a cooperative action pattern. In specific implementation, the atomic action label sequence is input into the interaction relationship inference layer in the target behavior recognition model. This layer first calculates the temporal overlap and spatial proximity between atomic action labels. Then, based on the temporal overlap and spatial proximity, a graph neural network or rule inference system is used to identify the cooperative action patterns between multiple moving targets and generate the interaction relationship description set. For example, the atomic action label sequence is obtained, with each label corresponding to the basic action of the moving target at a certain time period, such as student A "running" and student B "passing the ball," and the action time range and target position are recorded. Next, temporal overlap and spatial proximity are calculated. The temporal overlap of overlapping time ranges for different atomic action tags is determined. For example, if student A's "running" time is 8:00-8:10 and student B's "passing" time is 8:05-8:12, the overlap is 5 minutes. Spatial distance is calculated based on target locations to determine proximity; distances less than a threshold are considered proximity. When two atomic action tags have high temporal overlap and are spatially close, they are identified as a collaborative action pattern, such as "passing-receiving collaboration." Finally, these patterns are organized and described to form an interaction relationship description set such as "student A and student B have a passing-receiving collaboration."

[0105] Step S440: Input the set of interaction relationship descriptions into the intent inference layer in the target behavior recognition model, combine the scene context features corresponding to the time segment, analyze the logical matching degree between the collaborative action mode and the preset behavior intent template, and generate the behavior intent probability distribution.

[0106] The intent inference layer is a layer in the target behavior recognition model used to infer the behavioral intent of a moving target. It can include network layers of knowledge graphs and LSTM. Scene context features are relevant features of the scene corresponding to the time segment, such as scene type and environmental information. Preset behavioral intent templates are a series of predefined behavioral intent templates, each template corresponding to a specific behavioral intent, such as "play" or "fight". Logical matching degree is the degree of logical matching between the collaborative action pattern and the preset behavioral intent template. Behavioral intent probability distribution is the probability distribution obtained after matching the collaborative action pattern with the preset behavioral intent template, with each probability corresponding to the likelihood of a behavioral intent. In specific implementation, the set of interaction relationship descriptions is input into the intent inference layer in the target behavior recognition model. This layer combines the scene context features corresponding to the time segment with a deep learning model or knowledge graph to analyze the logical matching degree between the collaborative action pattern and the preset behavioral intent template. Then, based on the logical matching degree, the behavioral intent probability distribution is generated. For example, for a certain collaborative action pattern, its matching score with each preset behavioral intent template is calculated, and the score is normalized to obtain the behavioral intent probability distribution.

[0107] Step S450: Generate a structured behavioral semantic description tree containing action level, interaction level and intent level based on the atomic action tag sequence, interaction relationship description set and behavioral intent probability distribution.

[0108] A structured behavioral semantic description tree (MSD) is a tree-like structure that contains information at the action, interaction, and intent levels. The action level represents the basic actions of a moving target, the interaction level represents the interaction relationships between multiple moving targets, and the intent level represents the behavioral intent of the moving target. For example, each label in an atomic action label sequence can be mapped to a leaf node in the MSD, and the start timestamp and duration of each leaf node can be recorded. Then, based on the collaborative action patterns in the interaction relationship description set, horizontal association edges are established between the leaf nodes. These horizontal association edges contain interaction type identifiers and the number of participating moving targets. Based on intent categories whose probability values ​​exceed a preset threshold in the behavioral intent probability distribution, intent branch nodes are created below the root node of the MSD, and the leaf nodes corresponding to the horizontal association edges are clustered into matching intent branch nodes. Leaf nodes in the MSD where timestamp overlap exceeds a preset ratio are merged to generate composite action nodes, and vertical association edges are established between the composite action nodes and their corresponding intent branch nodes. Based on the temporal coverage of the vertically related edges and the logical consistency with the intent branch nodes, the composite action nodes undergo semantic generalization and are replaced with high-level behavioral description phrases. Through these steps, a structured behavioral semantic description tree containing action, interaction, and intent levels is generated.

[0109] Step S460: Perform time-axis compression and semantic aggregation on the structured behavioral semantic description tree, remove redundant description nodes and merge branches of the same type of action to generate a behavioral semantic description sequence.

[0110] Timeline compression compresses the structured behavioral semantic description tree along the timeline, removing redundant temporal information. Semantic aggregation merges nodes with the same semantic meaning, reducing the complexity of the description. Redundant description nodes are those in the structured behavioral semantic description tree that do not substantially contribute to the behavioral semantic description. Branches of the same type of action are branches with the same action type. The behavioral semantic description sequence is the sequence obtained after timeline compression and semantic aggregation, which more concisely describes the semantic information of the behavior.

[0111] In practice, the structured behavioral semantic description tree is first compressed along the timeline. This can be done by merging nodes with overlapping timestamps to remove redundant temporal information. Then, semantic aggregation is performed, merging nodes with the same semantic meaning. For example, multiple nodes representing "walking" are merged into one node. Redundant descriptive nodes, such as intermediate nodes without practical significance, are removed. Finally, the processed structured behavioral semantic description tree is converted into a sequence, generating a behavioral semantic description sequence.

[0112] As one implementation method, the process of generating a structured behavior semantic description tree and performing timeline compression and semantic aggregation may specifically include: step S451: mapping each tag in the atomic action tag sequence to a leaf node of the structured behavior semantic description tree, and recording the start timestamp and duration corresponding to each leaf node.

[0113] A structured behavior semantic description tree is a tree-like structure used to represent behavior semantics. Leaf nodes are the bottom-level nodes of the tree, and each leaf node corresponds to an atomic action label. The start timestamp is the time when the atomic action begins, and the duration is the duration of the atomic action.

[0114] In practice, the atomic action tag sequence is traversed, and each tag is mapped to a leaf node of the structured behavior semantic description tree. Simultaneously, based on the time segment information, the start timestamp and duration corresponding to each leaf node are recorded. For example, if the time segment corresponding to a certain atomic action tag is from frame 10 to frame 20, then the start timestamp of that leaf node is frame 10, and the duration is 10 frames.

[0115] Step S452: Based on the collaborative action patterns in the interaction relationship description set, establish horizontal association edges between leaf nodes. The horizontal association edges contain interaction type identifiers and information on the number of participating motion targets.

[0116] The interaction relationship description set contains information on the cooperative action patterns between multiple moving targets. Lateral association edges are edges established between the leaf nodes of the structured behavior semantic description tree, representing the interaction relationship between two atomic actions. Interaction type identifiers are symbols used to identify the interaction type, such as "together" or "collision." The number of participating moving targets is the number of moving targets participating in the cooperative action. In practice, the interaction relationship description set is traversed, and lateral association edges are established between the corresponding leaf nodes based on the information in the cooperative action patterns. For example, if a cooperative action pattern indicates that two moving targets are walking together, a lateral association edge is established between the leaf nodes corresponding to these two moving targets, and the edge is labeled with the interaction type identifier "together" and the number of participating moving targets, 2.

[0117] Step S453: Based on the intent categories whose probability values ​​exceed a preset threshold in the behavioral intent probability distribution, create intent branch nodes below the root node of the structured behavioral semantic description tree, and cluster the leaf nodes corresponding to the horizontal association edges to the matching intent branch nodes.

[0118] The behavioral intent probability distribution is the probability distribution of each behavioral intent. The preset threshold is a pre-defined probability value used to filter out behavioral intents with higher probabilities. Intent branch nodes are branch nodes below the root node in the structured behavioral semantic description tree; each intent branch node corresponds to a specific behavioral intent. Clustering groups related leaf nodes under the same intent branch node. Specifically, the behavioral intent probability distribution is traversed to determine intent categories with probability values ​​exceeding the preset threshold. Intent branch nodes corresponding to these intent categories are created below the root node of the structured behavioral semantic description tree. Then, based on the information from the horizontal association edges, the corresponding leaf nodes are clustered into the matching intent branch nodes. For example, if an intent category is "play," then the leaf nodes corresponding to the horizontal association edges related to "play" are clustered under the "play" intent branch node.

[0119] Step S454: Merge leaf nodes in the structured behavior semantic description tree where the timestamp overlaps by more than a preset ratio to generate composite action nodes, and establish vertical association edges between the composite action nodes and the corresponding intent branch nodes.

[0120] Timestamp overlap occurs when the time segments corresponding to two leaf nodes overlap. The preset ratio is a pre-defined value used to determine whether timestamp overlap meets the merging criteria. A composite action node is a node formed by merging multiple leaf nodes, representing a composite action. Vertical association edges are edges in the structured behavior semantic description tree that run from composite action nodes to intent branch nodes, representing the relationship between composite actions and behavioral intents.

[0121] In specific implementation, traverse the structured behavior semantic description tree to determine the leaf nodes where the timestamp overlap exceeds a preset ratio. Merge these leaf nodes into a composite action node. Then, according to the intention corresponding to the composite action node, establish a vertical association edge between it and the corresponding intention branch node. Exemplarily, if the intention corresponding to a certain composite action node is "play", then establish a vertical association edge between this composite action node and the "play" intention branch node.

[0122] Step S455: According to the time coverage range of the vertical association edge and the logical consistency of the intention branch node, perform semantic generalization processing on the composite action node and replace it with a high-level behavior description phrase.

[0123] The time coverage range of the vertical association edge is the time range corresponding to the composite action node and the intention branch node connected by the vertical association edge. Logical consistency is the degree of logical correspondence between the composite action node and the intention branch node. Semantic generalization processing is to abstract and generalize the description of the composite action node and replace it with a higher-level behavior description phrase.

[0124] In specific implementation, analyze the time coverage range of the vertical association edge and the logical consistency of the intention branch node. For the composite action nodes that meet the logic, perform semantic generalization processing. Exemplarily, if a composite action node contains atomic actions such as "running" and "jumping", and the corresponding intention branch node is "play", then this composite action node can be replaced with a high-level behavior description phrase such as "lively play".

[0125] Step S456: Generate a compressed behavior semantic description sequence based on the chronological order of the high-level behavior description phrases.

[0126] The high-level behavior description phrase is the behavior description phrase obtained after semantic generalization processing. The chronological order arrangement is to arrange according to the chronological order corresponding to the high-level behavior description phrases. The compressed behavior semantic description sequence is the sequence obtained by arranging the high-level behavior description phrases in chronological order, which removes redundant information and is more concise and clear.

[0127] In specific implementation, sort the high-level behavior description phrases according to the start timestamps corresponding to the high-level behavior description phrases. Then, connect the sorted high-level behavior description phrases in sequence to generate a compressed behavior semantic description sequence. Exemplarily, if there are three high-level behavior description phrases "read books quietly", "play lively", and "queue up orderly", and their start timestamps are t1, t2, t3 in sequence, and t1 < t2 < t3, then the compressed behavior semantic description sequence is "read books quietly -> play lively -> queue up orderly".

[0128] As one implementation method, in step S400, the real-time behavior monitoring result is determined based on the matching result between the behavior semantic description sequence and the preset abnormal behavior rule base. Specifically, it may include: step S470: performing semantic parsing on the behavior semantic description sequence to extract key behavior action nodes and the temporal dependencies between nodes.

[0129] Semantic parsing involves analyzing a sequence of semantic descriptions of behaviors to understand their semantic information. Key behavioral action nodes are those nodes in the sequence that significantly influence the semantics of the behavior; they represent the main actions of the behavior. Temporal dependencies between nodes are the sequential or concurrent relationships between key behavioral action nodes, reflecting the temporal order and logical relationships of the behaviors.

[0130] In practical implementation, the behavioral semantic description sequence can first be divided into multiple continuous basic semantic units according to a preset time granularity. Redundancy filtering of these basic semantic units is then performed based on a semantic coherence threshold to generate a de-redundant semantic unit set. Action keyword matching is then performed on each semantic unit in the de-redundant semantic unit set to identify candidate semantic units containing target action words from a preset behavioral action dictionary, and these candidate semantic units are mapped to initial behavioral action nodes. Based on the distribution density of the initial behavioral action nodes on the time axis, initial behavioral action nodes within adjacent time windows are clustered and merged to generate key behavioral action nodes with temporal clustering. Dependency analysis is performed on the time intervals between key behavioral action nodes. If the time interval between two key behavioral action nodes is less than a preset action association threshold, a temporal dependency edge is established between the two key behavioral action nodes. Based on the directionality of the temporal dependency edge, the sequential triggering relationship or concurrent relationship between the key behavioral action nodes is determined, generating a set of temporal dependency relationships with temporal tags. Perform context consistency verification on the temporal dependency edges in the temporal dependency set. If there are dependency edges that conflict with the temporal logic verified in the historical behavior pattern, remove the conflicting dependency edges and recalculate the temporal dependencies to generate optimized key behavior action nodes and temporal dependencies between nodes.

[0131] When segmenting the behavioral semantic description sequence according to a preset time granularity, the preset time granularity can be flexibly set according to the actual monitoring scenario and needs, such as in seconds or minutes. The semantic coherence threshold measures the degree of semantic association between basic semantic units. When the semantic coherence of a basic semantic unit with other units is lower than this threshold, it is considered redundant and should be filtered out. When performing action keyword matching, the preset behavioral action dictionary is a pre-constructed set containing various behavioral action words. By matching the semantic units in the deredundant semantic unit set with the target action words in this dictionary, semantic units with key behavioral actions can be accurately identified.

[0132] Clustering and merging based on the distribution density of initial action nodes on the timeline can combine initial action nodes that are close in time into a more representative key action node. For example, in a campus monitoring scenario, if the initial action node "running" appears multiple times in a short period of time and they are densely distributed on the timeline, then these "running" nodes can be merged into a key action node representing "continuous running".

[0133] When performing dependency analysis on the time intervals between key action nodes, a preset action association threshold is used to determine whether two key action nodes are related. If the time interval between two nodes is less than this threshold, a temporal dependency relationship is considered to exist between them, and this relationship is represented by establishing a temporal dependency edge. When determining whether a triggering relationship or a concurrent relationship exists, the directionality of the temporal dependency edge is used for judgment. For example, if the temporal dependency edge points from node A to node B, it means that node A is triggered before node B; if there is a bidirectional temporal dependency edge, it means that the two nodes may be concurrent.

[0134] When performing context consistency verification, historical behavior patterns are behavioral logic patterns obtained through analysis and learning of a large amount of historical data. When a temporal dependency edge in the set of temporal dependencies is found to conflict with the verified temporal logic in the historical behavior pattern, it indicates that the dependency edge may be incorrect. The conflicting dependency edge needs to be removed and the temporal dependencies recalculated to ensure the accuracy and reliability of key behavioral action nodes and the temporal dependencies between nodes.

[0135] Step S480: Construct a behavior state transition graph based on temporal dependencies, and calculate the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule base.

[0136] A behavior state transition graph is a graph structure used to represent changes and transitions in behavior states. It consists of state nodes and state transition edges. State nodes represent key behavior action nodes, and state transition edges represent temporal dependencies between nodes. A reference state transition graph is a predefined graph structure in the abnormal behavior rule base used to represent abnormal behavior state transitions. Topological similarity is an indicator used to measure the degree of similarity between two graph structures in terms of topological structure. By comparing the topological similarity between the behavior state transition graph and the reference state transition graph, it can be determined whether the current behavior matches an abnormal behavior rule.

[0137] In practical implementation, key behavioral action nodes are first mapped to state nodes in the state transition graph, and temporal dependency edges in the temporal dependency relationship are mapped to state transition edges, generating an initial behavioral state transition graph. Each state transition edge in the initial behavioral state transition graph is labeled with temporal constraints, and the time interval range of the corresponding temporal dependency edge in the temporal dependency relationship is extracted and used as the temporal constraint condition for the state transition edge. Based on the temporal constraint condition of the state transition edge and the state transition frequency in historical behavioral patterns, a transition probability weight is assigned to each state transition edge, generating a behavioral state transition graph with temporal constraints and probability attributes. A reference state transition graph matching the current monitoring scenario is extracted from the abnormal behavior rule base. The reference state transition graph contains a predefined set of abnormal state nodes and a set of abnormal state transition edges. The state nodes in the behavioral state transition graph are semantically aligned with the abnormal state nodes in the reference state transition graph, identifying pairs of state nodes with the same behavioral semantic description. Based on the semantic alignment results of the state node pairs, subgraph matching is performed on the behavioral state transition graph and the reference state transition graph, extracting the structural differences between the two in terms of state transition edge direction, temporal constraint conditions, and transition probability weights. The topological coverage between the behavioral state transition diagram and the reference state transition diagram is calculated based on the structural difference, and the topological similarity is generated based on the comparison result of the topological coverage and the preset anomaly threshold.

[0138] When mapping key action nodes to state nodes in a state transition graph, each key action node corresponds to a state node, and the attributes of the state node can include information such as the name of the action and the time of occurrence. When mapping temporal dependency edges in a temporal dependency relationship to state transition edges, the attributes of the state transition edges can include time constraints and transition probability weights. When annotating state transition edges with time constraints, the time constraints of the state transition edges are determined by analyzing the time interval range of the corresponding temporal dependency edges in the temporal dependency relationship. For example, it may be stipulated that a certain state transition edge must occur within a specific time range.

[0139] The transition probability weights of state transition edges are assigned based on the time constraints of the state transition edges and the state transition frequency in historical behavior patterns. The state transition frequency in historical behavior patterns can be obtained through statistical analysis of a large amount of historical data. For example, if the transition frequency from state A to state B is high in historical data, then a higher transition probability weight is assigned to that state transition edge.

[0140] During semantic alignment, state node pairs with identical semantics are identified by comparing the behavioral semantic descriptions of state nodes in the behavioral state transition graph with those of anomalous state nodes in the reference state transition graph. During subgraph matching, the differences between the behavioral state transition graph and the reference state transition graph in terms of state transition edge directions, time constraints, and transition probability weights are analyzed to calculate the structural difference degree. Topological coverage can be obtained by calculating the proportion of matching portions in the two graph structures. When the topological coverage exceeds a preset anomaly threshold, the behavioral state transition graph and the reference state transition graph are considered to have high topological similarity, and the current behavior may be anomalous.

[0141] Step S490: When the topological similarity exceeds a preset threshold, generate an abnormal behavior type identifier that matches the abnormal behavior rule base.

[0142] The preset threshold is a pre-defined value used to determine whether the topological similarity meets the anomaly criteria. The abnormal behavior type identifier is a symbol or code used to identify the type of abnormal behavior; it corresponds to a specific abnormal behavior type in the abnormal behavior rule base. When the topological similarity between the behavior state transition diagram and the reference state transition diagram exceeds the preset threshold, it indicates that the current behavior highly matches a certain abnormal behavior pattern in the abnormal behavior rule base. In this case, a corresponding abnormal behavior type identifier needs to be generated for subsequent processing.

[0143] In practice, when the calculated topological similarity exceeds a preset threshold, an abnormal behavior rule matching the current behavior's state transition graph is searched in the abnormal behavior rule base. Each abnormal behavior rule corresponds to a specific abnormal behavior type identifier, which is extracted and used as the abnormal behavior type identifier for the current behavior. For example, the abnormal behavior rule base may define abnormal behavior types such as "fighting" and "damaging public property," each with a corresponding identifier. When the current behavior matches the "fighting" abnormal behavior rule, the abnormal behavior type identifier corresponding to "fighting" is generated.

[0144] Step S4100: Generate real-time behavior monitoring results containing behavior warning strategies based on the risk level corresponding to the abnormal behavior type identifier.

[0145] The risk level corresponding to the abnormal behavior type identifier is a pre-defined risk level for each abnormal behavior type, which can be divided into high, medium, and low levels. Behavior warning strategies are response measures developed for different abnormal behavior types and risk levels. For example, for high-risk abnormal behavior, it may be necessary to immediately notify security personnel for on-site intervention; for low-risk abnormal behavior, it may only require handling through broadcast reminders. Real-time behavior monitoring results are a comprehensive summary of information including abnormal behavior type, risk level, and behavior warning strategies, providing a basis for subsequent behavior management and intervention.

[0146] In practice, the system searches for the corresponding risk level in the abnormal behavior rule base based on the abnormal behavior type identifier. The rule base explicitly defines the risk level for each abnormal behavior type; for example, "fighting" might be set as a high-risk level, while "loud noise" might be set as a low-risk level. Based on the risk level, a corresponding behavior warning strategy is selected from a pre-set behavior warning strategy library. This library contains specific intervention measures for different risk levels and abnormal behavior types. For example, for high-risk abnormal behavior, there might be a warning strategy of "immediately notifying security personnel to rush to the scene to stop and handle the situation"; for low-risk abnormal behavior, there might be a warning strategy of "reminding relevant personnel to maintain order via campus broadcasts."

[0147] By integrating information such as abnormal behavior types, risk levels, and behavioral warning strategies, real-time behavior monitoring results are generated. These results can be presented in report form for easy viewing and processing by relevant personnel. For example, the real-time behavior monitoring results might display "Abnormal behavior type: fighting, risk level: high, behavioral warning strategy: immediately notify security personnel to rush to the scene to stop and handle the situation."

[0148] Step S500: Receive real-time behavior monitoring results through edge computing nodes and adaptively adjust the acquisition parameters of the video data stream based on a dynamic priority strategy.

[0149] Edge computing nodes are devices with computing and data processing capabilities distributed throughout the campus monitoring area, responsible for receiving real-time behavior monitoring results sent by the central server. The dynamic priority strategy is a strategy that dynamically adjusts the priority of video data stream acquisition parameters based on information such as risk level in the real-time behavior monitoring results. Video data stream acquisition parameters include frame rate, resolution, and transmission bandwidth; adaptive adjustment automatically adjusts these acquisition parameters according to real-time conditions to meet different monitoring needs.

[0150] In practice, edge computing nodes receive real-time behavior monitoring results from the central server via the network. Based on the abnormal behavior type identifiers in the real-time behavior monitoring results, the risk level coefficient of the current monitored area is determined. The risk level coefficient is a quantitative indicator used to represent the degree of risk in the current monitored area; different abnormal behavior types correspond to different risk level coefficients. Based on the risk level coefficient and the resource load status of the edge computing nodes, the priority weights of the video data stream's acquisition frame rate, resolution, and transmission bandwidth are calculated.

[0151] Video acquisition tasks on multiple edge computing nodes are scheduled based on priority weights, with edge computing nodes at preset priority weights being allocated preset computing resources. The adjusted acquisition parameters are synchronized to all edge computing nodes, and the acquisition task queue for the video data stream is updated.

[0152] As one implementation method, step S500 involves receiving real-time behavior monitoring results through an edge computing node and adaptively adjusting the acquisition parameters of the video data stream based on a dynamic priority strategy. Specifically, step S510 involves determining the risk level coefficient of the current monitoring area based on the abnormal behavior type identifier in the real-time behavior monitoring results.

[0153] An abnormal behavior type identifier is a symbol or code used to identify the type of abnormal behavior. Different abnormal behavior types have different levels of risk. The risk level coefficient is a quantitative indicator used to represent the risk level of the current monitored area. It can be found in the abnormal behavior rule base based on the abnormal behavior type identifier.

[0154] In practice, after receiving real-time behavior monitoring results, the edge computing node extracts the abnormal behavior type identifier. It then searches the abnormal behavior rule base for the corresponding risk level coefficient. For example, the abnormal behavior rule base specifies a risk level coefficient of 0.8 for "fighting" and 0.2 for "loud noise." When the abnormal behavior type identifier in the real-time behavior monitoring results is "fighting," the risk level coefficient for the current monitored area is determined to be 0.8.

[0155] Step S520: Based on the risk level coefficient and the resource load status of the edge computing node, calculate the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream.

[0156] The risk level coefficient indicates the risk level of the current monitored area, while the resource load status of the edge computing node reflects the node's current computing and processing capabilities. Priority weights are numerical values ​​used to determine the importance of the video data stream's acquisition frame rate, resolution, and transmission bandwidth. By calculating priority weights, resources can be allocated rationally to meet different monitoring needs.

[0157] In practical implementation, a multi-objective optimization function is constructed with resource utilization, behavior detection accuracy, and response latency as constraints. Resource utilization refers to the usage of computing and storage resources of edge computing nodes; behavior detection accuracy refers to the accuracy of behavior detection and recognition; and response latency is the time delay from data acquisition to the output of the processing result. The multi-objective optimization function is solved using the gradient descent algorithm to obtain a Pareto optimal solution set for the acquisition frame rate, resolution, and transmission bandwidth. The Pareto optimal solution set is the set of solutions that, under the constraints, cannot improve a certain objective without reducing other objectives. Based on the ratio of behavior detection accuracy to resource consumption corresponding to each solution in the Pareto optimal solution set, solutions that satisfy preset balance conditions are selected as priority weights. The preset balance conditions can be set according to actual needs, such as requiring behavior detection accuracy to reach a certain level while keeping resource consumption within an acceptable range.

[0158] As one implementation method, step S520 may specifically include: step S521: constructing a multi-objective optimization function with resource utilization, behavior detection accuracy and response latency as constraints.

[0159] Resource utilization is the proportion of computing and storage resources used by an edge computing node; behavior detection accuracy is the ability to accurately detect and identify behavior; and response latency is the time taken from data acquisition to the output of the processing result. A multi-objective optimization function is a function that comprehensively considers resource utilization, behavior detection accuracy, and response latency, aiming to find the optimal video data stream acquisition parameters while satisfying these constraints.

[0160] In practical implementation, resource utilization can be set as U, behavior detection accuracy as A, and response latency as D. The multi-objective optimization function F can be expressed as: ,in These are weighting coefficients, representing the importance of resource utilization, behavior detection accuracy, and response latency in the function, and can be adjusted according to actual needs. These are sub-functions relating to resource utilization, behavior detection accuracy, and response latency, for example... It can be a function that is inversely proportional to resource utilization rate, to ensure that resource utilization rate is within a reasonable range; It can be a function that is proportional to the accuracy of behavior detection, in order to improve the accuracy of behavior detection; This could be a function inversely proportional to response latency, in order to reduce it. Simultaneously, constraints need to be set for resource utilization, behavior detection accuracy, and response latency; for example, requiring resource utilization to not exceed a certain threshold U. max The behavior detection accuracy is not lower than a certain threshold A. min The response delay does not exceed a certain threshold D max .

[0161] Step S522: Solve the multi-objective optimization function using the gradient descent algorithm to obtain the Pareto optimal solution set for the acquisition frame rate, resolution, and transmission bandwidth.

[0162] Gradient descent is an optimization algorithm used to find the minimum value of a function. The Pareto optimal set is the set of solutions in a multi-objective optimization problem where it is impossible to improve a particular objective without sacrificing other objectives. By solving a multi-objective optimization function using gradient descent, a Pareto optimal set satisfying the constraints can be found, thereby determining the optimal values ​​for the capture frame rate, resolution, and transmission bandwidth of a video data stream.

[0163] In practical implementation, the initial values ​​of the acquisition frame rate, resolution, and transmission bandwidth can be set as the initial solution for the algorithm. Then, the gradient of the multi-objective optimization function at the current solution is calculated; the gradient represents the rate of change of the function at that point. Based on the direction of the gradient, the values ​​of the acquisition frame rate, resolution, and transmission bandwidth are updated, causing the value of the multi-objective optimization function to gradually decrease. This process is repeated until a convergence condition is met, such as the magnitude of the gradient being less than a preset threshold. During the solution process, it is necessary to ensure that the updated solution satisfies constraints on resource utilization, behavior detection accuracy, and response latency. The final set of solutions is the Pareto optimal solution set for the acquisition frame rate, resolution, and transmission bandwidth.

[0164] Step S523: Based on the behavior detection accuracy and resource consumption ratio of each solution in the Pareto optimal solution set, select the solution that satisfies the preset balance condition as the priority weight.

[0165] The behavior detection accuracy to resource consumption ratio is the ratio of behavior detection accuracy to resource consumption for a given solution. It reflects the behavior detection accuracy achievable with a certain amount of resources. The preset balance condition is a pre-defined condition used to select solutions that satisfy the balance between behavior detection accuracy and resource consumption. Priority weights are determined based on the selected solutions and are used to adjust the acquisition frame rate, resolution, and transmission bandwidth of the video data stream. In practice, each solution in the Pareto optimal solution set is traversed, and its corresponding behavior detection accuracy to resource consumption ratio is calculated. The preset balance condition can require the behavior detection accuracy to resource consumption ratio to reach a certain threshold, or it can require that behavior detection accuracy and resource consumption remain balanced within a certain range. Solutions that satisfy the preset balance condition are selected as priority weights. For example, if the preset balance condition requires the behavior detection accuracy to resource consumption ratio to be no less than 0.8, then solutions with a behavior detection accuracy to resource consumption ratio of no less than 0.8 are selected from the Pareto optimal solution set, and their corresponding acquisition frame rate, resolution, and transmission bandwidth values ​​are used as priority weights.

[0166] Step S530: Schedule video acquisition tasks of multiple edge computing nodes according to priority weights, and allocate preset computing resources to edge computing nodes with preset priority weights.

[0167] Priority weights represent the importance of the video data stream's capture frame rate, resolution, and transmission bandwidth. By scheduling video capture tasks across multiple edge computing nodes according to priority weights, computing resources can be allocated rationally, ensuring clearer and more accurate video data is obtained in high-risk areas. Preset priority weights are pre-defined weight values ​​used to distinguish the priorities of different edge computing nodes, while preset computing resources are the amount of computing resources allocated to edge computing nodes based on their priority weights.

[0168] In practice, multiple edge computing nodes are sorted according to priority weights, with higher-priority edge computing nodes ranked first. Edge computing nodes within a preset priority weight range are allocated preset computing resources. For example, if the preset priority weights define the top 30% of edge computing nodes as high-priority nodes, then these 30% of edge computing nodes will be allocated more computing resources, such as higher CPU utilization and larger memory space, to ensure they can acquire video data at higher frame rates, resolutions, and transmission bandwidths.

[0169] Step S540: Synchronize the adjusted acquisition parameters to all edge computing nodes and update the acquisition task queue of the video data stream.

[0170] The adjusted acquisition parameters are determined based on priority weights and resource allocation, including the acquisition frame rate, resolution, and transmission bandwidth of the video data stream. Synchronization involves sending the adjusted acquisition parameters to all edge computing nodes, enabling them to acquire video data according to the new parameters. Updating the video data stream's acquisition task queue involves rearranging the video acquisition tasks of the edge computing nodes based on the adjusted acquisition parameters to ensure smooth acquisition. In practice, the adjusted acquisition parameters can be sent to all edge computing nodes via the network. Upon receiving the adjusted acquisition parameters, the edge computing nodes update their own acquisition configuration and acquire video data according to the new acquisition frame rate, resolution, and transmission bandwidth. Simultaneously, they update the video data stream's acquisition task queue, rearranging the execution order and timing of acquisition tasks according to the new acquisition parameters. For example, if the acquisition frame rate of an edge computing node increases, the time interval of its acquisition tasks needs to be adjusted accordingly to ensure acquisition at the new frame rate.

[0171] As one implementation, the method provided by the present invention may further include the following derivative steps: Step S600: Construct a behavior pattern evolution graph in the central server, the behavior pattern evolution graph including historical behavior pattern nodes, real-time behavior pattern nodes and pattern transition edges.

[0172] A behavior pattern evolution graph is a graphical structure used to represent the evolution of behavior patterns over time. It helps analyze the trends and patterns of behavior pattern changes. Historical behavior pattern nodes are nodes representing behavior patterns that have appeared over a past period, with each node representing a specific behavior pattern. Real-time behavior pattern nodes are nodes representing behavior patterns that have appeared at the current moment. Pattern transition edges are edges connecting historical and real-time behavior pattern nodes, representing the transition relationships between behavior patterns.

[0173] In the specific construction process, different behavioral patterns can be extracted from historical behavioral data and used as historical behavioral pattern nodes. Each historical behavioral pattern node can be represented by a feature vector, which contains information such as the spatiotemporal features and semantics of the behavioral pattern. For real-time behavioral pattern nodes, the current behavioral pattern is extracted by real-time monitoring and analysis of the current behavioral data and used as the real-time behavioral pattern node. Then, the relationship between historical and real-time behavioral pattern nodes is analyzed to determine the pattern transition edges between them. The pattern transition edges can be weighted, with the weight representing the probability or frequency of the behavioral pattern transition. For example, if the frequency of transitioning from historical behavioral pattern A to real-time behavioral pattern B is high, the weight of the pattern transition edge A->B can be set to a larger value.

[0174] Step S700: Update the real-time behavior pattern node based on the real-time behavior monitoring results, and calculate the pattern similarity between the real-time behavior pattern node and the historical behavior pattern node.

[0175] Real-time behavior monitoring results contain behavioral information at the current moment, which can be used to update the characteristics of real-time behavior pattern nodes. Pattern similarity is an indicator used to measure the degree of similarity between real-time behavior pattern nodes and historical behavior pattern nodes. By calculating pattern similarity, we can understand the differences and connections between current and historical behavior patterns.

[0176] In practical implementation, the feature vectors of real-time behavior pattern nodes are updated based on information such as the semantic description sequence of behavior and key behavior action nodes in the real-time behavior monitoring results. Various methods can be used to calculate the pattern similarity between real-time behavior pattern nodes and historical behavior pattern nodes, such as cosine similarity and Euclidean distance. Taking cosine similarity as an example, let the feature vector of the real-time behavior pattern node be X, and the feature vector of the historical behavior pattern node be Y, then their cosine similarity S can be expressed as... ,in This represents the dot product of vectors X and Y. and Let X and Y represent the magnitudes of vectors X and Y, respectively. By calculating the pattern similarity between real-time behavior pattern nodes and all historical behavior pattern nodes, a similarity list can be obtained.

[0177] Step S800: When the pattern similarity is continuously lower than the preset threshold, the incremental learning process of the target behavior recognition model is triggered.

[0178] The preset threshold is a pre-defined value used to determine whether the pattern similarity is too low. When the pattern similarity between the real-time behavior pattern node and the historical behavior pattern node is continuously lower than the preset threshold, it indicates that the current behavior pattern is significantly different from the historical behavior pattern. The target behavior recognition model may not be able to accurately identify these new behavior patterns. At this time, it is necessary to trigger the incremental learning process to update and optimize the target behavior recognition model.

[0179] Specifically, the incremental learning process may include: Step S810: Collect video data stream segments corresponding to the current real-time behavior pattern node.

[0180] The video data stream segments corresponding to the real-time behavior pattern nodes are video data segments containing the current real-time behavior patterns. These segments record detailed information about newly emerging behavior patterns and are an important data source for incremental learning.

[0181] In practice, based on the timestamp information of the real-time behavior pattern node, the corresponding video data segment is extracted from the video data stream collected by the edge computing node. The time period containing the real-time behavior pattern can be located by querying the timestamp index of the video data, and then a video data stream segment within that time period is extracted. For example, if the time range of the real-time behavior pattern node is from frame 100 to frame 200, then the segment from frame 100 to frame 200 is extracted from the video data stream as the video data stream segment corresponding to the current real-time behavior pattern node.

[0182] Step S820: Re-extract features from video data stream segments to generate an incremental training sample set.

[0183] Feature re-extraction involves re-extracting features from the collected video data stream segments to obtain more accurate and representative features. The incremental training sample set is a collection of samples obtained after feature re-extraction, which will be used for incremental learning of the target behavior recognition model.

[0184] In specific implementation, the same method as in step S100 is used to extract features from video data stream segments. Specifically, frame sequence segmentation and spatiotemporal feature extraction are performed on the video data stream segments to generate an initial behavioral feature set. Then, multi-dimensional correlation analysis is performed on the initial behavioral feature set according to the method in step S200 to extract spatiotemporal behavioral feature vectors corresponding to the target monitoring scene. The extracted spatiotemporal behavioral feature vectors are used as samples in the incremental training sample set. Each sample can carry a corresponding behavioral label, which can be determined based on the behavioral semantic description sequence in the real-time behavioral monitoring results.

[0185] Step S830: Fine-tune the parameters of the target behavior recognition model based on the incremental training sample set, and update the pattern transition edge weights in the behavior pattern evolution graph.

[0186] Parameter fine-tuning involves making minor adjustments to the parameters of the target behavior recognition model based on an incremental training sample set, enabling it to better identify newly emerging behavior patterns. Updating the pattern transition edge weights in the behavior pattern evolution graph involves adjusting the weights of the pattern transition edges based on the relationship between newly emerging behavior patterns and historical behavior patterns, thus reflecting the evolution of behavior patterns.

[0187] In practice, the incremental training sample set is input into the target behavior recognition model, and optimization algorithms, such as stochastic gradient descent, are used to fine-tune the model's parameters. During fine-tuning, the model's loss function is calculated, and the model's parameters are updated based on the gradient of the loss function, making the model's output closer to the true label of the sample. For the weights of pattern transition edges in the behavior pattern evolution graph, the weights of the pattern transition edges are adjusted based on the pattern similarity and transition frequency between newly emerging real-time behavior pattern nodes and historical behavior pattern nodes. For example, if a newly emerging real-time behavior pattern has a high similarity to a certain historical behavior pattern and the transition frequency increases, the weight of the pattern transition edge between these two nodes is increased accordingly. Through parameter fine-tuning, the target behavior recognition model can better adapt to newly emerging behavior patterns, improving the accuracy of behavior recognition; by updating the weights of the pattern transition edges, the behavior pattern evolution graph can more accurately reflect the evolution trend of behavior patterns.

[0188] Please see Figure 2 , Figure 2This is a schematic diagram illustrating the composition of a smart campus security early warning system according to an embodiment of the present invention. The system includes edge computing nodes 100 and a central server 300 that are interconnected via a network 200. The number of edge computing nodes 100 can be one or more. The memory of each edge computing node 100 and the central server 300 stores a computer program. When the processors of the edge computing nodes 100 and the central server 300 respectively load and execute the corresponding computer program, the smart campus security early warning method based on edge computing and big data provided in this embodiment of the invention is implemented.

Claims

1. A smart campus security early warning method based on edge computing and big data, characterized in that, The method includes: acquiring multiple video data streams in real time through edge computing nodes deployed in the campus monitoring area, performing frame sequence segmentation and spatiotemporal feature extraction on the video data streams to generate an initial behavioral feature set; performing multi-dimensional correlation analysis on the initial behavioral feature set based on preset behavioral semantic tags to extract spatiotemporal behavioral feature vectors corresponding to the target monitoring scene, and transmitting the spatiotemporal behavioral feature vectors to the central server; specifically including: extracting a set of semantic keywords matching the target monitoring scene from the preset behavioral semantic tags, and assigning a scene association weight coefficient to each semantic keyword; performing semantic similarity matching between each behavioral feature in the initial behavioral feature set and the set of semantic keywords, and filtering those with similarity exceeding the threshold. A subset of candidate behavioral features exceeding a preset threshold is selected. Cross-correlation analysis of the behavioral features within this subset is performed using temporal and spatial dimensions to generate a joint feature map of temporal continuity distribution and spatial density heatmap. Based on the scene correlation weight coefficient, the feature regions in the joint feature map are dynamically weighted, enhancing the weight values ​​of feature regions strongly correlated with the target monitoring scene and suppressing the weight values ​​of irrelevant feature regions. The weighted joint feature map is then input into a spatiotemporal convolution kernel for local feature aggregation, extracting feature segment sequences with temporal dependence and spatial clustering. Based on the cumulative distribution of the feature segment sequences on the time axis and the overlap ratio of the spatial regions, a spatiotemporal behavioral feature vector integrating multi-dimensional correlations is generated. The spatiotemporal behavior feature vector contains spatiotemporal feature information of behavior in the target monitoring scenario. In the central server, the initial behavior recognition model is trained in multiple stages based on the historical behavior data set to obtain the target behavior recognition model. The multi-stage training includes feature fusion based on spatiotemporal correlation and cross-scenario behavior pattern transfer. Specifically, it includes: extracting a first training set with spatiotemporal annotations and a second training set of unannotated cross-scenario data from the historical behavior data set. Each sample in the first training set contains an annotated spatiotemporal behavior feature vector and a corresponding normal behavior type label, while each sample in the second training set contains unannotated video segments from different monitoring scenarios. The first training set is then input into the central server. An initial behavior recognition model aligns the spatiotemporal behavior feature vectors across frames using a spatiotemporal correlation constraint module to generate a first intermediate feature representation with temporal continuity. This first intermediate feature representation is then input into a spatiotemporal feature fusion module, which dynamically assigns weights and superimposes features within adjacent time windows to generate a fused second intermediate feature representation. Based on the similarity loss between the second intermediate feature representation and the normal behavior type label, the initial behavior recognition model undergoes supervised training to obtain a first intermediate model. Finally, the second training set is input into the first intermediate model, and a cross-scene behavior pattern transfer module extracts scene-invariant features from unlabeled video segments to generate a third intermediate feature representation unrelated to the monitoring scene.The third intermediate feature representation is subjected to temporal dimension compression and feature reconstruction to generate a cross-scene behavior pattern embedding vector. The first intermediate model is then self-supervised based on the reconstruction error to obtain a second intermediate model. The feature extraction layer parameters of the first intermediate model and the feature reconstruction layer parameters of the second intermediate model are alternately weighted and fused to generate a target behavior recognition model with cross-scene adaptability. The spatiotemporal behavior feature vector is input into the target behavior recognition model to generate a behavior semantic description sequence corresponding to the video data stream. Specifically, this includes: performing time-axis sliding window segmentation on the spatiotemporal behavior feature vector to generate a feature segment sequence corresponding to continuous time segments; inputting the feature segment sequence into the basic action parsing layer of the target behavior recognition model to extract the limb displacement rate change pattern and direction consistency features of the moving target within the time segment, generating an atomic action label sequence; and inputting the atomic action label sequence into the interaction relationship inference layer of the target behavior recognition model, based on the atomic actions... The temporal overlap and spatial proximity between tags are used to identify collaborative action patterns among multiple moving targets, generating an interaction relationship description set. This interaction relationship description set is input into the intent inference layer of the target behavior recognition model. Combined with the scene context features corresponding to the time segments, the logical matching degree between the collaborative action patterns and preset behavioral intent templates is analyzed, generating a behavioral intent probability distribution. Based on the atomic action tag sequence, the interaction relationship description set, and the behavioral intent probability distribution, a structured behavioral semantic description tree containing action levels, interaction levels, and intent levels is generated. The structured behavioral semantic description tree undergoes temporal compression and semantic aggregation, removing redundant description nodes and merging similar action branches to generate the behavioral semantic description sequence. Based on the matching results of the behavioral semantic description sequence and a preset abnormal behavior rule base, real-time behavior monitoring results are determined. The real-time behavior monitoring results are received through the edge computing node, and the acquisition parameters of the video data stream are adaptively adjusted based on a dynamic priority strategy.

2. The method according to claim 1, characterized in that, The step of performing frame sequence segmentation and spatiotemporal feature extraction on the video data stream to generate an initial behavioral feature set includes: performing frame segmentation on the video data stream to obtain a continuous video frame sequence, and filtering noise from the video frame sequence based on the pixel change rate between adjacent video frames; performing dynamic region detection on the filtered video frame sequence to determine local image regions containing moving targets, and performing multi-scale spatiotemporal convolution processing on the local image regions; extracting motion trajectory features, posture change features, and environmental interaction features from the local image regions after multi-scale spatiotemporal convolution processing, and aligning the motion trajectory features, posture change features, and environmental interaction features on the time axis; and performing weighted fusion of the aligned features according to a preset spatiotemporal weight distribution matrix to generate the initial behavioral feature set.

3. The method according to claim 2, characterized in that, The step of extracting motion trajectory features, pose change features, and environmental interaction features of the local image region after multi-scale spatiotemporal convolution processing, and aligning these features along the time axis, includes: dividing the local image region after multi-scale spatiotemporal convolution processing into time segments to generate a set of time segments containing start and end timestamps; performing trajectory tracking processing on each time segment in the time segment set to obtain a continuous displacement coordinate sequence of the moving target within the local image region, and generating the motion trajectory features based on the rate of change of the vector direction of adjacent coordinates in the displacement coordinate sequence; and performing keypoint detection processing on each time segment in the time segment set to extract the spatial position sequence of the moving target on skeletal joints, and generating the pose change based on the relative displacement of the spatial position sequence between adjacent time segments. Features: For each time segment in the time segment set, target relationship parsing is performed to identify the spatial distance sequence and contact state change sequence between the moving target and surrounding static objects, and the environmental interaction features are generated based on the coupling relationship between the spatial distance sequence and the contact state change sequence; according to the start and end timestamps corresponding to each time segment in the time segment set, the motion trajectory features, posture change features and environmental interaction features are time-synchronized to generate time-aligned versions of the motion trajectory features, posture change features and environmental interaction features with a unified time reference; based on the temporal continuity of the motion trajectory features, posture change features and environmental interaction features in the time-aligned version, adaptive trajectory interpolation compensation is performed on the feature data corresponding to missing timestamps to generate the motion trajectory features, posture change features and environmental interaction features after complete time axis alignment.

4. The method according to claim 1, characterized in that, The step of aligning spatiotemporal behavior feature vectors across frames using a spatiotemporal correlation constraint module to generate a first intermediate feature representation with temporal continuity includes: extracting feature subsequences corresponding to consecutive time segments from the spatiotemporal behavior feature vectors and calculating the dynamic time warping distance between feature subsequences of adjacent time segments; performing nonlinear interpolation on the feature subsequences based on the dynamic time warping distance to generate a fourth intermediate feature representation aligned with the time axis; slicing the fourth intermediate feature representation temporally to generate local feature blocks within multiple overlapping time windows; assigning learnable spatiotemporal attention weights to each local feature block and performing feature enhancement on the local feature blocks based on the spatiotemporal attention weights to generate a fifth intermediate feature representation; inputting the fifth intermediate feature representation into a bidirectional temporal convolutional network to capture future contextual dependencies through forward propagation and historical contextual dependencies through backpropagation to generate a sixth intermediate feature representation with bidirectional temporal correlation; and performing a residual connection between the sixth intermediate feature representation and the original spatiotemporal behavior feature vector to generate the first intermediate feature representation.

5. The method according to claim 1, characterized in that, The step of determining the real-time behavior monitoring result based on the matching result between the behavior semantic description sequence and the preset abnormal behavior rule base includes: performing semantic parsing on the behavior semantic description sequence to extract key behavior action nodes and the temporal dependencies between nodes; constructing a behavior state transition graph based on the temporal dependencies, and calculating the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule base; when the topological similarity exceeds a preset threshold, generating an abnormal behavior type identifier that matches the abnormal behavior rule base; and generating the real-time behavior monitoring result containing a behavior warning strategy based on the risk level corresponding to the abnormal behavior type identifier.

6. The method according to claim 5, characterized in that, The step of semantically parsing the behavioral semantic description sequence to extract key behavioral action nodes and temporal dependencies between nodes includes: dividing the behavioral semantic description sequence into multiple continuous basic semantic units according to a preset time granularity, and performing redundancy filtering on the basic semantic units based on a semantic coherence threshold to generate a set of deredundant semantic units; performing action keyword matching on each semantic unit in the set of deredundant semantic units to identify candidate semantic units containing target action words from a preset behavioral action dictionary, and mapping the candidate semantic units to initial behavioral action nodes; and clustering and merging the initial behavioral action nodes within adjacent time windows based on the distribution density of the initial behavioral action nodes on the time axis to generate a set of temporally clustered nodes. The system identifies key behavioral action nodes; it performs dependency analysis on the time intervals between these nodes, and if the time interval between two key behavioral action nodes is less than a preset action association threshold, it establishes a temporal dependency edge between them; based on the directionality of the temporal dependency edge, it determines the sequential triggering relationship or concurrent relationship between the key behavioral action nodes, generating a set of temporal dependency relationships with temporal tags; it performs context consistency verification on the temporal dependency edges in the set of temporal dependency relationships, and if there are dependency edges that conflict with the temporal logic verified in the historical behavior pattern, it removes the conflicting dependency edges and recalculates the temporal dependency relationships, generating optimized key behavioral action nodes and the temporal dependency relationships between them.

7. The method according to claim 5, characterized in that, The step of constructing a behavior state transition graph based on the temporal dependency relationship and calculating the topological similarity between the behavior state transition graph and the reference state transition graph in the abnormal behavior rule base includes: mapping the key behavior action nodes to state nodes in the state transition graph, and mapping the temporal dependency edges in the temporal dependency relationship to state transition edges to generate an initial behavior state transition graph; labeling each state transition edge in the initial behavior state transition graph with time constraints, extracting the time interval range of the corresponding temporal dependency edge in the temporal dependency relationship, and using the time interval range as the time constraint condition of the state transition edge; assigning a transition probability weight to each state transition edge according to the time constraint condition of the state transition edge and the state transition frequency in the historical behavior pattern to generate a behavior with temporal constraints and probability attributes. A state transition graph is generated. A reference state transition graph matching the current monitoring scenario is extracted from the abnormal behavior rule base. The reference state transition graph includes a predefined set of abnormal state nodes and a set of abnormal state transition edges. Semantic alignment is performed between the state nodes in the behavior state transition graph and the abnormal state nodes in the reference state transition graph to identify pairs of state nodes with the same behavioral semantic description. Based on the semantic alignment results of the state node pairs, subgraph matching is performed between the behavior state transition graph and the reference state transition graph to extract their structural differences in state transition edge direction, time constraints, and transition probability weights. The topological coverage between the behavior state transition graph and the reference state transition graph is calculated based on the structural differences, and a topological similarity is generated based on the comparison result between the topological coverage and a preset abnormal threshold.

8. The method according to claim 1, characterized in that, The step of receiving the real-time behavior monitoring results through the edge computing nodes and adaptively adjusting the acquisition parameters of the video data stream based on a dynamic priority strategy includes: determining the risk level coefficient of the current monitoring area based on the abnormal behavior type identifier in the real-time behavior monitoring results; calculating the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream based on the risk level coefficient and the resource load status of the edge computing nodes; scheduling the video acquisition tasks of multiple edge computing nodes according to the priority weights, allocating preset computing resources to edge computing nodes with preset priority weights; synchronizing the adjusted acquisition parameters to all edge computing nodes and updating the acquisition task queue of the video data stream; wherein, calculating the priority weights of the acquisition frame rate, resolution, and transmission bandwidth of the video data stream includes: constructing a multi-objective optimization function with resource utilization, behavior detection accuracy, and response latency as constraints; solving the multi-objective optimization function using a gradient descent algorithm to obtain a Pareto optimal solution set for the acquisition frame rate, resolution, and transmission bandwidth; and selecting the solution that satisfies the preset balance condition as the priority weight based on the behavior detection accuracy and resource consumption ratio corresponding to each solution in the Pareto optimal solution set.

9. A smart campus security early warning system, characterized in that, The system includes edge computing nodes and a central server that are interconnected, and the memory of each edge computing node and the central server stores a computer program. When the processors of the edge computing nodes and the central server load and execute the corresponding computer programs, the smart campus security early warning method based on edge computing and big data as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Campus abnormal behavior analysis system and method

    CN115100572A

  • Method for training video label recommendation model, and method for determining video label

    WO2023273769A1