Network edge monitoring and early warning method based on video image AI analysis
By using the combination of the hypergraph Transformer network and the improved butterfly optimization algorithm in network edge monitoring and early warning, a spatio-temporal hypergraph structure is constructed and an abnormal behavior feedback mechanism is used to solve the problems of insufficient detection accuracy and high response delay in the existing technology, and efficient and accurate abnormal behavior detection and real-time early warning are achieved.
Patent Information
- Application Number
- CN202510300486.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as insufficient detection accuracy, high response delay and limited ability to capture complex spatiotemporal information in network edge monitoring and early warning, especially when processing large-scale real-time video data.
An innovative combination of hypergraph Transformer network and improved butterfly optimization algorithm is adopted to accurately extract the multi-scale spatiotemporal features of the target area in the video image by constructing a spatiotemporal hypergraph structure, and anomaly detection is achieved using anomaly behavior feedback mechanism.
It realizes comprehensive capture of complex space-time dynamic information, has the ability to respond in real time, accurately identify abnormal behaviors, and adaptive optimization, which significantly improves detection accuracy and response speed.
Smart Images

Figure CN120220062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AI technology, and in particular, to a network edge monitoring and early warning method based on AI analysis of video images. Background Art
[0002] With the rapid development of information technology and artificial intelligence technology, video image analysis has been increasingly widely used in fields such as security monitoring, intelligent transportation, industrial automation, and network edge monitoring. In the prior art, video image-based monitoring and early warning systems mainly rely on traditional convolutional neural networks (CNNs), region proposal networks (RPNs), and deep learning-based object detection algorithms such as Faster R-CNN, YOLO series, etc. These methods have achieved certain success in static image object detection. However, in the field of network edge monitoring and early warning, especially when high-efficiency analysis of large-scale real-time video data and detection of abnormal behaviors are required, traditional methods have problems such as insufficient detection accuracy, high response latency, and limited ability to capture complex spatio-temporal information.
[0003] When processing video image data, the prior art usually only focuses on single-frame or local temporal information, lacking effective modeling of the high-order spatio-temporal correlation relationships between multiple targets. Traditional methods are difficult to comprehensively capture the complex interactions and spatio-temporal dynamic changes between target regions in the video, resulting in missed detections and false detections in the identification of abnormal behaviors in the network edge monitoring environment. On the other hand, due to limited computing resources of network edge devices, existing deep learning models often have difficulty achieving real-time response while ensuring high accuracy, further limiting their application effects in network security early warning.
[0004] In addition, in the prior art, traditional gradient descent, genetic algorithms, particle swarm optimization, etc. are generally used to update model parameters and construct hyperedge structures in the optimization algorithms. Although these algorithms can achieve adaptive adjustment of the model to a certain extent, they are prone to falling into local optima, have a slow convergence speed, and lack an effective mechanism for dynamically adjusting the network structure when dealing with complex high-dimensional data. Therefore, how to improve the extraction accuracy of multi-target behavior features in video images and the response speed of abnormal behavior detection under limited edge computing resources has become an urgent technical problem to be solved.
[0005] In recent years, in response to the above problems, scholars have begun to explore ways to combine hypergraph theory with Transformer networks to construct a more accurate spatio-temporal hypergraph structure, so as to better capture the high-order spatio-temporal semantic relationships in the target regions of video images. At the same time, in view of the limitations of traditional optimization algorithms in the global search and parameter update processes, some studies have introduced swarm intelligence algorithms such as the butterfly optimization algorithm, and combined innovative mechanisms such as chaotic mapping, gradient sensitivity factors, and inertial perturbations to enhance the dynamics of model parameter updates and global search capabilities. These innovative technologies have to some extent improved the model's ability to express complex spatio-temporal features and the recognition effect of abnormal behaviors, but there are still deficiencies.
[0006] Traditional hypergraph construction methods usually only utilize simple spatial or temporal features and cannot fully integrate multi-scale spatio-temporal information simultaneously; while traditional attention mechanisms often ignore the high-order associations between nodes in the hypergraph structure when capturing long-range dependence relationships, resulting in insufficient information transmission. In addition, the butterfly optimization algorithm lacks effective integration of abnormal behavior feedback information in traditional applications and is difficult to achieve real-time adaptive adjustment of the hypergraph Transformer network model parameters and hyperedge selection strategies. Therefore, how to design an innovative network edge monitoring and early warning method based on video image AI analysis to make full use of the hypergraph Transformer network and the improved butterfly optimization algorithm, and achieve accurate detection and real-time early warning of target behaviors through multi-scale spatio-temporal information fusion and abnormal behavior feedback mechanisms has become an important direction of current technical research.
[0007] Therefore, how to provide a network edge monitoring and early warning method based on video image AI analysis is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0008] An object of the present invention is to propose a network edge monitoring and early warning method based on video image AI analysis. The present invention makes full use of the innovative combination of the hypergraph Transformer network and the butterfly optimization algorithm, accurately extracts the multi-scale spatio-temporal features of the target area in the video image by constructing a spatio-temporal hypergraph, and realizes anomaly detection by adopting an improved attention mechanism and an abnormal behavior feedback mechanism. Specifically, the present invention first converts video data into an image frame sequence, obtains the spatial and temporal features of each target area by using an object detection algorithm, then takes these target areas as spatio-temporal hypergraph nodes, and realizes the dynamic modeling of the high-order spatio-temporal semantic relationship between nodes by constructing a multi-scale hypergraph Laplacian matrix and spatio-temporal relative position encoding. At the same time, the present invention uses an improved butterfly optimization algorithm, calculates the butterfly perception value by introducing a hybrid chaos factor and a gradient sensitive factor, and combines an abnormal behavior clustering feedback mechanism to adaptively optimize the network parameters and the hyperedge selection strategy, so as to continuously improve the accuracy of abnormal behavior detection. In the implementation process of this method, it not only ensures the comprehensive capture of complex spatio-temporal dynamics in video images, but also has the capabilities of real-time response, accurate identification of abnormal behaviors, and self-adaptive optimization. Compared with the prior art, the present invention has significant advantages such as fast response speed, high detection accuracy, strong adaptability to complex network edge scenarios, and excellent model self-learning ability.
[0009] A network edge monitoring and early warning method based on video image AI analysis according to an embodiment of the present invention includes the following steps:
[0010] S1. Obtain video image data of the network edge scenario through a camera, perform frame segmentation, image enhancement, noise removal, and edge detection on the video image data to obtain an image frame sequence;
[0011] S2. Analyze the image frame sequence, identify the target areas in the image frame sequence, and extract spatial and temporal features;
[0012] S3. Take the target areas as hypergraph nodes, construct hyperedges according to the spatial and temporal features of the hypergraph nodes, and dynamically adjust the connection mode of the hyperedges through a genetic algorithm to form a spatio-temporal hypergraph structure;
[0013] S4. Construct a hypergraph Transformer network based on a multi-layer Transformer architecture, extract the multi-scale spatio-temporal features of the hypergraph nodes through a self-attention mechanism, calculate the semantic relationship between the hypergraph nodes, and obtain behavior feature data;
[0014] S5. Set the initial population of the butterfly optimization algorithm, each butterfly represents a different hypergraph Transformer network model parameter configuration, and dynamically optimize the hypergraph Transformer network model parameters through a global search and a local search mechanism;
[0015] S6. Construct an abnormal behavior detection model based on the behavioral characteristic data, conduct initial training using the normal behavior dataset, identify abnormal behaviors, and feedback the identification results to the butterfly optimization algorithm to dynamically adjust the hypergraph Transformer network model parameters and hyperedge selection strategy;
[0016] S7. Deploy the optimized hypergraph Transformer network model and abnormal behavior detection model on the network edge computing node. When an abnormal behavior is detected, trigger a warning signal through the warning system and push the abnormal event information to the management platform or security system.
[0017] Optionally, the S2 specifically includes:
[0018] S21. Select the CenterNet object detection model, load the pre-trained CenterNet object detection model weights, set the detection categories of the CenterNet object detection model, and configure the prediction threshold, learning rate, and batch size of the CenterNet object detection model;
[0019] S22. Input the image frame sequence frame by frame into the CenterNet object detection model, perform multi-scale feature analysis on the image through the feature extraction network, and generate the feature map data of the potential target regions in the image frame;
[0020] S23. According to the feature map data, the CenterNet object detection model predicts the center point position of the target region in the image through the key point detection method, combines the width and height offset information to generate the bounding box of the target region, and outputs the category, position coordinates, and confidence information of the target;
[0021] S24. Extract the spatial features of the identified target regions, including the coordinate position, size, shape features, and relative distance between different target regions of the target, to obtain the spatial feature data of the target regions;
[0022] S25. Track the target regions in the image frame sequence, and extract the temporal feature data of the target by analyzing the motion changes of the target in consecutive image frames.
[0023] Optionally, the S3 specifically includes:
[0024] S31. Take each identified target region as a hypergraph node. Each hypergraph node contains spatial features and temporal features, and assign a unique identifier ID to each hypergraph node;
[0025] S32. Calculate the spatial feature similarity between hypergraph nodes. By comparing the spatial distance, size difference, and shape features between hypergraph nodes, when the spatial feature similarity exceeds the preset spatial similarity threshold, construct the spatial hyperedges between hypergraph nodes;
[0026] S33. Analyze the time features of the hypergraph nodes. By calculating the similarity of the movement trajectories, the speed differences, and the changes in the movement directions of the nodes, when the time feature similarity exceeds the preset time similarity threshold, construct the time hyperedges between the spatio-temporal hypergraph nodes;
[0027] S34. Integrate the analysis results of the spatial features and the time features. When the spatial feature similarity between the hypergraph nodes exceeds the spatial similarity threshold T s and the time feature similarity exceeds the time similarity threshold T t calculate the spatio-temporal composite similarity S ij . When the spatio-temporal composite similarity exceeds the preset composite similarity threshold, construct the candidate spatio-temporal composite hyperedges:
[0028]
[0029] where d ij represents the spatial distance between the hypergraph nodes, d t,ij represents the time difference between the hypergraph nodes, T s and T t are the preset thresholds, λ and μ are the adjustment parameters, and exp() is the exponential function;
[0030] S35. Perform dynamic selection using the genetic algorithm. Randomly generate the initial population of the candidate hyperedges, then evaluate each candidate hyperedge individual according to the preset fitness function, select the hyperedges with fitness higher than the preset threshold, and add the hyperedges to the spatio-temporal hypergraph structure to finally generate the spatio-temporal hypergraph structure:
[0031]
[0032] where F(E) is the fitness function, E is the set of candidate hyperedges, K is the total number of hypergraph nodes, and ρ is the hyperedge number penalty factor.
[0033] Optionally, the specific steps of S4 include:
[0034] S41. Construct the spatio-temporal hypergraph node set N = {n1, n2, …, n K}, where each node n i has an initial feature vector f i ∈R d , forming the initial feature matrix F:
[0035]
[0036] where K represents the total number of spatio-temporal hypergraph nodes, d represents the total number of features, and R is the set of real numbers;
[0037] S42. Map the initial feature matrix F through a pre-trained linear projection to map the feature vectors of the nodes into three different subspaces respectively, generating a query matrix Q, a key matrix K, and a value matrix V;
[0038] S43. For each scale of the hypergraph Laplacian matrix extracted from the hypergraph structure, introduce a learnable weight δ i and a non-linear enhancement coefficient κ i , and use adaptive normalization to construct a multi-scale enhancement matrix:
[0039]
[0040] where 1 represents a vector of all 1s, ⊙ represents element-wise multiplication, ReLU() is the activation function, M represents the total number of scales of the hypergraph Laplacian matrix, softmax() represents the normalization operation, LN() represents the layer normalization function, L enh is the multi-scale enhancement matrix, κ i is the non-linear enhancement coefficient, δ i is the learnable weight;
[0041] S44. Let P S and P T be the spatial position and temporal position of the hypergraph nodes respectively. Generate a spatial encoding R S and a temporal encoding R T through non-linear mapping functions f S = f S (P S ) and f T = f T (P T ) respectively. Adopt a fusion strategy to construct a spatio-temporal relative position encoding matrix R:
[0042] R = LN(ηR S ⊙ sigmoid(R T ) + ζR S + ξR T );
[0043] where η, ζ, ξ are learnable fusion parameters, and sigmoid() is the activation function;
[0044] S45. Adopt the hypergraph attention mechanism to fuse and calculate the attention weights of the query matrix Q, the key matrix K, the value matrix V, the multi-scale enhancement matrix L enh and the spatio-temporal position encoding R:
[0045]
[0046] where β and γ are scaling factors, A is the attention weight matrix, dh is the projection dimension, and the output behavior feature matrix F is obtained through weighted fusion ′ :
[0047] F ′ = LN(AV + F)+ FFN(LN(AV + F));
[0048] Among them, FFN() is a feed-forward neural network for further non-linear mapping to obtain behavior feature data of multi-scale spatio-temporal fusion.
[0049] Optionally, the S5 specifically includes:
[0050] S51. The parameters of the hypergraph Transformer network are encoded into the candidate parameter vector p i ∈R n , where n represents the parameter dimension, R is the set of real numbers, and the initial population is randomly generated:
[0051] P = {p1, p2,..., p N};
[0052] Among them, N is the number of candidate parameters;
[0053] S52. For each candidate parameter vector p i , the fitness function is defined as F(p i ):
[0054]
[0055] Among them, L(p i ) represents the performance loss based on the hypergraph Transformer network, C(p i ) represents the penalty term, and α1 is the complexity penalty factor;
[0056] S53. For the candidate parameter vector, the chaos factor φ(t) is generated by the hybrid chaos mapping method in each iteration:
[0057] φ(t)= θ·sin(πφ(t - 1))+(1 - θ)·μφ(t - 1)(1 - φ(t - 1));
[0058] Among them, θ is the chaos weighting coefficient, μ is the logarithmic mapping parameter, φ(t - 1) represents the chaos factor at the (t - 1)-th iteration, sin() is the sine function, and the butterfly perception value f i is calculated by combining the candidate parameter fitness and the loss function gradient information:
[0059]
[0060] Among them, c is the scaling constant, α2 is the fitness sensitivity index, and γ is the gradient sensitivity factor, Denote the candidate parameter as p i is the gradient magnitude of the corresponding loss function, δ is the chaos factor modulation index, and exp() is the exponential function;
[0061] S54. Update the candidate parameter by fusing chaos factor modulation and inertia perturbation:
[0062]
[0063] wherein, is the updated value of the i-th candidate parameter at the (t + 1)-th iteration, is the candidate parameter at the t-th iteration, g is the candidate parameter vector with the highest fitness in the current population, α3 is a random variable, and ω is the inertia weight coefficient;
[0064] S55. After setting the maximum number of iterations T or satisfying the preset convergence condition, select the candidate parameter vector p with the highest fitness from the population P * , and update the model parameters of the hypergraph Transformer network.
[0065] Optionally, the S6 specifically includes:
[0066] S61. Calculate the anomaly error of each target by using the output behavioral feature data and the reference normal behavioral features obtained from the normal behavior dataset through the Euclidean distance and combined with the logarithmic compression method;
[0067] S62. Calculate the anomaly score A i based on the obtained anomaly error E i :
[0068]
[0069] wherein, τ is the anomaly sensitivity parameter, τ1 is the smoothing adjustment parameter, κ1, κ2, κ3 are the modulation indices, exp() is the exponential function, and tanh() is the hyperbolic tangent function;
[0070] S63. Construct the overall anomaly detection loss function to measure the anomaly deviation of all targets:
[0071]
[0072] wherein, K is the number of targets, μ1 is the balance parameter, μ2 is the anomaly error penalty coefficient, and log2 is the logarithmic function;
[0073] S64. Group the anomaly scores A of all targets i by using the clustering method, and the clustering result is C j , where each cluster C jContaining K j targets, calculate the average anomaly score for each cluster Then construct the feedback weight factor W f :
[0074]
[0075] where ε1 is the scaling parameter, ε2 is the gradient smoothing parameter, and U is the total number of clusters;
[0076] S65. Integrate the feedback weight factor W f into the adaptive update process of the hypergraph Transformer network parameters and the hyperedge selection strategy, update the network parameters, and feedback the anomaly behavior recognition result to the butterfly optimization algorithm to dynamically adjust the network structure and the hyperedge construction strategy.
[0077] The beneficial effects of the present invention are as follows:
[0078] The present invention adopts an innovative combination of a hypergraph Transformer network and a butterfly optimization algorithm, and has achieved remarkable beneficial effects in the field of network edge monitoring and early warning. By performing in-depth image analysis on video data, the present invention can accurately extract multi-scale spatio-temporal features of multiple target regions in the video and construct a spatio-temporal hypergraph structure reflecting the high-order semantic associations of the target regions, enabling the entire system to far exceed traditional methods in capturing complex spatio-temporal dynamic information. Using the hypergraph Transformer network, the present invention realizes the adaptive modeling of the subtle relationships between target regions, can not only take into account the fine expression of local features but also effectively integrate global information, thereby providing a more comprehensive and accurate feature representation for the subsequent detection of abnormal behaviors.
[0079] On this basis, the present invention further introduces an improved butterfly optimization algorithm. Through innovative mechanisms such as hybrid chaotic mapping, gradient-sensitive factor, and inertial perturbation, it realizes the dynamic adaptive optimization of the hypergraph Transformer network parameters and the hyperedge selection strategy. This optimization mechanism can respond in real time to the emergence of abnormal behaviors in video surveillance data and use the anomaly behavior clustering feedback mechanism to quickly adjust the global network parameters, enabling the system to still maintain high accuracy and high response speed in the changing network edge environment. At the same time, by constructing a complex anomaly detection loss function and a feedback weight factor, the present invention effectively feeds back the anomaly behavior recognition result into the butterfly optimization algorithm, further improving the self-learning ability of the network, so that it can self-correct and continuously improve the detection accuracy during the detection process.
[0080] Therefore, the present invention not only greatly improves the detection and feature extraction accuracy of the target area in video images, but also realizes the fast and accurate recognition of abnormal behaviors through the fusion of multi-scale spatio-temporal information and dynamic parameter optimization. In the application of network edge monitoring and early warning, the system can process a large amount of video data in real time, locate and identify abnormal behaviors in a short time, while reducing the false alarm rate and missed alarm rate, effectively improving the intelligent level and security protection ability of the entire monitoring system. In addition, since the system uses an edge computing architecture for data preprocessing and analysis, its response speed is much better than the traditional centralized processing method, and it can quickly send early warning signals to the management platform and security system, greatly shortening the response time.
[0081] In summary, through the close combination of the deep vision algorithm and the intelligent optimization algorithm, the present invention overcomes the deficiencies of the prior art in spatio-temporal feature capture, abnormal detection, and parameter adaptive update, achieving a real-time, high-precision, and self-learning early warning effect. This method not only improves the overall security and stability of the video monitoring system, but also shows good scalability and robustness in practical applications, providing strong technical support for the further development of the network edge security monitoring and early warning system. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0083] Figure 1 is a flowchart of a network edge monitoring and early warning method based on video image AI analysis proposed by the present invention;
[0084] Figure 2 is a schematic diagram of the feedback mechanism of the butterfly optimization algorithm in abnormal behavior detection of a network edge monitoring and early warning method based on video image AI analysis proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0086] Refer to Figure 1 and Figure 2 , a network edge monitoring and early warning method based on video image AI analysis, includes the following steps:
[0087] S1. Obtain video image data of the network edge scene through a camera, perform frame segmentation, image enhancement, noise removal, and edge detection on the video image data to obtain an image frame sequence;
[0088] S2. Analyze the image frame sequence, identify the target regions in the image frame sequence, and extract spatial and temporal features;
[0089] S3. Take the target regions as hypergraph nodes, construct hyperedges according to the spatial and temporal features of the hypergraph nodes, and dynamically adjust the connection mode of the hyperedges through a genetic algorithm to form a spatio-temporal hypergraph structure;
[0090] S4. Construct a hypergraph Transformer network based on a multi-layer Transformer architecture, extract multi-scale spatio-temporal features of the hypergraph nodes through a self-attention mechanism, calculate the semantic relationships between the hypergraph nodes, and obtain behavioral feature data;
[0091] S5. Set the initial population of the butterfly optimization algorithm, where each butterfly represents a different configuration of the hypergraph Transformer network model parameters, and dynamically optimize the hypergraph Transformer network model parameters through global search and local search mechanisms;
[0092] S6. According to the behavioral feature data, construct an abnormal behavior detection model, conduct initial training using a normal behavior dataset, identify abnormal behaviors, and feedback the identification results to the butterfly optimization algorithm to dynamically adjust the hypergraph Transformer network model parameters and hyperedge selection strategies;
[0093] S7. Deploy the optimized hypergraph Transformer network model and abnormal behavior detection model on network edge computing nodes. When an abnormal behavior is detected, trigger a warning signal through a warning system and push the abnormal event information to the management platform or security system.
[0094] In this embodiment, the S2 specifically includes:
[0095] S21. Select the CenterNet object detection model, load the pre-trained CenterNet object detection model weights, set the detection categories of the CenterNet object detection model, and configure the prediction threshold, learning rate, and batch size of the CenterNet object detection model;
[0096] S22. Input the image frame sequence frame by frame into the CenterNet object detection model, perform multi-scale feature analysis on the image through a feature extraction network, and generate feature map data of potential target regions in the image frames;
[0097] S23. According to the feature map data, the CenterNet object detection model predicts the center point positions of the target regions in the image through a key point detection method, combines the width and height offset information to generate the bounding boxes of the target regions, and outputs the category, position coordinates, and confidence information of the targets;
[0098] S24. Extract spatial features of the identified target regions, including the coordinate positions, sizes, shape features of the targets, and the relative distances between different target regions, to obtain the spatial feature data of the target regions;
[0099] S25. Track the target regions in the image frame sequence, and extract the temporal feature data of the targets by analyzing the motion changes of the targets in consecutive image frames.
[0100] In this embodiment, the specific steps of S3 are as follows:
[0101] S31. Take each identified target region as a hypergraph node. Each hypergraph node contains spatial features and temporal features, and assign a unique identifier ID to each hypergraph node;
[0102] S32. Calculate the spatial feature similarity between hypergraph nodes. By comparing the spatial distances, size differences, and shape features between hypergraph nodes, when the spatial feature similarity exceeds the preset spatial similarity threshold, construct spatial hyperedges between hypergraph nodes;
[0103] S33. Analyze the temporal features of hypergraph nodes. By calculating the similarity of motion trajectories, speed differences, and changes in motion directions of nodes, when the temporal feature similarity exceeds the preset temporal similarity threshold, construct temporal hyperedges between spatio-temporal hypergraph nodes;
[0104] S34. Integrate the analysis results of spatial features and temporal features. When the spatial feature similarity between hypergraph nodes exceeds the spatial similarity threshold T s and the temporal feature similarity exceeds the temporal similarity threshold T t calculate the spatio-temporal composite similarity S ij . When the spatio-temporal composite similarity exceeds the preset composite similarity threshold, construct candidate spatio-temporal composite hyperedges:
[0105]
[0106] where d ij represents the spatial distance between hypergraph nodes, d t,ij represents the temporal difference between hypergraph nodes, T s and T t are preset thresholds, λ and μ are adjustment parameters, and exp() is an exponential function;
[0107] S35. Perform dynamic selection using a genetic algorithm. Randomly generate an initial population of candidate hyperedges, then evaluate each candidate hyperedge individual according to the preset fitness function, select the hyperedges with fitness higher than the preset threshold, add the hyperedges to the spatio-temporal hypergraph structure, and finally generate the spatio-temporal hypergraph structure:
[0108]
[0109] Among them, F(E) is the fitness function, E is the set of candidate hyperedges, K is the total number of hypergraph nodes, and ρ is the hyperedge number penalty factor.
[0110] In this embodiment, the S4 specifically includes:
[0111] S41. Construct a spatio-temporal hypergraph node set N = {n1, n2,..., n K}, where each node n i has an initial feature vector f i ∈ R d to form an initial feature matrix F:
[0112]
[0113] Among them, K represents the total number of spatio-temporal hypergraph nodes, d represents the total number of features, and R is the set of real numbers;
[0114] S42. Pass the initial feature matrix F through a pre-trained linear projection to map the feature vectors of the nodes into three different subspaces respectively, generating a query matrix Q, a key matrix K, and a value matrix V;
[0115] S43. {L (i)} i M =1 are hypergraph Laplacian matrices of each scale extracted from the hypergraph structure. Introduce a learnable weight δ i and a non-linear enhancement coefficient κ i , and use adaptive normalization to construct a multi-scale enhancement matrix:
[0116]
[0117] Among them, 1 represents a vector of all 1s, ⊙ represents element-wise multiplication, ReLU() is the activation function, M represents the total number of scales of the hypergraph Laplacian matrix, softmax() represents the normalization operation, LN() represents the layer normalization function, L enh is the multi-scale enhancement matrix, κ i is the non-linear enhancement coefficient, and δ i is the learnable weight;
[0118] S44. Let P S and P T be the spatial position and temporal position of the hypergraph node respectively. Generate a spatial encoding R S and f T respectively through non-linear mapping functions f S = f S (P S) and the time encoding R T = f T (P T ), a spatio-temporal relative position encoding matrix R is constructed using a fusion strategy:
[0119] R = LN(ηR S ⊙ sigmoid(R T ) + ζR S + ξR T );
[0120] Among them, η, ζ, and ξ are learnable fusion parameters, and sigmoid() is an activation function;
[0121] S45. Adopt a hypergraph attention mechanism to fuse and calculate the attention weights of the query matrix Q, the key matrix K, the value matrix V, the multi-scale enhancement matrix L enh and the spatio-temporal position encoding R:
[0122]
[0123] Among them, β and γ are scaling factors, A is the attention weight matrix, d h is the projection dimension, and the behavior feature matrix F ′ is output through weighted fusion:
[0124] F ′ = LN(AV + F) + FFN(LN(AV + F));
[0125] Among them, FFN() is a feed-forward neural network for further non-linear mapping to obtain multi-scale spatio-temporal fusion behavior feature data.
[0126] In this embodiment, the S5 specifically includes:
[0127] S51. Encode the hypergraph Transformer network parameters into a candidate parameter vector p i ∈ R n , where n represents the parameter dimension, R is the set of real numbers, and an initial population is randomly generated:
[0128] P = {p1, p2,..., p N};
[0129] Among them, N is the number of candidate parameters;
[0130] S52. For each candidate parameter vector p i , define the fitness function as F(p i ):
[0131]
[0132] Among them, L(pi ) represents the performance loss based on the hypergraph Transformer network, C(p i ) represents the penalty term, and α1 is the complexity penalty factor;
[0133] S53. For the candidate parameter vector, generate the chaos factor φ(t) through the hybrid chaos mapping method in each iteration:
[0134] φ(t) = θ·sin(πφ(t - 1)) + (1 - θ)·μφ(t - 1)(1 - φ(t - 1));
[0135] where θ is the chaos weighting coefficient, μ is the logarithmic mapping parameter, φ(t - 1) represents the chaos factor at the (t - 1)-th iteration, sin() is the sine function, and combining the candidate parameter fitness and the loss function gradient information, calculate the butterfly perception value f i :
[0136]
[0137] where c is the scaling constant, α2 is the fitness sensitivity index, γ is the gradient sensitivity factor, represents the candidate parameter p i corresponding to the gradient magnitude of the loss function, δ is the chaos factor modulation index, and exp() is the exponential function;
[0138] S54. Update the candidate parameters by fusing chaos factor modulation and inertia perturbation:
[0139]
[0140] where, is the updated value of the i-th candidate parameter at the (t + 1)-th iteration, is the candidate parameter at the t-th iteration, g is the candidate parameter vector with the highest fitness in the current population, α3 is a random variable, and ω is the inertia weight coefficient;
[0141] S55. After setting the maximum number of iterations T or meeting the preset convergence condition, select the candidate parameter vector p with the highest fitness from the population P * , and update the model parameters of the hypergraph Transformer network.
[0142] In this embodiment, the S6 specifically includes:
[0143] S61. Use the output behavioral feature data and the reference normal behavioral features obtained from the normal behavior dataset to calculate the anomaly error of each target through the Euclidean distance and in combination with the logarithmic compression method;
[0144] S62. Based on the obtained anomaly error Ei , calculate the anomaly score A i :
[0145]
[0146] where τ is the anomaly sensitivity parameter, τ1 is the smoothing adjustment parameter, κ1, κ2, κ3 are modulation exponents, exp() is the exponential function, and tanh() is the hyperbolic tangent function;
[0147] S63. Construct the overall anomaly detection loss function Measure the anomaly deviation of all targets:
[0148]
[0149] where K is the number of targets, μ1 is the balance parameter, μ2 is the anomaly error penalty coefficient, and log2 is the logarithmic function;
[0150] S64. Use the clustering method to group the anomaly scores A of all targets i The clustering result is C j , where each cluster C j contains K j targets, calculate the average anomaly score of each cluster Then construct the feedback weight factor W f :
[0151]
[0152] where ε1 is the scaling parameter, ε2 is the gradient smoothing parameter, and U is the total number of clusters;
[0153] S65. Integrate the feedback weight factor W f into the adaptive update process of the hypergraph Transformer network parameters and the hyperedge selection strategy, update the network parameters, and feedback the anomaly behavior recognition result to the butterfly optimization algorithm to dynamically adjust the network structure and the hyperedge construction strategy.
[0154] Example 1:
[0155] To verify the feasibility of the present invention in implementation, the present invention is applied to a smart city network edge security monitoring and early warning system. This area includes the junction of the industrial park and the residential area at the edge of the city. In this area, the traditional monitoring system often fails to identify abnormal behaviors in real time due to signal transmission delay and insufficient device resources, resulting in frequent security hazards. The present invention adopts a network edge monitoring and early warning method based on video image AI analysis, uses a hypergraph Transformer network to efficiently extract multi-scale spatio-temporal features of multiple target regions in the video, and adaptively optimizes the model parameters and hyperedge selection strategy through an improved butterfly optimization algorithm, so as to achieve rapid identification and real-time early warning of abnormal behaviors.
[0156] In practical applications, first, high-resolution cameras are deployed at the junction of the industrial park and the residential area. These cameras are directly connected to the edge computing nodes, collect video data in real time, and convert the video into a continuous sequence of image frames. After image preprocessing, pedestrians, vehicles, and other suspicious targets in the video are identified through deep learning object detection algorithms, and multi-dimensional spatial and temporal features including target position, size, shape, as well as motion trajectory, speed, and direction are extracted. Subsequently, the system takes these target regions as spatio-temporal hypergraph nodes and constructs a hypergraph structure reflecting the high-order spatio-temporal correlations between nodes, and uses the hypergraph Transformer network to fuse the features of each node and extract high-order behavior features. Compared with the traditional system that only relies on static image detection, this method can capture the dynamic behavior features of the target, significantly improving the accuracy and real-time performance of abnormal behavior detection.
[0157] To verify the superiority of this method, in this embodiment, a three-month actual deployment test is carried out in the industrial park edge monitoring and early warning system. During the test, the method of the present invention is compared with the traditional method combining the YOLO object detection algorithm with simple time series analysis. The data table details the key performance indicators of the system in different time periods, including average response time, abnormal detection accuracy, false alarm rate, missed alarm rate, and successful intervention rate of abnormal events. The test results show that the method of the present invention is superior to the traditional method in all key indicators, especially in the scenarios of low illuminance at night and high traffic flow, with obvious advantages in detection accuracy and response speed.
[0158] Table 1 Performance comparison data table of the network edge monitoring and early warning system
[0159]
[0160] As can be seen from the table, the method of the present invention is significantly superior to the traditional method in all key performance indicators. During the daytime peak period, the average response time of the traditional method is 350 milliseconds, while the method of the present invention only requires 210 milliseconds, and the response speed is increased by about 40%. At the same time, the detection accuracy rate of the method of the present invention in this scenario reaches 97%, and the false alarm rate and missed alarm rate are controlled at 5% and 3% respectively. Compared with the 90% detection accuracy rate, 10% false alarm rate and 8% missed alarm rate of the traditional method, the method of the present invention shows higher detection accuracy and lower error rate. In addition, in terms of the successful intervention rate of abnormal events, the method of the present invention reaches 85%, which is significantly higher than 60% of the traditional method.
[0161] During the low-light period at night, the method of the present invention still maintains a high detection accuracy rate, reaching 95%, while the traditional method is only 88%. Due to insufficient light at night, traditional vision algorithms are usually vulnerable to noise interference, resulting in relatively high false alarm rate and missed alarm rate, which are 12% and 10% respectively. However, the method of the present invention effectively improves the robustness of abnormal behavior detection by fusing multi-scale spatio-temporal features and an adaptive optimization algorithm, reduces the false alarm rate and missed alarm rate to 5% and 3% respectively, and improves the abnormal intervention success rate to 82%, significantly improving the effect of night monitoring.
[0162] During the non-peak period, the response time of the traditional method is 300 milliseconds, while the method of the present invention is shortened to 190 milliseconds, showing excellent real-time response ability even in a low-load environment. In addition, in terms of the overall average performance, the indicators of the traditional method are relatively low, the detection accuracy rate is only 90%, and the false alarm rate and missed alarm rate are 10% and 8% respectively, while the method of the present invention shows great advantages with a detection accuracy rate of 96%, a false alarm rate of 5% and a missed alarm rate of 3%. The overall abnormal intervention success rate is also increased from 60% of the traditional method to 80%.
[0163] In summary, through the comparison data of different scenarios, it can be seen that the method of the present invention not only shows higher abnormal detection accuracy in peak, night and low-load scenarios, but also can complete the response in a shorter time. This result benefits from the present invention extracting multi-scale spatio-temporal features through a hypergraph Transformer network and combining an improved butterfly optimization algorithm for adaptive adjustment of model parameters and hyperedge selection strategies, enabling the system to always be in the best detection state, thus achieving a faster and more accurate network edge monitoring and early warning effect.
[0164] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. A network edge monitoring and early warning method based on video image AI analysis, characterized in that: The steps include: S1. Obtain video image data of a network edge scene through a camera, perform frame segmentation, image enhancement, noise removal and edge detection on the video image data to obtain an image frame sequence; S2, analyzing the image frame sequence, identifying the target area in the image frame sequence, and extracting spatial and temporal features; S3, taking the target area as a hypergraph node, constructing a hyperedge according to the spatial and temporal characteristics of the hypergraph node, and dynamically adjusting the connection mode of the hyperedge through a genetic algorithm to form a spatiotemporal hypergraph structure; S4. Build a hypergraph Transformer network based on a multi-layer Transformer architecture, extract multi-scale spatiotemporal features of hypergraph nodes through the self-attention mechanism, calculate the semantic relationship between hypergraph nodes, and obtain behavioral feature data; S5. Set the initial population of the butterfly optimization algorithm. Each butterfly represents a different hypergraph Transformer network model parameter configuration. The hypergraph Transformer network model parameters are dynamically optimized through global search and local search mechanisms. S6. Build an abnormal behavior detection model based on the behavior feature data, use the normal behavior data set for initial training, identify abnormal behaviors, and feed the identification results back to the butterfly optimization algorithm to dynamically adjust the hypergraph Transformer network model parameters and hyperedge selection strategy; S7. Deploy the optimized hypergraph Transformer network model and abnormal behavior detection model on the network edge computing node. When abnormal behavior is detected, the early warning signal is triggered through the early warning system, and the abnormal event information is pushed to the management platform or security system.
2. According to claim 1, a network edge monitoring and early warning method based on video image AI analysis is characterized in that: The S3 specifically includes: S21. Select the CenterNet target detection model, load the pre-trained CenterNet target detection model weights, set the detection category of the CenterNet target detection model, and configure the prediction threshold, learning rate, and batch size of the CenterNet target detection model. S22, inputting the image frame sequence into the CenterNet target detection model frame by frame, performing multi-scale feature analysis on the image through the feature extraction network, and generating feature map data of the potential target area in the image frame; S23. Based on the feature map data, the CenterNet target detection model predicts the center point position of the target area in the image through the key point detection method, combines the width and height offset information, generates the bounding box of the target area, and outputs the target category, position coordinates and confidence information; S24, extracting spatial features of the identified target area, including the coordinate position, size, shape features of the target and the relative distance between different target areas, to obtain spatial feature data of the target area; S25. Track the target area in the image frame sequence, and extract the time feature data of the target by analyzing the motion changes of the target in the continuous image frames.
3. According to claim 1, a network edge monitoring and early warning method based on video image AI analysis is characterized in that: The S3 specifically includes: S31, taking each identified target area as a hypergraph node, each hypergraph node contains spatial features and temporal features, and assigning a unique identifier ID to each hypergraph node; S32, calculating the spatial feature similarity between the hypergraph nodes, by comparing the spatial distance, size difference and shape features between the hypergraph nodes, when the spatial feature similarity exceeds a preset spatial similarity threshold, constructing the spatial hyperedge between the hypergraph nodes; S33, analyzing the time characteristics of the hypergraph nodes, by calculating the similarity of the motion trajectories, speed differences, and motion direction changes of the nodes, when the time characteristic similarity exceeds a preset time similarity threshold, constructing a time hyperedge between the spatiotemporal hypergraph nodes; S34, the analysis results of the comprehensive spatial features and temporal features, when the spatial feature similarity between the hypergraph nodes exceeds the spatial similarity threshold T s When the temporal feature similarity exceeds the temporal similarity threshold T t When , calculate the spatiotemporal composite similarity S ij , when the spatiotemporal composite similarity exceeds the preset composite similarity threshold, a candidate spatiotemporal composite hyperedge is constructed: Among them, d ij represents the spatial distance between hypergraph nodes, d t,ij represents the time difference between hypergraph nodes, T s and T t is the preset threshold, λ and μ are adjustment parameters, and exp() is the exponential function; S35. Use a genetic algorithm for dynamic selection to randomly generate an initial population of candidate hyperedges, then evaluate each candidate hyperedge individual according to a preset fitness function, select hyperedges with a fitness higher than a preset threshold, add the hyperedges to the spatiotemporal hypergraph structure, and finally generate a spatiotemporal hypergraph structure: Among them, F(E) is the fitness function, E is the candidate hyperedge set, K is the total number of hypergraph nodes, and ρ is the penalty factor for the number of hyperedges.
4. According to claim 1, a network edge monitoring and early warning method based on video image AI analysis is characterized in that: The S4 specifically includes: S41, construct a spatiotemporal hypergraph node set N = {n1, n2, ..., n K }, where each node n i With an initial feature vector f containing spatial and temporal features i ∈R d , forming the initial feature matrix F: Where K represents the total number of spatiotemporal hypergraph nodes, d represents the total number of features, and R is a set of real numbers; S42, mapping the feature vectors of the nodes into three different subspaces respectively through pre-trained linear projection of the initial feature matrix F, generating a query matrix Q, a key matrix K and a value matrix V; S43, For each scale, a learnable weight δ is introduced to extract the Laplacian matrix of the hypergraph at each scale from the hypergraph structure. i and the nonlinear enhancement coefficient κ i , and use adaptive normalization to construct a multi-scale enhancement matrix: Among them, 1 represents a vector of all 1s, ⊙ represents element-by-element multiplication, ReLU() is the activation function, M represents the total number of scales of the hypergraph Laplacian matrix, softmax() represents the normalization operation, LN() represents the layer normalization function, and L enh is the multi-scale enhancement matrix, κ i is the nonlinear enhancement coefficient, δ i is the learnable weight; S44, let P S and P T are the spatial position and temporal position of the hypergraph node respectively, through the nonlinear mapping function f S and f T Generate spatial code R S =f S (P S ) and time code R T =f T (P T ), and adopt the fusion strategy to construct the spatiotemporal relative position encoding matrix R: R=LN(ηR S ⊙sigmoid(R T )+ζR S +ξR T ); Among them, η, ζ, ξ are learnable fusion parameters, and sigmoid() is the activation function; S45, using the hypergraph attention mechanism, the query matrix Q, key matrix K, value matrix V, and multi-scale enhancement matrix L enh The attention weights are calculated by fusion of the spatiotemporal position encoding R: Among them, β and γ are scaling factors, A is the attention weight matrix, and d h is the projection dimension, and the behavior feature matrix F is output through weighted fusion ′ : F ′ =LN(AV+F)+FFN(LN(AV+F)); Among them, FFN() is a feedforward neural network, which is used for further nonlinear mapping to obtain multi-scale spatiotemporal fusion behavior feature data.
5. According to claim 1, a network edge monitoring and early warning method based on video image AI analysis is characterized in that: The S5 specifically includes: S51, Hypergraph Transformer network parameters are encoded as candidate parameter vector p i ∈R n , where n represents the parameter dimension, R is a set of real numbers, and the initial population is randomly generated: P={p1,p2,…,p N }; Where N is the number of candidate parameters; S52. For each candidate parameter vector p i , define the fitness function as F(p i ): Among them, L(p i ) represents the performance loss based on the hypergraph Transformer network, C(p i ) represents the penalty term, α1 is the complexity penalty factor; S53. For the candidate parameter vector, a chaotic factor φ(t) is generated by a hybrid chaotic mapping method in each iteration: φ(t)=θ·sin(πφ(t-1))+(1-θ)·μφ(t-1)(1-φ(t-1)); Among them, θ is the chaos weight coefficient, μ is the logarithmic mapping parameter, φ(t-1) represents the chaos factor at the t-1th iteration, sin() is the sine function, and the butterfly perception value f is calculated by combining the candidate parameter fitness and the loss function gradient information. i : Among them, c is the scaling constant, α2 is the fitness sensitivity index, and γ is the gradient sensitivity factor. represents the candidate parameter p i The gradient amplitude corresponding to the loss function, δ is the modulation index of the chaos factor, and exp() is the exponential function; S54. Update the candidate parameters by integrating chaos factor modulation and inertial perturbation: in, is the updated value of the i-th candidate parameter at the t+1-th iteration, is the candidate parameter at the tth iteration, g is the candidate parameter vector with the highest fitness in the current population, α3 is a random variable, and ω is the inertia weight coefficient; S55, after setting the maximum number of iterations T or satisfying the preset convergence condition, select the candidate parameter vector p with the highest fitness from the population P * , update the model parameters of the hypergraph Transformer network.
6. The network edge monitoring and early warning method based on video image AI analysis according to claim 1 is characterized in that: The S6 specifically includes: S61, using the output behavior feature data and the reference normal behavior feature obtained from the normal behavior data set, calculating the abnormal error of each target through Euclidean distance combined with a logarithmic compression method; S62, based on the abnormal error E i , calculate the anomaly score A i : Among them, τ is the abnormal sensitivity parameter, τ1 is the smoothing adjustment parameter, κ1, κ2, κ3 are modulation indexes, exp() is the exponential function, and tanh() is the hyperbolic tangent function; S63. Constructing the overall anomaly detection loss function Measure abnormal deviations from all targets: Among them, K is the target number, μ1 is the balance parameter, μ2 is the abnormal error penalty coefficient, and log2 is the logarithmic function; S64, use clustering method to calculate the abnormal score A of all targets i Grouping, the clustering result is C j , where each cluster C j Contains K j targets, calculate the average anomaly score for each cluster Then construct the feedback weight factor W f : Among them, ε1 is the scaling parameter, ε2 is the gradient smoothing parameter, and U is the total number of clusters; S65, feedback weight factor W f It is integrated into the adaptive update process of the hypergraph Transformer network parameters and hyperedge selection strategy, updates the network parameters, feeds back the abnormal behavior recognition results to the butterfly optimization algorithm, and dynamically adjusts the network structure and hyperedge construction strategy.
Citation Information
Cited By
Equipment support real-time control method and device based on random dynamic algorithm
CN120821198A
Target retrieval method and system
CN120910115A
A target retrieval method and system
CN120910115B
Mine monitoring video key frame extraction method
CN121459254A
Typhoon wind field wind speed space-time prediction method
CN121542747A