Intelligent video monitoring scene image classification method and system based on hypergraph structure learning

By employing a hypergraph structure learning method, a hypergraph association matrix is ​​constructed using the kNN algorithm and multi-view learning. Combined with graph regularization and density-aware attention layers, this approach addresses the issues of inaccurate graph structure and lack of high-order associations in existing technologies, thereby improving the accuracy and efficiency of image classification in intelligent video surveillance scenarios.

CN118429717BActive Publication Date: 2025-11-11SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410613855.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-11
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

Existing intelligent video surveillance scene image classification methods are prone to including incorrect node connections and noise information when constructing graph structures, and most of them only focus on pairwise connections, lacking the ability to mine higher-order relationships, resulting in poor classification performance.

Method used

We employ a hypergraph structure learning approach, constructing the original hypergraph association matrix using the kNN algorithm. By combining multi-view hypergraph learning and a learnable multi-head cosine similarity calculation function, we can mine higher-order associations between data. Furthermore, we utilize graph regularization techniques to optimize the hypergraph structure and enhance node encoding representations through a hypergraph density-aware attention layer and a contrastive learning module.

Benefits of technology

It significantly improves the classification performance of intelligent video surveillance scene image classification models, especially in the case of a few labeled samples, and can more accurately capture the complex correlation information between scene images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118429717B_ABST
    Figure CN118429717B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for classifying intelligent video surveillance scene images based on hypergraph structure learning. The method includes: obtaining a feature matrix and a label matrix from an acquired intelligent video surveillance scene image dataset; constructing an original hypergraph association matrix based on the feature matrix; inputting the feature matrix into a hypergraph structure learning module in a classification model to obtain an implicit hypergraph association matrix through hypergraph learning from multiple views; fusing the original hypergraph association matrix and the implicit hypergraph association matrix to obtain a fused hypergraph association matrix; inputting the feature matrix and the fused hypergraph association matrix into a hypergraph representation learning module in the classification model to obtain a classification prediction result for the scene image; training the classification model using the feature matrix and the label matrix, updating the parameters of the classification model based on a total loss function to obtain a trained classification model; and inputting the feature matrix corresponding to the intelligent video surveillance scene image set to be classified into the trained classification model to obtain a category prediction result. This invention can significantly improve the classification performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image classification technology, and in particular to an intelligent video surveillance scene image classification method, system, terminal device, and computer-readable storage medium based on hypergraph structure learning. Background Technology

[0002] With the rise of the smart city concept globally, my country has also been making every effort to deploy and develop it in recent years, striving to empower urban governance with digital technology and drive high-quality development through data. Simultaneously, leveraging the wave of artificial intelligence development, intelligent video surveillance technology has rapidly advanced, providing more innovative technological support for the construction of smart cities. The application scenarios of intelligent video surveillance are very broad. In the security field, it is a powerful tool for preventing and combating crime, promptly detecting anomalies and issuing alarms by monitoring and tracking the behavior of suspicious persons or objects in real time, effectively assisting security personnel in handling crises. In the field of traffic management, intelligent video surveillance assists traffic management departments in optimizing traffic lights and control strategies through real-time monitoring and analysis of traffic flow. In the field of customer service, intelligent video surveillance can detect and analyze customer behavior, providing enterprises with valuable customer behavior data and market feedback information, helping them optimize products and services. In the data identification, classification, and processing process of intelligent video surveillance, scene image classification plays a crucial role as a fundamental technology. It can intelligently identify and classify event and scene data collected from intelligent video surveillance, thereby improving monitoring efficiency. However, most current mainstream scene image classification technologies tend to process individual scene images independently, ignoring the correlation between image data. Therefore, these methods obtain limited image feature information, which severely limits the performance of scene image classification, especially when the number of labeled samples is small.

[0003] Before the advent of graph neural networks (Graph Neural Networks), traditional deep learning methods achieved great success in extracting features from Euclidean space data. However, many practical applications use data generated from non-Euclidean spaces. Furthermore, a core assumption of traditional deep learning methods is that data samples are independent of each other; however, this is not the case for graph data. With the introduction of Graph Neural Networks, the problems encountered by traditional deep learning with non-Euclidean data have been effectively solved. Graph Neural Networks can fully utilize the inherent topological information of non-Euclidean data to learn richer feature representations. Leveraging this advantage, Graph Neural Networks have rapidly become a popular area in deep learning and are now widely used in computer vision, recommender systems, traffic prediction, and molecular structure analysis. Therefore, applying Graph Neural Networks to scene image classification in intelligent video surveillance can effectively overcome the aforementioned limitations. Graph Neural Networks can not only process the feature information of individual scene image samples but also combine the topological information of the graph structure to mine the correlation information between scene images. Compared to traditional methods, Graph Neural Networks can extract richer information from scene images and effectively improve classification performance in scenarios with only a few labeled samples.

[0004] Currently, in the field of intelligent video scene image classification, there have been some studies based on graph neural networks, but these methods mostly face two shortcomings. First, most methods still use manually constructed graph structures, such as the k-nearest neighbor algorithm, and these initial input graph structures remain unchanged during training. However, the performance of graph neural networks heavily depends on the reliability of the graph structure. If the initial graph structure contains incorrect inter-node connections or other noise information, it will significantly affect the model's classification performance. Furthermore, k-nearest neighbor graphs are mainly based on a fixed, single, and non-learnable similarity metric function, which cannot accurately and effectively measure the similarity between complex data samples. Therefore, accurately constructing a graph structure suitable for scene image classification is a very challenging task. Second, existing graph neural network-based scene image classification methods mostly focus only on low-order relationships such as pairwise connections between data samples. However, in practical applications, the relationships between samples are not limited to pairwise relationships; sometimes they are more complex one-to-many or many-to-many higher-order relationships. If the graph is constructed only based on pairwise relationships between samples, it will lack the ability to mine higher-order semantic relationships between data. Hypergraphs are a generalization of ordinary graphs. Edges in a hypergraph can connect any number of nodes, thus effectively uncovering higher-order relationships between data. Furthermore, previous research based on ordinary graphs has shown that prior knowledge inherent in the graph structure itself can be used to learn and constrain the graph structure; this technique, known as graph regularization, has proven effective. However, applying ordinary graph regularization techniques to the hypergraph domain remains a problem with limited research. Therefore, intelligent video surveillance scene image classification methods based on hypergraph neural networks remain an area to be explored. Summary of the Invention

[0005] To address at least one of the technical problems in the prior art, this invention provides a method, system, terminal device, and computer-readable storage medium for intelligent video surveillance scene image classification based on hypergraph structure learning, which can significantly improve the classification performance of intelligent video surveillance scene image classification models.

[0006] The first objective of this invention is to provide an intelligent video surveillance scene image classification method based on hypergraph structure learning.

[0007] The second objective of this invention is to provide an intelligent video surveillance scene image classification system based on hypergraph structure learning.

[0008] The third objective of this invention is to provide a terminal device.

[0009] A fourth objective of this invention is to provide a computer-readable storage medium.

[0010] The first objective of this invention can be achieved by adopting the following technical solution:

[0011] A method for intelligent video surveillance scene image classification based on hypergraph structure learning, the method comprising:

[0012] Obtain an intelligent video surveillance scene image dataset; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data;

[0013] Based on the scene images and label data in the intelligent video surveillance scene image dataset, the feature matrix and label matrix are obtained;

[0014] The kNN algorithm is used to calculate the similarity between eigenvectors in the feature matrix, and the original hypergraph association matrix is ​​constructed based on the similarity.

[0015] The feature matrix is ​​input into the hypergraph structure learning module in the intelligent video surveillance scene image classification model. The implicit hypergraph association matrix is ​​obtained by learning the hypergraph from multiple views. In each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function.

[0016] The original hypergraph incidence matrix and the implicit hypergraph incidence matrix are merged to obtain the merged hypergraph incidence matrix;

[0017] The feature matrix and the fused hypergraph association matrix are input into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and the node encoding representation of the hypergraph is output. Based on the node encoding representation, the classification prediction result of the scene image is obtained.

[0018] A smart video surveillance scene image classification model is trained using feature matrices and label matrices. The parameters of the smart video surveillance scene image classification model are updated based on the total loss function to obtain the trained smart video surveillance scene image classification model. The total loss function includes the loss function of the hypergraph structure learning module. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, including a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views.

[0019] Input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction result.

[0020] Furthermore, the process of obtaining the implicit hypergraph association matrix through hypergraph learning from multiple views includes:

[0021] In each view, the feature vectors in the feature matrix are mapped to a low-dimensional space, and the similarity between feature vectors in the low-dimensional space is calculated using a learnable multi-head cosine similarity calculation function to obtain the final similarity matrix.

[0022] The similarity matrix is ​​sparsified, and the sparsified similarity matrix is ​​used as a view to learn the hypergraph association matrix.

[0023] The hypergraph incidence matrices learned from all views are fused to obtain the implicit hypergraph incidence matrix.

[0024] Furthermore, before inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module, message propagation on the hypergraph is performed first to obtain the node feature matrix and hyperedge feature matrix of the hypergraph.

[0025] Furthermore, the hypergraph representation learning module employs a hypergraph density-aware attention layer, which consists of a node aggregation module and a hyperedge aggregation module.

[0026] The step of inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module of the intelligent video surveillance scene image classification model to obtain the node encoding representation of the hypergraph includes:

[0027] The hyperedge feature matrix and node feature matrix are input into the node aggregation module. The node information is aggregated onto the hyperedge based on the density-aware attention mechanism to obtain the updated hyperedge feature matrix.

[0028] The hyperedge feature matrix and node feature matrix are input into the hyperedge aggregation module. Hyperedge features are aggregated based on the updated hyperedge feature matrix to obtain the updated node feature matrix. The updated node feature matrix is ​​the output of a density-aware attention layer.

[0029] By employing multiple attention heads, the outputs of each density-aware attention layer are concatenated to obtain the final node encoding representation.

[0030] Furthermore, the step of inputting the hyperedge feature matrix and node feature matrix into the node aggregation module, and using a density-aware attention mechanism to aggregate node information onto the hyperedge to obtain an updated hyperedge feature matrix, specifically includes:

[0031] Calculate the node density; where the node density is the sum of the similarities of neighboring nodes whose similarity to the current node is greater than a preset threshold.

[0032] A shared attention mechanism is used to calculate the attention weights between nodes and hyperedges;

[0033] The node density information is fused with the attention weights to obtain the attention weights; based on the attention weights of all nodes, the attention weight matrix is ​​obtained.

[0034] The node features are aggregated using the node attention weight matrix to obtain the updated hyperedge feature matrix.

[0035] Furthermore, the total loss function also includes a contrastive learning loss function; the contrastive learning loss function is obtained based on the node encoding representation of the original hypergraph association matrix and the node encoding representation of the fused hypergraph association matrix, so as to improve the classification effect of the intelligent video surveillance scene image classification model; wherein, the node encoding representation of the original hypergraph association matrix is ​​obtained by inputting the feature matrix and the original hypergraph association matrix into the hypergraph representation learning module, and the node encoding representation of the fused hypergraph association matrix is ​​obtained by inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module.

[0036] Furthermore, the total loss function also includes the loss function of the hypergraph representation learning module, expressed as:

[0037]

[0038] In the formula, Y ij Let L be the label value of the i-th labeled data sample in the label matrix, corresponding to the j-th label category. Let L be the set of labeled samples in the label matrix, C be the total number of label categories, and O be the total number of categories. ij Let be the predicted value of the i-th data sample for the j-th label category.

[0039] Furthermore, the step of obtaining the feature matrix and label matrix based on the scene images and label data in the intelligent video surveillance scene image dataset includes:

[0040] Feature extraction is performed on each scene image in the intelligent video surveillance scene image dataset to obtain the corresponding feature vector;

[0041] Stack the feature vectors corresponding to all scene images to obtain the feature matrix;

[0042] Each label data is converted into a label vector with the dimension of the total number of label categories and containing only 0s and 1s through one-hot encoding; only the position corresponding to the category is 1, and the rest are 0s;

[0043] Stack the label vectors corresponding to all the label data to obtain a label matrix.

[0044] The second objective of this invention can be achieved by adopting the following technical solution:

[0045] An intelligent video surveillance scene image classification system based on hypergraph structure learning, the system comprising:

[0046] The sample acquisition module is used to acquire a dataset of intelligent video surveillance scene images; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data; based on the scene images and label data in the intelligent video surveillance scene image dataset, a feature matrix and a label matrix are obtained.

[0047] The first construction module is used to calculate the similarity between feature vectors in the feature matrix using the kNN algorithm, and to construct the original hypergraph association matrix based on the similarity.

[0048] The second construction module is used to input the feature matrix into the hypergraph structure learning module in the intelligent video surveillance scene image classification model, and obtain the implicit hypergraph association matrix by learning the hypergraph from multiple views; wherein, in each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function.

[0049] The fusion module is used to fuse the original hypergraph incidence matrix and the implicit hypergraph incidence matrix to obtain the fused hypergraph incidence matrix;

[0050] The classification prediction module is used to input the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and output the node encoding representation of the hypergraph; based on the node encoding representation, the classification prediction result of the scene image is obtained.

[0051] The training module is used to train the intelligent video surveillance scene image classification model using the feature matrix and label matrix. It updates the parameters of the intelligent video surveillance scene image classification model based on the total loss function to obtain the trained intelligent video surveillance scene image classification model. The total loss function includes the loss function of the hypergraph structure learning module. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, and includes a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views.

[0052] The category output module is used to input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction results.

[0053] The third objective of this invention can be achieved by adopting the following technical solution:

[0054] A terminal device includes a processor and a memory for storing processor-executable programs. When the processor executes the program stored in the memory, it implements the above-described intelligent video surveillance scene image classification method based on hypergraph structure learning.

[0055] The fourth objective of this invention can be achieved by adopting the following technical solution:

[0056] A computer-readable storage medium storing a program that, when executed by a processor, implements the above-described intelligent video surveillance scene image classification method based on hypergraph structure learning.

[0057] The present invention has the following advantages over the prior art:

[0058] This invention provides an intelligent video surveillance scene image classification method, system, terminal device, and computer-readable storage medium based on hypergraph structure learning. It utilizes the kNN algorithm to preserve high-order semantic information contained in the feature matrix, and introduces a learnable multi-head cosine similarity calculation function in the multi-view hypergraph structure learning module to comprehensively and accurately calculate the similarity between feature vectors in the feature matrix, thereby mining high-order correlations between data and learning the optimal hypergraph structure representation. This significantly improves the classification performance of the intelligent video surveillance scene image classification model. Furthermore, it continuously optimizes the hypergraph's correlation matrix using the loss function of the hypergraph structure learning module, making it more suitable for downstream scene image classification tasks. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0060] Figure 1 This is a flowchart of the intelligent video surveillance scene image classification method based on hypergraph structure learning according to Embodiment 1 of the present invention;

[0061] Figure 2 This is a schematic diagram of the structure of the intelligent video surveillance scene image classification model in Embodiment 1 of the present invention;

[0062] Figure 3 This is a schematic diagram of the hypergraph structure learning module in Embodiment 1 of the present invention;

[0063] Figure 4 This is a schematic diagram of the structure of the hypergraph representation learning module in Embodiment 1 of the present invention;

[0064] Figure 5 This is a schematic diagram of the structure of the hypergraph comparison learning module in Embodiment 1 of the present invention;

[0065] Figure 6 This is a structural block diagram of the intelligent video surveillance scene image classification system based on hypergraph structure learning according to Embodiment 2 of the present invention;

[0066] Figure 7 This is a structural block diagram of the terminal device according to Embodiment 3 of the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be understood that the specific embodiments described are merely used to explain this application and are not intended to limit this application.

[0068] Example 1:

[0069] The intelligent video surveillance scene image classification method based on hypergraph structure learning provided in this embodiment is based on the PyTorch deep learning framework version 2.0.0 and the VS Code development environment version 1.87, and the development language is Python version 3.10. PyTorch deep learning can easily and quickly build deep learning neural networks, and at the same time, it can accelerate the model training process by taking advantage of the machine's built-in GPU hardware resources.

[0070] like Figure 1 As shown, the intelligent video surveillance scene image classification method based on hypergraph structure learning provided in this embodiment specifically includes the following steps:

[0071] S101. Based on the acquired intelligent video surveillance scene image dataset, obtain the feature matrix and label matrix.

[0072] Obtain a dataset of intelligent video surveillance scene images, in which only some scene images have corresponding label data; based on the scene images and label data in the intelligent video surveillance scene image dataset, obtain the feature matrix and label matrix.

[0073] In this embodiment, the intelligent video surveillance scene image dataset includes a training set, a validation set, and a test set. The samples in the training set include intelligent video surveillance scene images and their corresponding original labels. The validation set and the test set only include intelligent video surveillance scene images and do not contain corresponding label data.

[0074] In this embodiment, the ratio of the number of images in the training set, validation set, and test set is 1:2:7.

[0075] The trained LC-KSVD network is used to extract features from each image in the dataset to obtain the corresponding feature vector. The feature vectors of all images are stacked together to obtain a feature matrix X. The original labels of each image in the training set are converted into label vectors with a dimension equal to the total number of label categories and containing only 0 and 1 through one-hot encoding. All label vectors are stacked together to obtain a label matrix.

[0076] In this embodiment, each scene image is given a 3000-dimensional feature vector.

[0077] In this embodiment, the intelligent video surveillance scene image dataset contains label values ​​for 15 categories. Therefore, the one-hot vector label value for each data point is a 15-dimensional vector, where only the position corresponding to the category is 1, and the rest are 0. For example, if the label of a data sample is the 10th category, then its one-hot vector will only have the element value of the 10th position as 1.

[0078] S102. Input the feature matrix into the intelligent video surveillance scene image classification model and output the classification prediction results of the scene image.

[0079] like Figure 2 As shown, the intelligent video surveillance scene image classification model includes a hypergraph structure learning module and a hypergraph representation learning module.

[0080] Before inputting the feature matrix into the hypergraph structure learning module, the module also includes constructing the original hypergraph association matrix based on the feature matrix.

[0081] In order to preserve the high-order semantic information contained in the original feature matrix, the kNN algorithm is specifically used to calculate the similarity of the feature vectors of the initial feature matrix X, and the original hypergraph association matrix H is constructed using this similarity.

[0082] Specifically, the kNN algorithm is used to calculate the k nearest neighbors of each eigenvector in the feature matrix. The current eigenvector and its k neighbors together form a hyperedge, and finally the initial hypergraph correlation matrix H is constructed.

[0083] In this embodiment, the k value is set to 15 when performing the topk operation in the kNN algorithm.

[0084] Further, step S102 includes:

[0085] (1) Input the feature matrix into the hypergraph structure learning module to construct the implicit hypergraph correlation matrix; merge the original hypergraph correlation matrix and the implicit hypergraph correlation matrix to obtain the fused hypergraph correlation matrix.

[0086] This module adopts the idea of ​​multi-view learning and introduces a learnable similarity metric function on this basis, with the aim of obtaining the optimal hypergraph structure representation of intelligent video surveillance scene images.

[0087] like Figure 3 As shown, the input to this module is the feature matrix X, and the output is the learned implicit hypergraph correlation matrix.

[0088] Further, step (1) includes:

[0089] (1-1) Perform hypergraph learning on the feature matrix from multiple views to obtain the similarity matrix of each view.

[0090] This module performs hypergraph learning across multiple views. In the hypergraph learning for each view, firstly, a parameter matrix M is used to map scene image features from the original feature space to a low-dimensional embedding space; then, a multi-head learnable cosine similarity function is used to learn the similarity between sample data.

[0091] Specifically, in order to reduce noise in the original data features and reduce computational overhead, the original image features are mapped to a lower-dimensional space using a parameter matrix M, as shown in the following formula:

[0092]

[0093] in, It is a representation of scene image data embedded in a low-dimensional space.

[0094] In this embodiment, the first dimension of matrix M is the same as the second dimension of feature matrix X, while its second dimension is 80.

[0095] Specifically, a learnable similarity metric function is used to learn the similarity between scene image feature data, as shown in the following formula:

[0096]

[0097] sti,j∈[0,N-1]

[0098] Where w represents the learnable weight vector, and ⊙ represents the dot product operation of the vectors. and represents the embedding vectors of eigenvectors i and j in the low-dimensional space, respectively, and N is the total number of eigenvectors in the feature matrix.

[0099] The function as a whole still uses the cosine similarity calculation method, but inside the function, a learnable weight vector is used to perform a dot product operation with each feature vector.

[0100] To enhance the learning ability of the learnable similarity metric function and stabilize the model's training process to make it more likely to converge, the above formula is improved into a multi-head learnable cosine similarity calculation function, as follows:

[0101]

[0102] Among them, w h This represents the learnable weight vector under the h-th head. Then, it represents the similarity between eigenvectors i and j in the feature matrix of the h-th head, m represents the total number of heads in a multi-head configuration, and S ij The final similarity is calculated by summing and averaging the similarities of each head.

[0103] The improved formula shows that the function uses m learnable weight vectors. Analogous to the attention mechanism in Transformers, this method can be viewed as using m learnable heads to learn the similarity between any two feature vectors in the feature matrix from m semantic spaces respectively. Finally, the similarities obtained from different perspectives are fused by summing and averaging to obtain the final similarity value S. ij A similarity matrix S is obtained from all the similarity values.

[0104] In this embodiment, the number of heads m is set to 16.

[0105] This embodiment designs a multi-view hypergraph structure learning network based on a multi-head learnable cosine similarity calculation function to calculate the similarity between samples on multiple views, thereby mining higher-order relationships between data and significantly improving the model's classification performance.

[0106] (1-2) Sparsify the similarity matrix and use the sparsified similarity matrix as a hypergraph association matrix learned by a view.

[0107] The similarity matrix is ​​sparsified by filtering the values ​​in the matrix using a pre-defined threshold δ, resulting in a sparsified similarity matrix.

[0108]

[0109] In this embodiment, the threshold δ is set to 0.7.

[0110] Sparsification can reduce subsequent computational overhead.

[0111] The hypergraph association matrix is ​​obtained by learning the sparsified similarity matrix as a view.

[0112] (1-3) Fuse the hypergraph incidence matrices learned from all views to obtain the implicit hypergraph incidence matrix.

[0113] By fusing the hypergraph correlation matrices learned from multiple views, the hypergraph correlation matrix output by the structure learning module is obtained.

[0114]

[0115] Where P represents the number of views used, q represents the q-th view, and H... (q) Let be the hypergraph incidence matrix learned under the q-th view.

[0116] The hypergraph association matrix output by the hypergraph structure learning module is the learned implicit hypergraph structure.

[0117] To ensure the consistency of the hypergraph association matrix learned under different views, a consistency loss function is specifically introduced:

[0118]

[0119] Where, ||·|2 is the L2 norm.

[0120] In this embodiment, P is set to 2.

[0121] (1-4) The original hypergraph incidence matrix and the implicit hypergraph incidence matrix are merged to obtain the merged hypergraph incidence matrix.

[0122] The original hypergraph incidence matrix H and the hypergraph incidence matrix output by the hypergraph structure learning module are compared. The merging process, using the weight parameter η, yields the fused hypergraph correlation matrix.

[0123]

[0124] In this embodiment, the weight parameter η is set to 0.4.

[0125] (1-5) Based on graph regularization and the fused hypergraph association matrix, construct the loss function of the hypergraph structure learning module.

[0126] The loss function of the hypergraph structure learning module utilizes graph regularization techniques commonly used in the field of ordinary graph structure learning: sparsity and connectivity. Sparsity is used to penalize overly dense hypergraph incidence matrices. However, sparsity alone might result in a hypergraph incidence matrix consisting entirely of zeros. Therefore, connectivity is introduced as an additional constraint, which indirectly requires the incidence matrix to have as many edges as possible through the degree of the nodes. The loss function of the hypergraph structure learning module consists of the two regularization terms mentioned above, plus a consistency loss.

[0127]

[0128] Where A is the adjacency matrix of the fused hypergraph, representing the connection relationship between all nodes in the fused hypergraph, α, β, and γ are the weight coefficients controlling the three loss terms, and 1 represents an all-1 vector. T For the transpose operation, |·| F It is the Frobenius norm of the matrix, and log(.) represents the natural logarithm operation.

[0129] In this embodiment, during the model training process, the loss function of the hypergraph structure learning module is used to continuously optimize the hypergraph association matrix, making it more suitable for downstream scene image classification tasks. At the same time, based on multi-view learning, a learnable multi-head cosine similarity calculation function is combined to perform a more comprehensive and accurate calculation of the similarity between samples, so that the learned hypergraph association matrix has higher-order semantic information that is more suitable for downstream tasks.

[0130] (2) Input the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module, and output the node encoding representation of the hypergraph; based on the node encoding representation, obtain the classification prediction result of the scene image.

[0131] Before inputting the feature matrix and the fused hypergraph association matrix into this module, a message propagation on the hypergraph is performed to obtain the node feature matrix V and the hyperedge feature matrix E of the hypergraph. The specific formula is as follows:

[0132]

[0133] Among them, D e and D v Let represent the hyperedge degree matrix and node degree matrix of the hypergraph, respectively. The node degree is calculated as follows: The method for calculating hypermargin is as follows:

[0134] like Figure 4 As shown, the input of this module is the hyperedge feature matrix and the node feature matrix. It learns the scene image encoding representation that can be used for prediction and performs scene image classification based on this encoding representation.

[0135] Furthermore, step (2) includes:

[0136] (2-1) Input the node feature matrix and the hyperedge feature matrix into the hypergraph representation learning module to obtain the label prediction matrix of the scene image.

[0137] The overall structure of the hypergraph representation learning module follows the design of the representation learning part in the DualHGNN model, that is, it adopts a hypergraph density-aware attention layer. Each density-aware attention layer uses a density-aware attention mechanism to transfer and aggregate information on node features and hyperedge features.

[0138] Each density-aware attention layer consists of a node aggregation module and a hyperedge aggregation module. Since the hypergraph message passing method used in the hypergraph representation learning part is still node-to-hyperedge aggregation, and then hyperedge-to-node aggregation, the node aggregation module's role is to aggregate node information onto the hyperedge based on the density-aware attention mechanism. The hyperedge aggregation module's purpose is to aggregate hyperedge information onto the nodes, thus completing the full message passing process.

[0139] The hyperedge feature matrix E and the node feature matrix V are input into the density-aware attention layer, and hypergraph node feature aggregation is performed in the node aggregation module. The specific process is as follows:

[0140] First, the node density is calculated.

[0141] Node density is defined as the sum of the similarities of neighboring nodes whose similarity to the current node is greater than a preset threshold ε, as shown in the following formula:

[0142]

[0143] in, Represents node v i density, Represents all nodes v i For neighboring nodes within the same hyperedge, W is a learnable weight matrix. This represents the similarity calculation function.

[0144] In this embodiment, ε is set to 0.4.

[0145] Secondly, a shared attention mechanism is used to compute node v. i and super edge e k The attention weights between them are calculated using the following formula:

[0146]

[0147] Next, density-aware attention weights are obtained by fusing density information with attention weights:

[0148]

[0149] in, It is the standardized node density, a V Attention weight The set, Indicates the superedge e k The set of all connected nodes, θ V Let represent the learnable weight matrix. T For transpose, || represents tensor concatenation, exp(.) represents the exponentiation function, and LeakyReLU(.) is the activation function.

[0150] Finally, calculate the values ​​of all nodes. Then, the node attention weight matrix Z can be obtained. v The node features are aggregated using this matrix to obtain the updated hyperedge feature matrix, defined as follows:

[0151]

[0152] Here, ELU(.) is the exponential linear unit activation function.

[0153] The hyperedge feature matrix E and the node feature matrix V are input into the hyperedge aggregation module, where hypergraph hyperedge feature aggregation is performed. The specific process is as follows:

[0154] First, calculate the density of the hyperedge based on the node density.

[0155] The density of a hyperedge is defined as the sum of the densities of all nodes connected by the hyperedge, as shown in the formula:

[0156]

[0157] in, For the superedge e k The density.

[0158] Then, similar to the attention calculation method in the node aggregation module, the attention weight calculation formula in the hyperedge aggregation module is as follows:

[0159]

[0160] in, It is the standardized node density, a E Attention weight The set, Representative node v i The set of all superedges that it belongs to.

[0161] Finally, after calculating all hyperedges Then, the hyperedge attention weight matrix z can be obtained. E Then, the attention weight matrix is ​​used to update the hyperedge feature matrix. The updated node feature matrix is ​​obtained by performing hyperedge feature aggregation.

[0162] Furthermore, in this embodiment, multiple attention heads are used, and the updated node feature matrices output by each attention head are concatenated to obtain the final hypergraph node encoding representation. Specifically defined as:

[0163]

[0164] Among them, Z E,h and Z V,h This represents the super-edge attention weight matrix and the node attention weight matrix under the h-th attention head. This represents the tensor concatenation operation, where H represents the total number of attention heads.

[0165] The node encoding representation is then processed using softmax(.) to obtain the label prediction matrix O of the scene image.

[0166] In this embodiment, the number of layers in the hypergraph density-sensing attention layer, i.e. the number of heads in the multi-head attention, is set to 2.

[0167] (2-2) Construct the loss function for the hypergraph representation learning module.

[0168] The loss function is defined as the cross-class loss function for multi-class classification:

[0169]

[0170] Among them, Y ij Let L be the label value of the i-th labeled data sample in the j-th label category, L be the set of labeled samples, C be the total number of label categories, and O be the total number of labels. ij Let be the predicted value of the i-th data sample for the j-th label category.

[0171] This embodiment utilizes the density-aware hypergraph attention layer of the DualHGNN model to construct a hypergraph representation learning module, which can learn node encoding representations with richer semantic information.

[0172] Furthermore, such as Figure 5As shown, this embodiment also incorporates a hypergraph contrast learning module in the intelligent video surveillance scene image classification model. This module primarily constrains the consistency between the original hypergraph structure and the learned implicit hypergraph structure, while mitigating the oversmoothing problem that may arise from hypergraph representation learning. Compared to the structure learning module, which simply uses the L2 norm of the difference between the correlation matrices for constraint, the hypergraph contrast learning module further constrains the consistency of the two correlation matrices at the node feature representation level. The core idea is that the two encoded representations obtained after learning the hypergraph representation of the same node using two different correlation matrices should have the same or similar semantic information. Furthermore, a common problem in graph neural networks is node oversmoothing, where adjacent nodes tend to have the same encoded representation during training. In the hypergraph domain, this means that the encoded representations of neighboring nodes within the same hyperedge are nearly identical. However, even neighboring nodes should be distinguishable as much as possible. To address this, the hypergraph contrast learning module constructs a loss function to implement the above constraints, effectively solving the aforementioned problem. The specific process is as follows:

[0173] First, the feature matrix and the original hypergraph association matrix, as well as the feature matrix and the fused hypergraph association matrix, are input into the hypergraph representation learning module, respectively, and the node encoding representations learned under the two different association matrices are output respectively:

[0174]

[0175] Where F represents the hypergraph representation learning module, This represents the hypergraph node encoding representation learned by the hypergraph representation learning module from the feature matrix X and the original hypergraph association matrix H. Represents the feature matrix X and the fused hypergraph correlation matrix The hypergraph node encoding representation is learned through the hypergraph representation learning module.

[0176] Then, based on the node encoding representations learned under two different correlation matrices, the contrastive learning loss function is constructed as follows:

[0177]

[0178] in, Let φ represent all neighboring nodes that are on the same hyperedge as node i, φ be the similarity calculation function, and exp(.) be the exponential operation function. and The positive sample pairs in the contrastive learning loss represent respectively and The encoded representation of node i in the middle. yes The encoding representation of node j in the middle. That is The encoding representation of node j is shown in the figure.

[0179] The formula clearly shows the encoding representation of the same node under different association matrices. i and The closer the elements are, the smaller the loss function value. At the same time, the greater the difference between the encoded representations of adjacent nodes within the same hyperedge, the smaller the loss function value.

[0180] This embodiment designs a hypergraph contrast learning module, which can constrain the consistency between the original hypergraph structure and the learned implicit hypergraph structure in the feature space of the nodes. At the same time, by constraining the distance between different node representations within the same hyperedge to be as large as possible, it can further alleviate the oversmoothing problem that may be caused by hypergraph representation learning, and ultimately improve the classification performance of the model.

[0181] The overall loss function of the model is the weighted sum of the loss functions of the hypergraph structure learning module, the hypergraph representation learning module, and the hypergraph contrast learning module:

[0182]

[0183] Where λ and μ are the weighting coefficients when merging the loss function.

[0184] The hypergraph neural network designed in this embodiment performs end-to-end joint optimization of the hypergraph structure learning module, the hypergraph representation learning module, and the hypergraph comparison learning module, thereby simultaneously learning the hypergraph structure and node representation most suitable for image classification tasks in intelligent video surveillance scenarios, ultimately improving classification performance.

[0185] In this embodiment, the similarity calculation function is the cosine similarity calculation function.

[0186] S103. Train the intelligent video surveillance scene image classification model using the feature matrix and label matrix.

[0187] The model is trained using the feature matrix and label matrix obtained in step S101, the model parameters are optimized, and the parameter information of the model when it achieves the best classification effect in the intelligent video surveillance scene image classification task is obtained.

[0188] This embodiment is a semi-supervised node classification task in the hypergraph domain. Therefore, in the actual training process, the feature vectors of all images in the training set, validation set and test set are used for training, but only the scene images in the training set have label information as supervision signals for training.

[0189] In this embodiment, the total loss function of the intelligent video surveillance scene image classification model includes the loss function of the hypergraph structure learning module and the loss function of the hypergraph representation learning module.

[0190] Furthermore, the overall loss function of the intelligent video surveillance scene image classification model also includes the loss function of the hypergraph contrastive learning module, expressed as:

[0191]

[0192] Where λ and μ are the weighting coefficients when merging the loss function.

[0193] In this embodiment, the value of λ is set to 1.0, and the value of μ is set to 0.8.

[0194] The feature matrix obtained in step S101 is used as the input of the model. The model's total loss function is used to guide the model training, and the model finally converges on the dataset. Then, the model parameters at the point of achieving the best classification effect are saved.

[0195] S104. Input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the category prediction result corresponding to the intelligent video surveillance scene image.

[0196] The model is loaded using the model parameter file saved in step S103. Then, the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified is input into the model, and the category prediction result corresponding to the intelligent video surveillance scene images to be classified is output, thus completing the end-to-end prediction process.

[0197] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium.

[0198] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the order of execution of the described steps may be changed. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0199] Example 2:

[0200] like Figure 6As shown, this embodiment provides an intelligent video surveillance scene image classification system based on hypergraph structure learning. The system includes a sample acquisition module 601, a first construction module 602, a second construction module 603, a fusion module 604, a classification prediction module 605, a training module 606, and a category output module 607, wherein:

[0201] The sample acquisition module 601 is used to acquire a smart video surveillance scene image dataset; only some scene images in the smart video surveillance scene image dataset have corresponding label data; based on the scene images and label data in the smart video surveillance scene image dataset, a feature matrix and a label matrix are obtained.

[0202] The first construction module 602 is used to calculate the similarity between feature vectors in the feature matrix using the kNN algorithm, and to construct the original hypergraph association matrix based on the similarity.

[0203] The second construction module 603 is used to input the feature matrix into the hypergraph structure learning module in the intelligent video surveillance scene image classification model, and obtain the implicit hypergraph association matrix by learning the hypergraph from multiple views; wherein, in each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function.

[0204] The fusion module 604 is used to fuse the original hypergraph incidence matrix and the implicit hypergraph incidence matrix to obtain the fused hypergraph incidence matrix;

[0205] The classification prediction module 605 is used to input the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and output the node encoding representation of the hypergraph; based on the node encoding representation, the classification prediction result of the scene image is obtained.

[0206] Training module 606 is used to train the intelligent video surveillance scene image classification model using the feature matrix and label matrix, and to update the parameters of the intelligent video surveillance scene image classification model based on the total loss function to obtain the trained intelligent video surveillance scene image classification model. The total loss function includes the loss function of the hypergraph structure learning module. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, including a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views.

[0207] The category output module 607 is used to input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction results.

[0208] The specific implementation of each module in this embodiment can be found in Embodiment 1 above, and will not be repeated here. It should be noted that the system provided in this embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.

[0209] Example 3:

[0210] This embodiment provides a terminal device, which can be a computer, such as... Figure 7 As shown, the processor 702, memory, input device 703, display 704, and network interface 705 are connected via system bus 701. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium 706 and internal memory 707. The non-volatile storage medium 706 stores the operating system, computer programs, and database. The internal memory 707 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. When the processor 702 executes the computer programs stored in the memory, it implements the intelligent video surveillance scene image classification method based on hypergraph structure learning in Embodiment 1, as follows:

[0211] Obtain an intelligent video surveillance scene image dataset; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data;

[0212] Based on the scene images and label data in the intelligent video surveillance scene image dataset, the feature matrix and label matrix are obtained;

[0213] The kNN algorithm is used to calculate the similarity between eigenvectors in the feature matrix, and the original hypergraph association matrix is ​​constructed based on the similarity.

[0214] The feature matrix is ​​input into the hypergraph structure learning module in the intelligent video surveillance scene image classification model. The implicit hypergraph association matrix is ​​obtained by learning the hypergraph from multiple views. In each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function.

[0215] The original hypergraph incidence matrix and the implicit hypergraph incidence matrix are merged to obtain the merged hypergraph incidence matrix;

[0216] The feature matrix and the fused hypergraph association matrix are input into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and the node encoding representation of the hypergraph is output. Based on the node encoding representation, the classification prediction result of the scene image is obtained.

[0217] A smart video surveillance scene image classification model is trained using feature matrices and label matrices. The parameters of the smart video surveillance scene image classification model are updated based on the total loss function to obtain the trained smart video surveillance scene image classification model. The total loss function includes the loss function of the hypergraph structure learning module. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, including a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views.

[0218] Input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction result.

[0219] Example 4:

[0220] This embodiment provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the intelligent video surveillance scene image classification method based on hypergraph structure learning described in Embodiment 1 above, as follows:

[0221] Obtain an intelligent video surveillance scene image dataset; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data;

[0222] Based on the scene images and label data in the intelligent video surveillance scene image dataset, the feature matrix and label matrix are obtained;

[0223] The kNN algorithm is used to calculate the similarity between eigenvectors in the feature matrix, and the original hypergraph association matrix is ​​constructed based on the similarity.

[0224] The feature matrix is ​​input into the hypergraph structure learning module in the intelligent video surveillance scene image classification model. The implicit hypergraph association matrix is ​​obtained by learning the hypergraph from multiple views. In each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function.

[0225] The original hypergraph incidence matrix and the implicit hypergraph incidence matrix are merged to obtain the merged hypergraph incidence matrix;

[0226] The feature matrix and the fused hypergraph association matrix are input into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and the node encoding representation of the hypergraph is output. Based on the node encoding representation, the classification prediction result of the scene image is obtained.

[0227] A smart video surveillance scene image classification model is trained using feature matrices and label matrices. The parameters of the smart video surveillance scene image classification model are updated based on the total loss function to obtain the trained smart video surveillance scene image classification model. The total loss function includes the loss function of the hypergraph structure learning module. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, including a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views.

[0228] Input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction result.

[0229] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0230] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A method for intelligent video surveillance scene image classification based on hypergraph structure learning, characterized in that, The method includes: Obtain an intelligent video surveillance scene image dataset; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data; Based on the scene images and label data in the intelligent video surveillance scene image dataset, the feature matrix and label matrix are obtained; The kNN algorithm is used to calculate the similarity between eigenvectors in the feature matrix, and the original hypergraph association matrix is ​​constructed based on the similarity. The feature matrix is ​​input into the hypergraph structure learning module in the intelligent video surveillance scene image classification model. The implicit hypergraph association matrix is ​​obtained by learning the hypergraph from multiple views. In each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function. The original hypergraph incidence matrix and the implicit hypergraph incidence matrix are merged to obtain the merged hypergraph incidence matrix; The feature matrix and the fused hypergraph association matrix are input into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and the node encoding representation of the hypergraph is output. Based on the node encoding representation, the classification prediction result of the scene image is obtained. A smart video surveillance scene image classification model is trained using feature matrices and label matrices. The parameters of the model are then updated based on a total loss function to obtain the trained model. The total loss function includes the loss function of the hypergraph structure learning module and a contrastive learning loss function. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, and includes a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views. The contrastive learning loss function is obtained based on the node encoding representations of the original and fused hypergraph association matrices to improve the classification performance of the smart video surveillance scene image classification model. Input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction result. The learnable multi-head cosine similarity calculation function is as follows: in, w represents the similarity between eigenvectors i and j in the feature matrix of the h-th head. h This represents the learnable weight vector under the h-th head. and Let represent the embedding vectors of eigenvectors i and j in the low-dimensional space, respectively; ⊙ represents the dot product operation of the vectors; m represents the total number of heads in the multi-head sequence; S ij This represents the final similarity score after summing and averaging the similarities of each head.

2. The intelligent video surveillance scene image classification method according to claim 1, characterized in that, The process of obtaining the implicit hypergraph association matrix through hypergraph learning from multiple views includes: In each view, the feature vectors in the feature matrix are mapped to a low-dimensional space, and the similarity between feature vectors in the low-dimensional space is calculated using a learnable multi-head cosine similarity calculation function to obtain the final similarity matrix. The similarity matrix is ​​sparsified, and the sparsified similarity matrix is ​​used as a view to learn the hypergraph association matrix. The hypergraph incidence matrices learned from all views are fused to obtain the implicit hypergraph incidence matrix.

3. The intelligent video surveillance scene image classification method according to claim 1, characterized in that, Before inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module, message propagation on the hypergraph is performed first to obtain the node feature matrix and hyperedge feature matrix of the hypergraph.

4. The intelligent video surveillance scene image classification method according to claim 3, characterized in that, The hypergraph representation learning module adopts a hypergraph density-aware attention layer, which consists of a node aggregation module and a hyperedge aggregation module. The process of inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module of the intelligent video surveillance scene image classification model, and outputting the node encoding representation of the hypergraph, includes: The hyperedge feature matrix and node feature matrix are input into the node aggregation module. The node information is aggregated onto the hyperedge based on the density-aware attention mechanism to obtain the updated hyperedge feature matrix. The hyperedge feature matrix and node feature matrix are input into the hyperedge aggregation module. Hyperedge features are aggregated based on the updated hyperedge feature matrix to obtain the updated node feature matrix. The updated node feature matrix is ​​the output of a density-aware attention layer. By employing multiple attention heads, the outputs of each density-aware attention layer are concatenated to obtain the final node encoding representation.

5. The intelligent video surveillance scene image classification method according to claim 4, characterized in that, The process of inputting the hyperedge feature matrix and node feature matrix into the node aggregation module, and then using a density-aware attention mechanism to aggregate node information onto the hyperedge to obtain an updated hyperedge feature matrix, specifically includes: Calculate the node density; where the node density is the sum of the similarities of neighboring nodes whose similarity to the current node is greater than a preset threshold. A shared attention mechanism is used to calculate the attention weights between nodes and hyperedges; The node density information is fused with the attention weights to obtain the attention weights; based on the attention weights of all nodes, the attention weight matrix is ​​obtained. The node features are aggregated using the node attention weight matrix to obtain the updated hyperedge feature matrix.

6. The intelligent video surveillance scene image classification method according to claim 1, characterized in that, The node encoding representation of the original hypergraph association matrix is ​​obtained by inputting the feature matrix and the original hypergraph association matrix into the hypergraph representation learning module, and the node encoding representation of the fused hypergraph association matrix is ​​obtained by inputting the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module.

7. The intelligent video surveillance scene image classification method according to any one of claims 1 or 6, characterized in that, The total loss function also includes the loss function of the hypergraph representation learning module, expressed as: In the formula, Y ij Let L be the label value of the i-th labeled data sample in the label matrix, corresponding to the j-th label category. Let L be the set of labeled samples in the label matrix, C be the total number of label categories, and O be the total number of categories. ij Let be the predicted value of the i-th data sample for the j-th label category.

8. The intelligent video surveillance scene image classification method according to any one of claims 1 to 6, characterized in that, The step of obtaining the feature matrix and label matrix based on the scene images and label data in the intelligent video surveillance scene image dataset includes: Feature extraction is performed on each scene image in the intelligent video surveillance scene image dataset to obtain the corresponding feature vector; Stack the feature vectors corresponding to all scene images to obtain the feature matrix; Each label data is converted into a label vector with the dimension of the total number of label categories and containing only 0s and 1s through one-hot encoding; only the position corresponding to the category is 1, and the rest are 0s; Stack the label vectors corresponding to all the label data to obtain a label matrix.

9. An intelligent video surveillance scene image classification system based on hypergraph structure learning, characterized in that, The system includes: The sample acquisition module is used to acquire a dataset of intelligent video surveillance scene images; only some scene images in the intelligent video surveillance scene image dataset have corresponding label data; based on the scene images and label data in the intelligent video surveillance scene image dataset, a feature matrix and a label matrix are obtained. The first construction module is used to calculate the similarity between feature vectors in the feature matrix using the kNN algorithm, and to construct the original hypergraph association matrix based on the similarity. The second construction module is used to input the feature matrix into the hypergraph structure learning module in the intelligent video surveillance scene image classification model, and obtain the implicit hypergraph association matrix by learning the hypergraph from multiple views; wherein, in each view, the similarity between feature vectors in the feature matrix is ​​learned using a learnable multi-head cosine similarity calculation function. The fusion module is used to fuse the original hypergraph incidence matrix and the implicit hypergraph incidence matrix to obtain the fused hypergraph incidence matrix; The classification prediction module is used to input the feature matrix and the fused hypergraph association matrix into the hypergraph representation learning module in the intelligent video surveillance scene image classification model, and output the node encoding representation of the hypergraph; based on the node encoding representation, the classification prediction result of the scene image is obtained. The training module is used to train an intelligent video surveillance scene image classification model using feature matrices and label matrices. It updates the parameters of the intelligent video surveillance scene image classification model based on a total loss function to obtain a trained model. The total loss function includes the loss function of the hypergraph structure learning module and a contrastive learning loss function. The loss function of the hypergraph structure learning module is constructed using graph regularization techniques based on the fused hypergraph association matrix, including a consistency loss function. The consistency loss function is used to constrain the consistency of the hypergraph association matrix learned under different views. The contrastive learning loss function is obtained based on the node encoding representations of the original and fused hypergraph association matrices to improve the classification performance of the intelligent video surveillance scene image classification model. The category output module is used to input the feature matrix corresponding to the set of intelligent video surveillance scene images to be classified into the trained intelligent video surveillance scene image classification model, and output the corresponding category prediction results. The learnable multi-head cosine similarity calculation function is as follows: in, w represents the similarity between eigenvectors i and j in the feature matrix of the h-th head. h This represents the learnable weight vector under the h-th head. and Let represent the embedding vectors of eigenvectors i and j in the low-dimensional space, respectively; ⊙ represents the dot product operation of the vectors; m represents the total number of heads in the multi-head sequence; S ij This represents the final similarity score after summing and averaging the similarities of each head.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent video surveillance scene image classification method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-omics association phenotype prediction method based on hypergraph representation and Dirichlet distribution

    CN114927162A

  • Continuous task data classification method and device based on dynamic hypergraph continuous learning

    CN117556292A