Campus Security Control Method and System Integrating Multiple Systems
By arranging high-definition cameras and sensor arrays in the campus area, and combining the multi-system data fusion method to detect and locate abnormal events, the problem of limited perception and analysis capabilities of abnormal events in the existing technology is solved, and more efficient campus safety management is achieved.
Patent Information
- Application Number
- CN202510322827.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing campus safety control methods mainly rely on a single video surveillance data, making it difficult to make full use of the data sources of multiple campus systems, resulting in limited perception and analysis capabilities of abnormal events.
The multi-system data fusion method is adopted, including laying high-definition cameras and sensor arrays in the campus area, adopting a human posture estimation model based on attention mechanism, building a campus security knowledge graph and training a graph neural network model, and combining video streams and audio signals for abnormal event detection and positioning.
It realizes campus personnel safety monitoring based on multi-system data, improves the detection accuracy and positioning accuracy of abnormal events, and enhances the foresight and accuracy of campus safety management.
Smart Images

Figure CN119850386B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a campus security control method and system integrating multiple systems. Background Art
[0002] With the continuous improvement of the intelligence level of campuses, the application of video surveillance, audio monitoring, and sensor technologies in campus security has received extensive attention. As a place with dense personnel flow and frequent activities, it is particularly important to ensure the safety of students and teaching staff. In this context, the real-time detection and accurate positioning of abnormal events have become the key to ensuring campus security.
[0003] The Chinese patent application with the publication number CN118153953A discloses a campus security event early warning method based on spatio-temporal analysis and video AI. The method includes constructing a spatial data model, constructing a road network, deploying cameras, and generating an early warning model. By combining the specific school schedule, three-dimensional spatial structure characteristics, and real-time video intelligent analysis results of the campus, a three-dimensional and intelligent early warning system is constructed, thereby effectively improving the foresight and accuracy of campus security management.
[0004] The existing campus security control methods mainly rely on single video surveillance data and are difficult to make full use of the data sources of multiple campus systems, resulting in limited perception and analysis capabilities for abnormal events. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems in the related technologies to some extent. For this reason, an object of this application is to propose a campus security control method and system integrating multiple systems, which realizes the security monitoring of campus personnel based on multi-system data.
[0006] One aspect of this application provides a campus security control method integrating multiple systems, including:
[0007] Step S100: Deploy high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals. When an abnormal audio signal is captured, obtain the video streams before and after the abnormal occurrence time;
[0008] Step S200: Use a human pose estimation model based on the attention mechanism to obtain the human pose information in the video stream, calculate the abnormal behavior data, and design an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data;
[0009] Step S300: Construct a campus security knowledge graph based on multi-system data of the campus, construct a graph neural network model based on the campus security knowledge graph. When an abnormal event occurs, use the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event;
[0010] Step S400: Obtain a sensor combination based on the abnormal audio signal, calculate the signal similarity between each pair of sensors in the sensor combination, calculate the candidate sound source position with the highest matching degree score according to the signal similarity, and use it as the occurrence position of the abnormal event;
[0011] Step S500: Send the occurrence position of the abnormal event and the reasoning reason of the abnormal event to the campus security personnel;
[0012] The specific method for deploying high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals and obtaining the video streams before and after the abnormal occurrence moment when an abnormal audio signal is captured is as follows: Deploy high-definition cameras and sensor arrays in the campus area, and deploy the sensor array near the camera using a uniformly distributed topological structure. The sensor spacing of the sensor array satisfies: , where c is the speed of sound, is the maximum frequency of the audio signal; when the sensor array captures an abnormal audio signal at the abnormal occurrence moment , obtain the video stream with a continuous duration before and after the abnormal occurrence moment ;
[0013] The specific acquisition method of the human pose estimation model based on the attention mechanism is as follows:
[0014] Step S210: Use the human pose estimation model based on the attention mechanism to perform human pose estimation, extract the spatial coordinates of human key points, and introduce the attention mechanism in the time dimension to weight and aggregate the human pose features of different frames. The attention weight of the t-th frame is ;
[0015] Step S220: Introduce the attention mechanism in the spatial dimension, extract the feature vectors of different parts of the human body, and the attention weight of part j1 is ;
[0016] Step S230: The human pose estimation model based on the attention mechanism uses the Attention-LSTM network as the initial network, including a CNN backbone network, an attention mechanism, and an LSTM time series module. Use the continuous-frame human body images obtained from the video stream as input data, and use the human key points of each frame of human body image as output data. Obtain the human pose information estimated by the Attention-LSTM network according to the spatial position of the human key points, the relative position and distance between the human key points, and the movement trajectory of the human key points. The movement trajectory of the human key points is obtained by connecting the coordinates of the same human key point in the human body images of different frames;
[0017] The training method of the human body pose estimation model is as follows:
[0018] Using the Attention-LSTM network as the initial network, design the Attention-LSTM network. The end of the Attention-LSTM network is connected to a fully connected layer, which maps the hidden state output by the LSTM time series module to the two-dimensional coordinates of the human body key points. The output of the entire model is the coordinate sequence of the human body key points corresponding to each frame of the human body image;
[0019] Collect a human body pose dataset. The dataset includes human body images and the corresponding coordinate annotations of human body key points. Use the human body images as input data and the corresponding coordinate annotations of human body key points as output data to form training samples. Use the training samples to train the Attention-LSTM network. Adopt the mean square error of the human body key point coordinates as the loss function for training the human body pose estimation model, and use minimizing the value of the loss function as the training objective. When the loss function converges, the training is completed, and a trained human body pose estimation model is obtained;
[0020] The specific method for obtaining the human body pose information in the video stream, calculating the abnormal behavior data, and designing an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data is as follows:
[0021] Step S240: The abnormal behavior data includes the static duration of the human body, the height difference of the center of gravity position, the center of gravity descent speed and acceleration, the change in the human body falling angle and angular velocity, the average key point density, and the average movement amplitude of all human body key points; the abnormal events include abnormal fainting events, abnormal aggregation events, and abnormal limb movements. The abnormal event judgment function includes an abnormal fainting event judgment function, an abnormal aggregation event judgment function, and an abnormal limb movement judgment function;
[0022] Step S250: Set a static threshold , calculate the inter-frame displacement of the human body key points. For the k-th human body key point, calculate its inter-frame displacement between the t-th frame and the (t + 1)-th frame , when the inter-frame displacement of the human body key point is less than the static threshold, it is considered that the human body key point is in a static state. If all human body key points remain static within consecutive frames, the static duration is , and the static duration exceeds the preset time threshold , then it is considered that the person enters a static state, where is the frame interval time;
[0023] Step S260: Calculate the coordinates of the human body center of gravity position according to the coordinates of the human body key points , calculate the height difference of the human body's center of gravity position between the t-th frame and the (t-1)-th frame , calculate the center of gravity descent speed and acceleration , calculate the change in the human body's falling angle and angular velocity ;
[0024] Step S270: Divide the video stream screen into grids, calculate the density of human body key points in each grid , obtain the average key point density of the entire screen , set the density threshold , define the abnormal aggregation event judgment function as: when the average key point density continuously exceeds the density threshold within the time threshold , it is determined that an abnormal aggregation event has occurred;
[0025] Step S280: Calculate the average movement amplitude of the k-th human body key point within consecutive frames , calculate the average movement amplitude of all human body key points , set the amplitude threshold , define the abnormal limb movement judgment function as: when the average movement amplitude of all human body key points exceeds the amplitude threshold, it is determined that an abnormal limb movement has occurred;
[0026] Step S290: Set the center of gravity height change threshold, center of gravity descent speed threshold, acceleration threshold, falling angle change threshold, and angular velocity threshold. Set condition one as: the height difference of the center of gravity continuously is less than the negative center of gravity height change threshold within consecutive frames; condition two as: the center of gravity descent speed continuously exceeds the center of gravity descent speed threshold and the acceleration continuously exceeds the acceleration threshold within consecutive frames; condition three as: the change in the falling angle continuously exceeds the falling angle change threshold and the angular velocity continuously exceeds the angular velocity threshold within consecutive frames; condition four as: all human body key points remain stationary within consecutive frames, and the stationary duration exceeds the preset time threshold ; design the abnormal fainting event judgment function to include the above conditions one, two, three, and four. When any three of the four conditions are met, it is determined that an abnormal fainting event has occurred;
[0027] The specific method for constructing the campus security knowledge graph based on campus multi-system data and constructing a graph neural network model based on the campus security knowledge graph is as follows:
[0028] Step S310: Define the top-level concepts of the campus security knowledge graph to form an ontology architecture;
[0029] Step S320: Extract different entities, entity attributes, and the relationships between entities from the multi-system database of the campus according to the top-level concepts. Use the entities as the nodes of the graph and the relationships as the edges of the graph to construct a campus security knowledge graph;
[0030] Step S330: Construct a graph neural network model based on the campus security knowledge graph. The graph neural network model includes a graph convolutional layer and a graph attention layer. The graph convolutional layer is used to perform a weighted sum of the embedded representations of the neighbors of the node, and the weights of the weighted sum are determined by the edge weights between the nodes ;
[0031] Step S340: Introduce an attention mechanism on the basis of the graph convolutional layer to obtain a graph attention layer, calculate the attention weight of different neighbor nodes j to node i , and obtain the embedded representation of node i based on the attention weight ;
[0032] Step S350: Train the graph neural network model in an unsupervised manner. Based on the relationships between nodes in the campus security knowledge graph, form triples (head node, relationship, tail node). The triples are used as positive samples, and negative samples are constructed through negative sampling. The positive samples and negative samples are used as training data, and an unsupervised training method is adopted, with the nodes in the campus security knowledge graph as the input and the embedded representations of the nodes as the output;
[0033] Step S360: Use maximizing the likelihood probability of positive samples and minimizing the likelihood probability of negative samples as the training objective. Based on the likelihood probabilities of positive samples and negative samples, obtain the loss function for training the graph neural network model , train the graph neural network model, minimize the loss function, and when the loss function converges, obtain the trained graph neural network model;
[0034] The specific method for using the graph neural network model to learn the embedded representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event when an abnormal event occurs is as follows:
[0035] Step S370: When an abnormal event occurs, extract the abnormal event entity of the abnormal event and input it into the graph neural network model to output the embedded representation of the abnormal event entity;
[0036] Step S380: In the campus security knowledge graph, with the node where the abnormal event entity is located as the center, search for other entities with the smallest inner product with the abnormal event entity in the embedded space of the campus security knowledge graph. Select the top K* entities with the smallest inner product as candidate reasons, and calculate their semantic correlation probability based on the embedded vectors of the abnormal event entity and the candidate reasons , integrating the semantic correlation probability and the corresponding candidate reasons into the reasoning reasons for abnormal events;
[0037] The specific method for obtaining the sensor combination based on the abnormal audio signal, calculating the signal similarity between each pair of sensors in the sensor combination, and calculating the candidate sound source position with the highest matching degree score according to the signal similarity, and taking it as the occurrence position of the abnormal event is as follows:
[0038] Step S410: Obtain the short-time energy of each sensor when the abnormal event occurs and the signal-to-noise ratio , and select the sensors with short-time energy greater than the energy threshold and signal-to-noise ratio greater than the signal-to-noise ratio threshold as the sensor combination;
[0039] Step S420: For each pair of sensors in the sensor combination , calculate the signal similarity between the sensor pairs according to their generalized cross-correlation function at the moment when the abnormal event occurs ;
[0040] Step S430: Calculate the matching degree score of the candidate sound source position P by accumulating the signal similarity of each pair of sensor pairs ;
[0041] Step S440: Take the candidate sound source position P with the maximum matching degree score as the occurrence position of the abnormal event.
[0042] One aspect of the present application provides a campus security control system integrating multiple systems, including:
[0043] A video and audio acquisition module, used to deploy high-definition cameras and sensor arrays in the campus area, collect video streams and audio signals, and obtain the video streams before and after the abnormal event occurs when an abnormal audio signal is captured;
[0044] An abnormal event judgment module, used to adopt a human pose estimation model based on the attention mechanism to obtain the human pose information in the video stream, calculate the abnormal behavior data, and design an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data;
[0045] An abnormal reason reasoning module, used to construct a campus security knowledge graph based on campus multi-system data, construct a graph neural network model based on the campus security knowledge graph, and when an abnormal event occurs, use the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyze the reasoning reasons for the abnormal event;
[0046] An occurrence location calculation module, configured to obtain a sensor combination based on an abnormal audio signal, calculate the signal similarity between each pair of sensors in the sensor combination, calculate the candidate sound source location with the highest matching degree score according to the signal similarity, and use it as the occurrence location of the abnormal event;
[0047] A security information receiving module, configured to send the occurrence location of the abnormal event and the reasoning reason of the abnormal event to the campus security personnel.
[0048] One aspect of the present application provides a readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the campus security control method for integrating multiple systems.
[0049] The campus security control method and system for integrating multiple systems proposed in the present application have the following advantages compared with the prior art:
[0050] The present application adopts a multi-modal perception method of a camera and a sensor array, which can capture more comprehensive abnormal event information. Especially, by analyzing the audio signal to trigger the detection of the abnormal event location, it can make up for the deficiency of single video analysis and improve the sensitivity of early warning.
[0051] The traditional video analysis method has insufficient accuracy in estimating human postures and is difficult to accurately judge abnormal changes in human postures. However, the present application adopts a human posture estimation model enhanced by an attention mechanism, which can more finely depict the motion characteristics of different parts of the human body and design a more scientific abnormal event determination criterion through an abnormal event judgment function.
[0052] The present application realizes the semantic fusion of campus multi-system data through knowledge graph technology, and trains a graph neural network model on this basis for event reason reasoning, which is a brand-new paradigm for the integrated utilization of campus data.
[0053] The present application innovatively proposes a sound source localization method based on sensor signal similarity, which can solve the problem that video analysis is easily interfered by factors such as occlusion and illumination, and is especially suitable for the situation of limited vision in the campus environment.
[0054] The present application integrates campus multi-system data, real-time monitors abnormal events, and locates abnormal events, forming a complete solution from data perception, feature extraction, event discrimination to positioning and alarming, providing new ideas and methods for improving the campus security management level. Description of the Drawings
[0055] Figure 1 It is the method flow chart of the campus security control method for integrating multiple systems provided by the present application;
[0056] Figure 2Flowchart of the method for analyzing the reasoning reasons of abnormal events provided by this application;
[0057] Figure 3 Flowchart of the method for estimating the occurrence location of abnormal events provided by this application;
[0058] Figure 4 Functional module diagram of the campus security control system integrating multiple systems provided by this application. Detailed implementation manners
[0059] To better understand this application, more detailed descriptions of various aspects of this application will be made with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of the exemplary embodiments of this application, and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.
[0060] In the accompanying drawings, for the convenience of illustration, the sizes, dimensions and shapes of the elements have been slightly adjusted. The drawings are only examples and are not drawn strictly to scale. As used herein, terms such as "substantially", "about" and similar terms are used as terms indicating approximation, rather than terms indicating degree, and are intended to illustrate the inherent deviations in measured or calculated values that would be recognized by those of ordinary skill in the art. Additionally, in this application, the order of description of each step process does not necessarily represent the order in which these processes occur in actual operation, unless there are clear other limitations or can be deduced from the context.
[0061] It should also be understood that expressions such as "including", "including having", "having", "containing" and / or "containing having" in this specification are open-ended rather than closed-ended expressions, which mean that there are the stated features, elements and / or components, but do not exclude the existence of one or more other features, elements, components and / or their combinations. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features, rather than just an individual element in the list. In addition, when describing the embodiments of this application, the use of "may" means "one or more embodiments of this application". And the term "exemplary" is intended to refer to an example or illustration.
[0062] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as the ordinary understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that unless there is a clear explanation in this application, words defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the related art, and should not be interpreted in an idealized or overly formal sense.
[0063] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0064] Embodiment 1
[0065] As Figure 1 shown, the campus security control method integrating multiple systems provided by the present application includes:
[0066] Step S100: Deploy high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals. When an abnormal audio signal is captured, obtain the video streams before and after the moment of the abnormality.
[0067] The specific method for deploying high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals and obtaining the video streams before and after the moment of the abnormality when an abnormal audio signal is captured is as follows: Deploy high-definition cameras and sensor arrays inside the campus area, and arrange the sensor arrays in a uniformly distributed topological structure near the cameras. The sensor spacing of the sensor arrays satisfies: , where c is the speed of sound, is the maximum frequency of the audio signal; when the sensor array captures an abnormal audio signal at the moment of the abnormality , obtain the video streams with a duration before and after the moment of the abnormality ;
[0068] The value is set by those skilled in the art according to experience, which refers to the duration of the abnormal event and is usually a short time window;
[0069] Step S200: Adopt a human pose estimation model based on the attention mechanism to obtain the human pose information in the video stream, calculate the abnormal behavior data, and design an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data.
[0070] The specific acquisition method of the human pose estimation model based on the attention mechanism is as follows:
[0071] Step S210: Adopt a human pose estimation model based on the attention mechanism for human pose estimation, extract the spatial coordinates of human key points, and introduce the attention mechanism in the time dimension to weight and aggregate the human pose features of different frames. The attention weight of the t-th frame is ;
[0072] The calculation formula for the attention weight of the t-th frame is: , , where is the attention score for the t-th frame, , are the weight matrix and bias term of the attention mechanism in the time dimension respectively, is the hidden state of the t-th frame, and T1 represents the number of frames in the input sequence, represents the attention score of the i1-th frame;
[0073] Step S220: In the spatial dimension, introduce the attention mechanism to extract the feature vectors of different parts of the human body. The attention weight of part j1 is ;
[0074] The calculation formula of the attention weight of the part j1 is: , , where is the attention score of part j1, , are the weight matrix and bias term of the attention mechanism in the spatial dimension respectively, represents the feature vector of part j1, represents the attention score of part k1, and J1 represents the total number of all parts;
[0075] Step S230: The human pose estimation model based on the attention mechanism uses the Attention-LSTM network as the initial network, including a CNN backbone network, an attention mechanism, and an LSTM time series module. The continuous-frame human body images obtained from the video stream are used as input data, and the human body key points of each frame of the human body image are used as output data. The human pose information estimated by the Attention-LSTM network is obtained according to the spatial position of the human body key points, the relative position and distance between the human body key points, and the movement trajectory of the human body key points. The movement trajectory of the human body key points is obtained by connecting the coordinates of the same human body key point in the human body images of different frames;
[0076] The training method of the human pose estimation model is:
[0077] Using the Attention-LSTM network as the initial network, design the Attention-LSTM network. The end of the Attention-LSTM network is connected to a fully connected layer to map the hidden state output by the LSTM time series module to the two-dimensional coordinates of the human body key points. The output of the entire model is the coordinate sequence of the human body key points corresponding to each frame of the human body image;
[0078] Collect a human body pose dataset. The dataset includes human body images and corresponding annotations of the coordinates of human body key points. Use the human body images as input data and the corresponding annotations of the coordinates of human body key points as output data to form training samples. Use the training samples to train the Attention-LSTM network. Adopt the mean square error of the coordinates of human body key points as the loss function for training the human body pose estimation model, and use minimizing the value of the loss function as the training objective. When the loss function converges, the training is completed, and a trained human body pose estimation model is obtained;
[0079] Among them, the CNN backbone network is used to extract the high-level feature representation of the human body image. Use a pre-trained deep convolutional neural network, remove the last fully connected layer, retain the convolutional layer and the pooling layer, take a single-frame human body image as input, and output a feature map;
[0080] The attention mechanism is used to adaptively focus on the key parts of the human body pose. Calculate the attention weights according to the content of the feature map, and perform weighted combination on the feature map. Use the feature map output by the CNN backbone network as input and the weighted combined attention feature map as output, and its size is the same as that of the feature map;
[0081] The LSTM time series module is used to model the evolution and dependence relationship of the human body pose in the time dimension. Use a long short-term memory network, take the sequence of attention feature maps output by the attention mechanism as input, and take the hidden state at each time step as output;
[0082] The specific method for obtaining the human body pose information in the video stream, calculating the abnormal behavior data, and designing an abnormal event judgment function based on the abnormal behavior data to judge whether an abnormal event occurs is as follows:
[0083] Step S240: The abnormal behavior data includes the static duration of the human body, the height difference of the center of gravity position, the center of gravity descent speed and acceleration, the change in the human body falling angle and angular velocity, the average key point density, and the average movement amplitude of all human body key points; the abnormal events include abnormal fainting events, abnormal aggregation events, and abnormal limb movements, and the abnormal event judgment function includes an abnormal fainting event judgment function, an abnormal aggregation event judgment function, and an abnormal limb movement judgment function;
[0084] Step S250: Set a static threshold , calculate the inter-frame displacement of the human body key points. For the k-th human body key point, calculate its inter-frame displacement between the t-th frame and the (t + 1)-th frame , when the inter-frame displacement of the human body key point is less than the static threshold, it is considered that the human body key point is in a static state. If all human body key points remain static within frames, the static duration is , and the static duration exceeds the preset time threshold If so, it is considered that the person enters a stationary state, where is the frame interval time;
[0085] The calculation formula for the displacement between frames is: where and respectively represent the abscissa and ordinate of the k-th human body key point in the (t + 1)-th frame, and respectively represent the abscissa and ordinate of the k-th human body key point in the t-th frame;
[0086] The and the stationary threshold are set by those skilled in the art according to experience;
[0087] Step S260: Calculate the coordinates of the human body center of gravity according to the coordinates of the human body key points , calculate the height difference of the human body center of gravity position between the t-th frame and the (t - 1)-th frame , calculate the center of gravity descent speed and the acceleration , calculate the change in the human body falling angle and the angular velocity ;
[0088] The calculation formula for the coordinates of the human body center of gravity position is: , where K is the total number of human body key points;
[0089] The calculation formula for the height difference of the human body center of gravity position between the t-th frame and the (t - 1)-th frame is: , where and are respectively the ordinates of the human body center of gravity position in the t-th frame and the (t + 1)-th frame;
[0090] The calculation formulas for the center of gravity descent speed and the acceleration are respectively: , , where is the center of gravity descent speed at the (t + 1)-th frame;
[0091] The specific method for calculating the change in the human body falling angle and the angular velocity is: obtain the direction vector of the upper body of the human body according to the coordinates of the human body key points, calculate the direction vector of the upper body, and calculate the change in the human body falling angle and the angular velocity of the upper body between the t-th frame and the -th frame;
[0092] The calculation formula for the change in the human body falling angle is: , where is the direction vector of the upper body in the (t + 1)-th frame, is the modulus length;
[0093] The calculation formula for the angular velocity of the upper body is: ;
[0094] Furthermore, the direction vector of the human upper body selects the vector connecting the midpoint of the shoulders and the midpoint of the hips;
[0095] Step S270: Divide the picture of the video stream into grids, calculate the density of human key points in each grid , obtain the average key point density of the entire picture , set the density threshold , define the abnormal aggregation event judgment function as: when the average key point density continuously exceeds the density threshold within the time threshold , it is determined that an abnormal aggregation event has occurred;
[0096] The calculation formula for the density of human key points in each grid is: , where is the number of human key points in the grid , is the area of the grid ;
[0097] The calculation formula for the average key point density of the entire picture is: ;
[0098] Step S280: Calculate the average movement amplitude of the k-th human key point within consecutive frames, calculate the average movement amplitude of all human key points, set the amplitude threshold , define the abnormal limb movement judgment function as: when the average movement amplitude of all human key points exceeds the amplitude threshold, it is determined that an abnormal limb movement has occurred;
[0099] The calculation formula for the average movement amplitude of the k-th human key point within consecutive frames is: , where , respectively represent the human key point k in the t-th frame and the (t + 1)-th frame, represents the Euclidean distance between the two points;
[0100] The average movement amplitude of all human key points is calculated as: where K represents the total number of all human key points;
[0101] Step S290: Set the threshold of the change in the center of gravity height, the threshold of the descending speed of the center of gravity, the acceleration threshold, the threshold of the change in the falling angle, and the angular velocity threshold. Set Condition 1 as follows: within consecutive frames, the difference in the center of gravity height continuously remains less than the negative threshold of the change in the center of gravity height. Condition 2 is: within consecutive frames, the descending speed of the center of gravity continuously remains greater than the threshold of the descending speed of the center of gravity and the acceleration continuously remains greater than the acceleration threshold. Condition 3 is: within consecutive frames, the change in the falling angle continuously remains greater than the threshold of the change in the falling angle and the angular velocity continuously remains greater than the angular velocity threshold. Condition 4 is: all human key points remain stationary within consecutive frames, and the stationary duration exceeds the preset time threshold ; Design an abnormal fainting event judgment function including the above Conditions 1, 2, 3, and 4. When any three of the four conditions are satisfied, it is determined that an abnormal fainting event has occurred;
[0102] The values of the threshold of the change in the center of gravity height, the threshold of the descending speed of the center of gravity, the acceleration threshold, the threshold of the change in the falling angle, and the angular velocity threshold are set by those skilled in the art according to experience;
[0103] Through the above steps and the optimized design of the abnormal event judgment function, the accuracy of abnormal event detection can be significantly improved, and false alarms and missed alarms can be reduced. At the same time, the multi-dimensional abnormal behavior data also provides richer information for subsequent abnormal event analysis.
[0104] Step S300: Construct a campus security knowledge graph based on multi-system data of the campus, construct a graph neural network model based on the campus security knowledge graph. When an abnormal event occurs, use the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event;
[0105] The specific method of constructing a campus security knowledge graph based on multi-system data of the campus and constructing a graph neural network model based on the campus security knowledge graph is as follows:
[0106] Step S310: Define the top-level concepts of the campus security knowledge graph to form an ontology architecture;
[0107] The top-level concepts include but are not limited to: personnel, events, and places;
[0108] Step S320: Extract different entities, entity attributes, and the relationships between entities from the multi-system database of the campus according to the top-level concepts. Use the entities as the nodes of the graph and the relationships as the edges of the graph to construct a campus security knowledge graph;
[0109] Exemplarily, knowledge elements such as entities, attributes, and relationships are extracted from the business data of the campus multi-system. For example, person entities and their attributes are extracted from the student information system, event and location entities and their attributes are extracted from the campus management system, and the past medical history relationships of persons are extracted from the electronic medical record system. The formed knowledge elements include: person entities: "Zhang San", "Li Si", etc., attributes including name, gender, identity, etc.; event entities: "fainting", "falling and getting injured", "getting sick", etc., attributes including time, location, severity, etc.; location entities: "Teaching Building A", "Student Dormitory B", etc., attributes including name, location, function, etc.; relationships: "Zhang San studies in the Computer Department", "Zhang San has hypoglycemia", etc.
[0110] Step S330: Construct a graph neural network model based on the campus security knowledge graph. The graph neural network model includes a graph convolutional layer and a graph attention layer. The graph convolutional layer is used to perform a weighted sum of the embedded representations of the neighbors of a node, and the weights of the weighted sum are determined by the edge weights between nodes;
[0111] The formula for the embedded representation is: , where represents the embedded representation of node i at the l-th layer, is the set of neighbor nodes of node i, j is a neighbor node, is the embedded representation of neighbor node j at the previous layer, is the weight matrix at the l-th layer, is the bias term at the l-th layer, is the weight of the edge between node i and neighbor node j, is the activation function;
[0112] The weight of the edge between node i and neighbor node j is used to balance the importance of different nodes and can usually be simply set to ;
[0113] Step S340: Introduce an attention mechanism on the basis of the graph convolutional layer to obtain a graph attention layer, calculate the attention weight of different neighbor nodes j to node i, and obtain the embedded representation of node i based on the attention weight;
[0114] The formula for the attention weight is: , where is the LeakyReLU activation function, is the parameter vector of the attention mechanism, T is the transpose symbol, is the parameter matrix of the attention mechanism, , , They are the embedding vectors of node i, neighbor node j, and neighbor node m respectively, and || is the concatenation operation of the embedding vectors;
[0115] The calculation formula for the embedding representation of node i is: , where is the embedding representation of node i obtained based on the attention weights;
[0116] When actually obtaining the node embedding representation, usually choose to use the graph convolution formula or the graph attention formula according to the requirements of the task and the characteristics of the data. The graph convolution formula has a lower computational complexity and is suitable for large-scale graph data, while the graph attention formula can better capture the relationships between nodes but has a higher computational complexity.
[0117] Step S350: Train the graph neural network model in an unsupervised manner. Based on the relationships between nodes in the campus security knowledge graph, form triples (head node, relation, tail node). The triples are used as positive samples, and negative samples are constructed through negative sampling. The positive samples and negative samples are used as training data, and in an unsupervised training manner, the nodes in the campus security knowledge graph are used as inputs, and the embedding representations of the nodes are used as outputs;
[0118] The negative sampling refers to randomly replacing the head node or the tail node of the positive sample;
[0119] Step S360: Use maximizing the likelihood probability of positive samples and minimizing the likelihood probability of negative samples as the training objective. Based on the likelihood probabilities of positive samples and negative samples, obtain the loss function for training the graph neural network model , train the graph neural network model, minimize the loss function, and when the loss function converges, obtain the trained graph neural network model;
[0120] The calculation formula for the loss function is: , where S is the set of positive samples, are the head node i and tail node j in the positive sample, is the head node in the negative sample and tail node , Q is the number of negative samples, is the negative sampling probability distribution, is the expectation operation;
[0121] In the loss function is the likelihood probability term of positive samples, is the likelihood probability term of negative samples;
[0122] The number of negative samples refers to the number of negative samples corresponding to each positive sample;
[0123] When an abnormal event occurs, the specific method of using the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event is as follows:
[0124] Step S370: When an abnormal event occurs, extract the abnormal event entity of the abnormal event and input it into the graph neural network model to output the embedding representation of the abnormal event entity;
[0125] Step S380: In the campus security knowledge graph, with the node where the abnormal event entity is located as the center, search for other entities with the smallest inner product with the abnormal event entity in the embedding space of the campus security knowledge graph, select the top K* entities with the smallest inner product as candidate causes, and calculate their semantic correlation probabilities based on the embedding vectors of the abnormal event entity and the candidate causes , and integrate the semantic correlation probability and the corresponding candidate cause into the reasoning cause of the abnormal event;
[0126] The calculation formula of the semantic correlation probability is: , where is the embedding vector of the abnormal event entity, is the embedding vector of the candidate cause, Cq is the candidate cause set, q is the candidate cause, is the embedding vector of the candidate cause q;
[0127] Figure 2 This is the flowchart of the reasoning cause analysis method for the abnormal event provided by this application;
[0128] Exemplarily, assume that an "event of Zhang San fainting" occurred on campus. Student Zhang San suddenly fainted and lost consciousness during the break exercises at 3 pm. Zhang San has a history of hypoglycemia and did not have lunch on time that day. According to the event description, all entities related to the "abnormal fainting event of Zhang San" are obtained in the knowledge graph: the person entity "Zhang San", the location entity "playground", and the event entity "fainting". Using the pre-trained graph neural network model, the abnormal event entity of the "abnormal fainting event of Zhang San" is input into the model, and the embedding representation of the event entity is obtained according to the relevant attributes. Centered on the embedding representation, other entities semantically related to it are searched in the embedding space of the knowledge graph. For example: the person entity "Zhang San" is related to the event entity "history of hypoglycemia", the event entity "fainting" is related to the event entity "not having lunch on time", and the behavior entity "vigorous exercise" is related to the location entity "playground". Identify "history of hypoglycemia" and "not having lunch on time" in the semantically related entities as candidate causes, denoted as cause1 and cause2 respectively, and denote "fainting" as event. Calculate the semantic correlation probabilities P(cause1|event) = 0.7 and P(cause2|event) = 0.3 for them to cause "fainting", and output the final reasoning cause of "fainting": "history of hypoglycemia" is the most likely cause of "fainting", with a probability of 70%; "not having lunch on time" is also a potential cause, but the possibility is relatively low, with a probability of 30%.
[0129] By integrating multi-dimensional semantic information such as personnel, environment, and activities on campus, the above steps can analyze the causes of abnormal events more comprehensively and deeply, upgrade event alarms from simple event notifications to scenario-based intelligent warnings, and provide more targeted reference information for the decision-making of campus security personnel.
[0130] Step S400: Obtain a sensor combination based on the abnormal audio signal, calculate the signal similarity between each pair of sensors in the sensor combination, and calculate the candidate sound source position with the highest matching degree score according to the signal similarity, and use it as the occurrence position of the abnormal event;
[0131] The specific method for obtaining a sensor combination based on the abnormal audio signal, calculating the signal similarity between each pair of sensors in the sensor combination, and calculating the candidate sound source position with the highest matching degree score according to the signal similarity and using it as the occurrence position of the abnormal event is as follows:
[0132] Step S410: Obtain the short-time energy of each sensor when the abnormal event occurs and the signal-to-noise ratio , and select the sensors with short-time energy greater than the energy threshold and signal-to-noise ratio greater than the signal-to-noise ratio threshold as the sensor combination;
[0133] The calculation formula for the short-time energy is: , where is the short-time energy of the a-th sensor, is the moment when the anomaly occurs, is the duration of the anomaly occurrence, is the abnormal audio signal of the a-th sensor;
[0134] The calculation formula of the signal-to-noise ratio is: , where is the signal-to-noise ratio of the a-th sensor, is the noise signal before and after the abnormal event;
[0135] The values of the energy threshold and the signal-to-noise ratio threshold are set by those skilled in the art according to actual needs;
[0136] Step S420: For the sensor pair in the sensor combination, calculate the signal similarity between the sensor pairs according to their generalized cross-correlation function at the moment when the abnormal event occurs ;
[0137] The calculation formula of the generalized cross-correlation function is: , where represents the propagation delay of the audio signal from sensor ai to sensor aj, , respectively represent the signals of sensors ai and aj in the frequency domain, represents the complex conjugate of the audio signal of sensor aj, is a complex number, where r is the imaginary unit and w is the frequency;
[0138] Step S430: Calculate the matching degree score of the candidate sound source position P by accumulating the signal similarity of each pair of sensor pairs;
[0139] The calculation formula of the matching degree score of the candidate sound source position P is: , where M is the sensor combination, is the propagation delay of the audio signal from sensor ai to sensor aj under the given candidate sound source position P;
[0140] Step S440: Take the candidate sound source position P with the maximum matching degree score as the occurrence position of the abnormal event;
[0141] As Figure 3 shown, it is the flow chart of the method for estimating the occurrence position of the abnormal event provided by this application;
[0142] Compared with single video localization, the above steps incorporate the spatial distribution information of audio signals, enabling more rapid and accurate determination of the location of abnormal events, which is of great significance for campus security personnel to arrive at the scene in a timely manner and provide efficient assistance.
[0143] Step S500: Send the occurrence location of the abnormal event and the reasoning reason for the abnormal event to campus security personnel;
[0144] The specific method of sending the occurrence location of the abnormal event and the reasoning reason for the abnormal event to campus security personnel is as follows: Organize the occurrence location of the abnormal event and the reasoning reason for the abnormal event into alarm information and send it to campus security personnel through the campus communication channel. Campus security personnel make timely responses based on the received information.
[0145] Embodiment 2
[0146] As Figure 4 shown, the campus security control system integrating multiple systems provided by this application includes:
[0147] A video and audio acquisition module, which is used to deploy high-definition cameras and sensor arrays in the campus area, collect video streams and audio signals, and when abnormal audio signals are captured, obtain the video streams before and after the abnormal occurrence time;
[0148] An abnormal event judgment module, which is used to adopt a human pose estimation model based on the attention mechanism to obtain the human pose information in the video stream, calculate abnormal behavior data, and design an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data;
[0149] An abnormal reason reasoning module, which is used to construct a campus security knowledge graph based on campus multi-system data, construct a graph neural network model based on the campus security knowledge graph, and when an abnormal event occurs, use the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyze the reasoning reason for the abnormal event;
[0150] An occurrence location calculation module, which is used to obtain a sensor combination based on the abnormal audio signal, calculate the signal similarity between each pair of sensors in the sensor combination, calculate the candidate sound source location with the highest matching degree score based on the signal similarity, and use it as the occurrence location of the abnormal event;
[0151] A security information receiving module, which is used to send the occurrence location of the abnormal event and the reasoning reason for the abnormal event to campus security personnel.
[0152] Embodiment 3
[0153] According to an embodiment of the present application, a computer-readable storage medium is also provided. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are run by a processor, the campus security control method integrating multiple systems according to the embodiments of the present application can be executed. The storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0154] In addition, according to the embodiments of the present application, the processes described above can be implemented as computer software programs. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application, such as: arranging high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals, and when abnormal audio signals are captured, obtaining the video streams before and after the abnormal occurrence time; using a human pose estimation model based on the attention mechanism to obtain the human pose information in the video stream, calculating abnormal behavior data, and designing an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data; constructing a campus security knowledge graph based on campus multi-system data, constructing a graph neural network model based on the campus security knowledge graph, and when an abnormal event occurs, using the graph neural network model to learn the embedding representation of the abnormal event entity in the abnormal event and analyzing the reasoning reason of the abnormal event; obtaining a sensor combination based on the abnormal audio signal, calculating the signal similarity between each pair of sensors in the sensor combination, calculating the candidate sound source position with the highest matching degree score according to the signal similarity, and using it as the occurrence position of the abnormal event; sending the occurrence position of the abnormal event and the reasoning reason of the abnormal event to the campus security personnel. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0155] The method and system, storage medium of the present application can be implemented in many ways. For example, the method and system, storage medium of the present application can be implemented by software, hardware, firmware or any combination of software, hardware, firmware. The above order of steps for the method is only for illustration, and the method steps of the present application are not limited to the specific order described above unless otherwise specifically stated. In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.
[0156] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0157] As described above in the specific embodiments, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A campus security management and control method integrating multiple systems, characterized in that: include: High-definition cameras and sensor arrays are deployed in the campus area to collect video streams and audio signals. When an abnormal audio signal is captured, the video streams before and after the abnormality occurs are obtained; A human posture estimation model based on the attention mechanism is used to obtain human posture information in the video stream, calculate abnormal behavior data, and design an abnormal event judgment function based on the abnormal behavior data to determine whether an abnormal event has occurred; Build a campus security knowledge graph based on campus multi-system data, and build a graph neural network model based on the campus security knowledge graph. When an abnormal event occurs, use the graph neural network model to learn the embedded representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event; Based on the abnormal audio signal, a sensor combination is obtained, the signal similarity between each sensor pair in the sensor combination is calculated, and the candidate sound source position with the highest matching score is calculated according to the signal similarity, which is used as the occurrence position of the abnormal event; The location of the abnormal event and the reasoning cause of the abnormal event are sent to campus security personnel.
2. The campus security management and control method integrating multiple systems as claimed in claim 1, characterized in that: The specific method of arranging high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals, and when an abnormal audio signal is captured, obtaining the video streams before and after the abnormality occurs is as follows: deploying high-definition cameras and sensor arrays in the campus area, arranging the sensor arrays in a uniformly distributed topological structure near the cameras, and the sensor spacing of the sensor arrays satisfy: , where c is the speed of sound, is the maximum frequency of the audio signal; when the sensor array is abnormal When an abnormal audio signal is captured, obtain the time when the abnormality occurs Duration video stream.
3. The campus security management and control method integrating multiple systems as claimed in claim 2 is characterized in that: The specific method for obtaining the human posture estimation model based on the attention mechanism is: The human posture estimation model based on the attention mechanism is used to estimate the human posture and extract the spatial coordinates of the key points of the human body. In the time dimension, the attention mechanism is introduced to weightedly aggregate the human posture features of different frames. The attention weight of the tth frame is ; In the spatial dimension, the attention mechanism is introduced to extract the feature vectors of different parts of the human body. The attention weight of part j1 is ; The human posture estimation model based on the attention mechanism uses the Attention-LSTM network as the initial network, including the CNN backbone network, the attention mechanism and the LSTM timing module, obtains continuous frames of human body images based on the video stream as input data, uses the human body key points of each frame of the human body image as output data, and obtains the human body posture information estimated by the Attention-LSTM network according to the spatial position of the human body key points, the relative position and distance between the human body key points, and the motion trajectory of the human body key points. The motion trajectory of the human body key points is obtained by connecting the coordinates of the same human body key points in human body images of different frames; The training method of the human posture estimation model is: The Attention-LSTM network is used as the initial network. The end of the Attention-LSTM network is connected to the fully connected layer. The hidden state output by the LSTM timing module is mapped to the two-dimensional coordinates of the key points of the human body. The output of the entire model is the coordinate sequence of the key points of the human body corresponding to each frame of the human body image. A human posture dataset is collected, where the dataset includes human body images and corresponding human body key point coordinate annotations. The human body images are used as input data, and the corresponding human body key point coordinate annotations are used as output data to form training samples. The Attention-LSTM network is trained using the training samples, and the mean square error of the human body key point coordinates is used as the loss function for training the human posture estimation model. The value of minimizing the loss function is used as the training goal. When the loss function converges, the training is completed, and a trained human posture estimation model is obtained.
4. The campus security management and control method integrating multiple systems as claimed in claim 3 is characterized in that: The specific method of obtaining the human body posture information in the video stream, calculating the abnormal behavior data, and designing an abnormal event judgment function for judging whether an abnormal event occurs based on the abnormal behavior data is: The abnormal behavior data include the static time of the human body, the height difference of the center of gravity, the descending speed and acceleration of the center of gravity, the change of the angle of the human body falling and the angular velocity, the average density of key points, and the average movement amplitude of all key points of the human body; the abnormal events include abnormal fainting events, abnormal gathering events, and abnormal limb movements, and the abnormal event judgment function includes abnormal fainting event judgment function, abnormal gathering event judgment function, and abnormal limb movement judgment function; Setting the Inactivity Threshold , calculate the inter-frame displacement of human key points, and for the kth human key point, calculate its inter-frame displacement between the tth frame and the t+1th frame When the inter-frame displacement of a human key point is less than the static threshold, the human key point is considered to be in a static state. Keep still within the frame for a period of , and the static time exceeds the preset time threshold , then the person is considered to be in a stationary state, where is the frame interval time; Calculate the coordinates of the center of gravity of the human body based on the coordinates of the key points of the human body , calculate the height difference between the center of gravity of the human body in the tth frame and the t-1th frame , calculate the center of gravity descent speed and acceleration , calculate the change in the angle of the human body falling to the ground and angular velocity ; Divide the video stream into Grid, calculate the density of key points of the human body in each grid , get the average key point density of the entire picture , set the density threshold , define the abnormal aggregation event judgment function as: when the time threshold When the internal average key point density is continuously greater than the density threshold, it is determined that an abnormal aggregation event has occurred; Calculate the kth human key point in the continuous Average motion within a frame , calculate the average movement amplitude of all human key points , set the amplitude threshold ,The abnormal limb movement judgment function is defined as: when the average movement amplitude of all human key points is greater than the amplitude threshold, it is determined that an abnormal limb movement has occurred; Set the center of gravity height change threshold, center of gravity descent speed threshold and acceleration threshold, falling angle change threshold and angular velocity threshold. Set condition 1 as follows: The center of gravity height difference within the frame is continuously less than the negative center of gravity height change threshold. Condition 2 is: The center of gravity descent speed within the frame is continuously greater than the center of gravity descent speed threshold and the acceleration is continuously greater than the acceleration threshold. Condition three is: The falling angle change in the frame is continuously greater than the falling angle change threshold and the angular velocity is continuously greater than the angular velocity threshold. Condition 4 is: All human key points are continuously The frame remains still and the stillness duration exceeds the preset time threshold ; Design an abnormal fainting event judgment function including the above conditions one, two, three, and four. When any three of the four conditions are met, it is determined that an abnormal fainting event has occurred.
5. The campus security management and control method integrating multiple systems as claimed in claim 4 is characterized in that: The specific method of constructing a campus security knowledge graph based on campus multi-system data and constructing a graph neural network model based on the campus security knowledge graph is: Define the top-level concepts of the campus security knowledge graph and form an ontology architecture; Extract different entities, entity attributes, and relationships between entities from the multi-system database of the campus based on the top-level concepts, use entities as graph nodes and relationships as graph edges to build a campus security knowledge graph; A graph neural network model is constructed based on the campus security knowledge graph. The graph neural network model includes a graph convolution layer and a graph attention layer. The graph convolution layer is used to perform weighted summation on the embedded representations of the neighbors of the node. The weight of the weighted summation is determined by the edge weights between the nodes. Decide; The attention mechanism is introduced on the basis of the graph convolution layer to obtain the graph attention layer, which calculates the attention weights of different neighbor nodes j on node i , based on the attention weight, we get the embedding representation of node i ; The graph neural network model is trained in an unsupervised manner. The relationship between nodes in the campus security knowledge graph constitutes a triple (head node, relationship, tail node). The triple is used as a positive sample, and a negative sample is constructed through negative sampling. The positive sample and the negative sample are used as training data. An unsupervised training method is adopted, and the nodes in the campus security knowledge graph are used as input, and the embedded representation of the node is used as output. The training objectives are to maximize the likelihood of positive samples and minimize the likelihood of negative samples. The loss function of the training graph neural network model is obtained based on the likelihood of positive samples and the likelihood of negative samples. , train the graph neural network model and minimize the loss function. When the loss function converges, the trained graph neural network model is obtained.
6. The campus security management and control method integrating multiple systems as claimed in claim 5, characterized in that: When an abnormal event occurs, the specific method of using the graph neural network model to learn the embedded representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event is: When an abnormal event occurs, the abnormal event entity of the abnormal event is extracted and input into the graph neural network model, and the embedded representation of the abnormal event entity is output; In the campus security knowledge graph, with the node where the abnormal event entity is located as the center, search for other entities with the smallest inner product with the abnormal event entity in the embedding space of the campus security knowledge graph, select the top K* entities with the smallest inner product as candidate causes, and calculate their semantic relevance probabilities based on the embedding vectors of the abnormal event entity and the candidate causes. , the semantically relevant probabilities and the corresponding candidate causes are integrated into the inferred causes of abnormal events.
7. The campus security management and control method integrating multiple systems as claimed in claim 6, characterized in that: The specific method of obtaining a sensor combination based on the abnormal audio signal, calculating the signal similarity between each sensor pair in the sensor combination, and obtaining the candidate sound source position with the highest matching score according to the signal similarity calculation, and using it as the occurrence position of the abnormal event is as follows: Obtain the short-term energy of each sensor when an abnormal event occurs and signal-to-noise ratio , select sensors whose short-time energy is greater than the energy threshold and whose signal-to-noise ratio is greater than the signal-to-noise ratio threshold as the sensor combination; For sensor pairs in a sensor combination , the signal similarity between sensor pairs is calculated based on their generalized cross-correlation function at the time of abnormal event ; The matching score of the candidate sound source position P is calculated by accumulating the signal similarity of each pair of sensors ; The candidate sound source position P with the largest matching score is taken as the occurrence position of the abnormal event.
8. A campus safety management and control system integrating multiple systems, which is used to implement the campus safety management and control method integrating multiple systems according to any one of claims 1 to 7, characterized in that: include: The video and audio acquisition module is used to arrange high-definition cameras and sensor arrays in the campus area to collect video streams and audio signals. When an abnormal audio signal is captured, the video streams before and after the abnormality occurs are obtained; An abnormal event judgment module is used to adopt a human posture estimation model based on an attention mechanism to obtain human posture information in a video stream, calculate abnormal behavior data, and design an abnormal event judgment function based on the abnormal behavior data to determine whether an abnormal event has occurred; The abnormal cause reasoning module is used to build a campus security knowledge graph based on campus multi-system data, and a graph neural network model based on the campus security knowledge graph. When an abnormal event occurs, the graph neural network model is used to learn the embedded representation of the abnormal event entity in the abnormal event and analyze the reasoning cause of the abnormal event; An occurrence position calculation module is used to obtain a sensor combination based on the abnormal audio signal, calculate the signal similarity between each sensor pair in the sensor combination, and obtain the candidate sound source position with the highest matching score based on the signal similarity calculation, and use it as the occurrence position of the abnormal event; The security information receiving module is used to send the location of the abnormal event and the reasoning cause of the abnormal event to the campus security personnel.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which is suitable for being loaded by a processor to execute the steps in the campus security management and control method integrating multiple systems as described in any one of claims 1-7.
Citation Information
Patent Citations
Campus security event early warning method based on space-time analysis and video AI
CN118153953A
Human motion posture recognition and evaluation method and system
CN111401270A
Campus abnormal event monitoring system based on deep network learning
CN115035439A