Hospital security early warning analysis method and system based on AI
By introducing AI-based early warning analysis methods in hospital security systems, the existing systems rely on manual monitoring and lack of intelligent early warning are solved, and intelligent analysis and automatic early warning of surveillance videos are realized, and security capabilities and management efficiency are improved.
Patent Information
- Application Number
- CN202510070296.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing hospital security system relies on manual monitoring and cannot intelligently analyze behavior patterns, resulting in the inability to identify potential risks in a timely manner and lack of intelligent early warning and behavioral analysis capabilities.
Using AI-based hospital security early warning analysis method, surveillance videos are analyzed in real time through artificial intelligence technology, potential threats are identified and multi-level early warnings are automatically triggered. Specific steps include video data decoding, noise reduction processing, object detection, Kalman filter prediction, behavior classification and abnormal detection.
It realizes intelligent analysis of surveillance videos, can automatically identify suspicious behaviors and trigger early warnings, improves the ability to adapt to hospital security scenarios, reduces false alarm problems, and provides an efficient and intelligent security management solution.
Smart Images

Figure CN119992653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security warning technology, and more specifically, to an AI-based hospital security warning analysis method and system. Background Art
[0002] As an important public place, hospitals have extremely high security protection requirements. The security system needs to promptly identify scenes such as abnormal crowd gathering, unauthorized entry into key areas (pharmacies, computer rooms, etc.), and suspicious behaviors (such as violence, abandoned items, etc.): Traditional monitoring systems mainly rely on video recording and simple motion detection functions, and are unable to intelligently analyze behavior patterns, resulting in the inability to timely identify potential risks. The existing hospital security system has the following problems: 1. Reliance on manual monitoring: Security monitoring mainly relies on real-time viewing by security personnel, which is prone to omissions and misjudgments; 2. Lack of intelligent early warning: The system cannot automatically identify suspicious behaviors in the monitoring screen and cannot respond quickly; 3. Insufficient behavior analysis capabilities: The existing system's identification of abnormal behaviors is mostly based on simple rules, which makes it difficult to cope with complex scenarios. Summary of the invention
[0003] The purpose of the present invention is to provide an AI-based hospital security early warning analysis method and system, which realizes real-time analysis of surveillance videos through artificial intelligence technology, identifies potential threats and automatically triggers multi-level early warnings.
[0004] The above technical objectives of the present invention are achieved through the following technical solutions:
[0005] In the first aspect, the present application provides an AI-based hospital security early warning analysis method, comprising the following specific steps:
[0006] Decode the collected real-time video data into continuous image frames, and perform noise reduction on the continuous image frames through a Gaussian filter;
[0007] Based on each frame of the continuous image frames after noise reduction processing, a preset target detection algorithm is used to perform target detection processing on the image, and when a tracking target exists in the image, state information of the corresponding tracking target is obtained, and the state information includes position information and speed information;
[0008] The prediction equation of Kalman filtering is used to obtain the predicted state of each tracking target in each frame image, the predicted state of the current frame image is matched with the state information, and the matching result of the corresponding tracking target is obtained;
[0009] For the tracking target whose matching result is successfully matched, the state of the tracking target is updated based on the predicted state and state information of the current frame to obtain the target information of the corresponding tracking target;
[0010] Based on the target information of each tracking target, the behavior of each tracking target is classified through a preset three-dimensional convolutional neural network, and the behavior category of each tracking target is obtained;
[0011] Based on the behavior category of each tracking target, the anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target. The warning level of each tracking target is obtained according to the anomaly score combined with the corresponding behavior category. The warning levels include low risk, medium risk and high risk.
[0012] Based on the above technical solution, the present invention can also be improved as follows.
[0013] Furthermore, the above-mentioned noise reduction process is performed on the continuous image frames by using a Gaussian filter, specifically:
[0014]
[0015] Wherein, the pixel point (x, y) in the image after denoising is expressed as G(x, y), and σ represents the standard deviation of the Gaussian distribution in the Gaussian filter.
[0016] Furthermore, the predicted state of each tracking target in each frame of the above image is specifically:
[0017]
[0018] In the formula, represents the predicted state of one of the tracked targets in the t-th frame image, x t-1 Indicates the predicted state of the tracked target in the t-1th frame image, u t represents the control quantity in the prediction equation of Kalman filter, A represents the state transfer matrix, and B represents the control input matrix of the prediction equation.
[0019] Furthermore, the state of the tracking target is updated based on the predicted state and state information of the current frame to obtain the target information of the corresponding tracking target, specifically:
[0020] in:
[0021] K t =P t|t-1 H T (HP t|t-1 H T +R) -1 ;P t|t =(IK t H)P t|t-1
[0022] In the formula, represents the target information of the corresponding tracking target obtained after the state is updated in the t-th frame image, represents the state information of the tracked target in the t-th frame image, and the predicted value of the t-th frame based on the information of the t-1 frame, K t represents the Kalman gain, z t represents the state information of the tracked target in the t-th frame image, H represents the measurement matrix in the target detection algorithm, and P t|t It represents the prediction covariance matrix of the Kalman filter prediction equation when processing the t-th frame image, P t|t-1 It represents the prediction covariance matrix when the prediction equation of the Kalman filter processes the t-1th frame image. The subscript T represents the vector transpose operation. R represents the measurement noise covariance matrix in the target detection algorithm. I represents the unit matrix with the same dimension as the prediction covariance matrix.
[0023] Furthermore, the above anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target, specifically:
[0024] in,
[0025] Where s(x) represents the anomaly score of the data point x of the tracking target, h(x) represents the path length of the data point x in the isolation forest detection algorithm, E(h(x)) represents the expected value of the path length h(x), C(n) represents the normalization constant, which is related to the number of samples n, and H(i) represents the harmonic number, H(i)=ln(i)+0.5772156649.
[0026] Furthermore, the above method further includes: obtaining true feedback of each tracking target, and updating the model parameters of the anomaly detection model according to the true feedback, specifically:
[0027]
[0028] In the formula, θ t+1 represents the model parameters of the anomaly detection model at the t+1th frame image, θ t represents the model parameters of the anomaly detection model at the t-th frame image, η represents the learning rate of the anomaly detection model, The loss function L of the anomaly detection model is expressed as t gradient.
[0029] Furthermore, the above method also includes:
[0030] The audio data and environmental sensor data in the monitoring environment are obtained, and audio features and sensor features are obtained according to the audio data and environmental sensor data; wherein the audio features of the audio data are extracted using a convolutional neural network, specifically:
[0031] Where M(f, t) represents the audio feature, X k represents the frequency domain signal of the audio data, H(f, k) represents the Mel filter, k represents the index variable, and K represents the number of items involved in the summation;
[0032] The Transformer model is used to align audio features and sensor features with video features in real-time video data through timestamps to obtain a feature matrix including multiple feature types, and early warning analysis is performed based on the feature matrix.
[0033] Furthermore, the above method also includes:
[0034] Obtain the feature map of each layer of the deep neural network-based anomaly detection model, and generate a heat map in the monitoring environment based on the weight of the feature map to explain the model decision process.
[0035] Furthermore, the weights of the above feature maps are specifically:
[0036] In the formula, α k represents the weight of the kth channel of the feature map, Z represents the normalization processing constant, represents the activation value of the kth channel of the feature map, i and j represent the spatial position of the feature map, and y c represents the anomaly score of behavior category c obtained by the anomaly detection model;
[0037] Heatmaps used to explain the model's decision-making process, specifically:
[0038] In the formula, represents the heat map, ReLU is the rectified linear unit, A k represents the activation value of the kth channel, α k Represents the weight of the kth channel of the feature map.
[0039] In a second aspect, the present application provides an AI-based hospital security early warning analysis system, which is applied to an AI-based hospital security early warning analysis method according to any one of the first aspects, including:
[0040] The data acquisition and preprocessing module is used to decode the acquired real-time video data into continuous image frames and perform noise reduction on the continuous image frames through a Gaussian filter;
[0041] A state information recognition module is used to perform target detection processing on each frame of the continuous image frames after noise reduction using a preset target detection algorithm, and obtain state information of the corresponding tracking target when there is a tracking target in the image, the state information including position information and speed information;
[0042] The prediction matching module is used to obtain the predicted state of each tracking target in each frame image by using the prediction equation of the Kalman filter, match the predicted state of the current frame image with the state information, and obtain the matching result of the corresponding tracking target;
[0043] A state update module is used to update the state of the tracking target for which the matching result is successful, based on the predicted state and state information of the current frame, and obtain the target information of the corresponding tracking target;
[0044] A behavior category recognition module is used to classify the behavior of each tracking target through a preset three-dimensional convolutional neural network based on the target information of each tracking target, and obtain the behavior category of each tracking target;
[0045] The anomaly detection and warning module is used to calculate the anomaly score of each tracking target based on the behavior category of each tracking target using the anomaly detection model based on the isolation forest detection algorithm, and obtain the warning level of each tracking target based on the anomaly score combined with the corresponding behavior category. The warning levels include low risk, medium risk and high risk.
[0046] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods in the first aspect when executing the computer program.
[0047] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute any one of the methods in the first aspect.
[0048] Compared with the prior art, the present invention has at least the following beneficial effects:
[0049] In the present application, firstly, by combining the target detection algorithm YOLO with the three-dimensional convolutional neural network 3D-CNN behavior classification, possible targets are detected in each frame of the image, and information such as the location and category of the target is given, and the final target information of each tracking target is obtained through the prediction equation of the Kalman filter; secondly, the behavior category is identified and classified according to the final target information of each tracking target, so as to determine the behavior category of the tracking target, and then the abnormal score of each tracking target is calculated according to the abnormal detection model; finally, the warning level is divided according to the abnormal score and based on the preset level score, and the relevant staff can perform the corresponding warning behavior according to the obtained warning level. For example, when the warning level is low risk, the security personnel can be notified by SMS; when the warning level is medium risk, the alarm can be triggered and the alarm can be sent to the management center; when the warning level is high risk, the video recording can be linked and the real-time picture can be sent to the security platform; by integrating multiple deep learning methods (target detection, anomaly detection, etc.) into a unified security warning process, the dynamic warning classification mechanism is combined with the graded response strategy to avoid the false alarm problem of the simple alarm system and improve the adaptability to key areas and high-risk scenarios.
[0050] In this application, in response to the specific security needs of the hospital (such as key area protection, equipment area intrusion, and abnormal behavior identification), a large number of medical scene features are incorporated into the early warning analysis process to improve the ability to identify specific behaviors; among them: by using a Gaussian filter to reduce the noise of continuous image frames, the Gaussian noise in the image can be effectively removed to obtain a smooth image after Gaussian filtering; dynamic risk assessment and classification are achieved through the isolation forest algorithm; in the incremental learning of the anomaly detection model, the model will be continuously updated with the arrival of new data. The model updates its parameters by calculating the loss function gradient based on the current data to better adapt to the new data. This method enables the model to continuously learn and improve without retraining the entire model. For example, in an online recommendation system, as new user behavior data is continuously generated, the system can use incremental learning to update the recommendation model to provide more accurate recommendations; the above method supports multiple behavior patterns, adapts to complex hospital scenarios, and realizes comprehensive perception and intelligent response to hospital security scenarios through artificial intelligence technology, providing an efficient and intelligent solution for hospital safety management, with high application value and innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0052] Figure 1 A method flow chart of the early warning analysis method in an embodiment of the present invention;
[0053] Figure 2 A connection diagram of an early warning analysis system in an embodiment of the present invention;
[0054] Figure 3 Schematic diagram of the connection of electronic equipment in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0056] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0058] Embodiment 1: In order to realize real-time analysis of surveillance videos through artificial intelligence technology to identify potential threats and automatically trigger multi-level warnings, this embodiment provides an AI-based hospital security warning analysis method, including the following specific steps:
[0059] S1, decoding the collected real-time video data into continuous image frames, and performing noise reduction processing on the continuous image frames through a Gaussian filter.
[0060] Among them, decoding the real-time video data can be expressed as: F t =Decode(V), where F t is the t-th frame image, V is the original video stream, and Decode is the decoding function. The function of the Gaussian filter is to perform weighted averaging on each pixel in the image, and the weight is determined by the Gaussian function.
[0061] Specifically, Gaussian filtering is widely used in image processing, such as in computer vision, medical image processing and other fields. It can effectively remove Gaussian noise (a common type of noise whose grayscale value distribution conforms to Gaussian distribution) in images. For example, in medical X-ray images, Gaussian filtering can reduce the noise in the image, allowing doctors to observe the lesions more clearly; in the target detection task of computer vision, Gaussian filtering can pre-process the image and improve the performance of subsequent algorithms. The closer the pixel is to the center pixel, the greater the weight, and the farther the pixel is, the smaller the weight.
[0062] Optionally, the above-mentioned noise reduction process is performed on the continuous image frames by using a Gaussian filter, specifically:
[0063]
[0064] In the formula, the pixel point (x, y) in the image after noise reduction processing is expressed as G(x, y), σ represents the standard deviation of the Gaussian distribution in the Gaussian filter, which determines the width of the Gaussian filter. A smaller σ value will make the filter narrower and retain more details of the image, but the noise reduction effect may be poor; a larger σ value will make the filter wider and the noise reduction effect better, but it may over-blur the image; is a normalization factor that ensures that the integral of the Gaussian function over the entire two-dimensional plane is equal to 1. This step is important because it ensures that the filtering process does not change the overall brightness of the image.
[0065] Among them, for each pixel point (i, j) in the image, the calculation process of Gaussian filtering is as follows: first determine a neighborhood centered on (i, j) (usually a square area, such as 3x3, 5x5, etc.); then for each pixel point (x, y) in the neighborhood, calculate its Gaussian weight G(xi, yj); finally, multiply the grayscale values of all pixels in the neighborhood by the corresponding Gaussian weight, and sum them up to get the new grayscale value of the center pixel point (i, j). By performing such operations on each pixel in the image, a smoothed image after Gaussian filtering can be obtained, thereby removing noise.
[0066] S2, based on each frame of the continuous image frames after noise reduction processing, a preset target detection algorithm is used to perform target detection processing on the image, and when a tracking target exists in the image, state information of the corresponding tracking target is obtained, and the state information includes position information and speed information.
[0067] Among them, the YOLO algorithm can be used to detect targets (such as people and objects) in the monitoring screen, and its algorithm formula can be expressed as: Among them, P(Object) is the probability of whether the target exists, To predict the intersection-over-union ratio of the bounding box and the true box; specifically, target detection can usually be completed by other target detection algorithms (such as Faster R-CNN, etc.), which detect possible targets in each frame of the image and provide information such as the target's position (bounding box) and category. When the target detection algorithm detects a target, it provides the target's initial position (bounding box), and this position information is used as the initial measurement value of the Kalman filter.
[0068] S3, using the prediction equation of Kalman filtering to obtain the predicted state of each tracking target in each frame image, matching the predicted state of the current frame image with the state information, and obtaining the matching result of the corresponding tracking target.
[0069] Among them, the target tracking adopts the SORT algorithm, and the trajectory prediction is realized through the Kalman filter. The SORT (Simple Online and Realtime Tracking) algorithm is a target tracking algorithm based on the Kalman filter. It is mainly used to track the target in real time in the video sequence; the Kalman filter is a recursive estimator used to estimate the state of the dynamic system from a series of measurements containing noise. In target tracking, the Kalman filter is used to predict the state information such as the position and speed of the target. The application of Kalman filtering in target tracking, Kalman filtering is divided into two main steps: prediction and update, and its prediction process is mainly reflected in: the predicted state of each tracking target in each frame of the above image, specifically:
[0070]
[0071] In the formula, represents the predicted state of one of the tracked targets in the t-th frame image, x t-1 Indicates the predicted state of the tracked target in the t-1th frame image, u t represents the control quantity in the prediction equation of Kalman filter, A represents the state transfer matrix, and B represents the control input matrix of the prediction equation.
[0072] The state transition matrix A described above describes how the target state evolves from one moment to the next; for example, in a two-dimensional plane, if the target state vector (where x, y are positions, is the speed), for the uniform linear motion model, the state transfer matrix A can be expressed as: Δt is the time interval.
[0073] Among them, the control input matrix B and the control quantity ut are usually used to consider the impact of external control factors on the target state. In a simple target tracking scenario, if there is no external control factor, B can be a zero matrix and ut can be a zero vector.
[0074] S4, for the tracking target whose matching result is successful, update the state of the tracking target based on the predicted state and state information of the current frame to obtain the target information of the corresponding tracking target.
[0075] As mentioned above, Kalman filtering is divided into two main steps: prediction and update. The update process is:
[0076] 1. When a new measurement is obtained (for example, the target position obtained by the target detection algorithm), the Kalman filter will be updated. The update step involves the measurement equation: t =Hx t +v t , z t is the state information of the tracked target in the t-th frame image, H represents the measurement matrix in the target detection algorithm, and v t is the measurement noise;
[0077] 2. Kalman gain K t Used to combine predicted values and measured values, K t =P t|t-1 H T (HP t|t-1 H T +R) -1 , the updated state estimate is: Updated covariance matrix: P t|t =(IK t H)P t|t-1 .
[0078] Specifically, the Kalman filter predicts the position of the target based on the prediction equation, and then updates the state by combining the update equation with the new measurement value; if in a certain frame, the target detection algorithm does not detect the target, but the Kalman filter predicts that the target exists (that is, the predicted state covariance is within a certain range), the target is considered to be temporarily blocked or lost, but its position is still predicted until the target is detected again or the predicted covariance exceeds a certain threshold, indicating that the target may have left the scene; the complete process of the SORT algorithm is:
[0079] 1. Initialization: When the first frame of image comes in, the target detection algorithm is used to detect the target, and a Kalman filter is initialized for each detected target, whose state vector includes information such as position and velocity.
[0080] 2. Prediction: For each new frame, use the Kalman filter prediction equation to predict the new position of all tracked targets.
[0081] 3. Matching: Match the target detected by the target detection algorithm in the current frame with the target predicted by the Kalman filter. The matching method can use methods such as the Hungarian algorithm to match based on features such as the position and appearance of the target.
[0082] 4. Update: For successfully matched targets, the status is updated using the Kalman filter update equation combined with the new position information provided by the target detection algorithm.
[0083] 5. Deletion and addition: If a tracking target is not matched successfully within a certain number of frames (it may have left the scene), the tracking target will be deleted; if the target detection algorithm detects a new target, a new Kalman filter will be initialized for it and tracking will begin.
[0084] Through the above steps, the SORT algorithm combined with the Kalman filter can effectively track the target in the video sequence and handle situations such as occlusion and temporary loss of the target.
[0085] Optionally, the state of the tracking target is updated based on the predicted state and state information of the current frame to obtain target information corresponding to the tracking target, specifically:
[0086] in:
[0087] K t =P t|t-1 H T (HP t|t-1 H T +R) -1 ;P t|t =(IK t H)P t|t-1
[0088] In the formula, represents the target information of the corresponding tracking target obtained after the state is updated in the t-th frame image, represents the state information of the tracked target in the t-th frame image, and the predicted value of the t-th frame based on the information of the t-1 frame, K t represents the Kalman gain, z t represents the state information of the tracked target in the t-th frame image, H represents the measurement matrix in the target detection algorithm, and P t|t It represents the prediction covariance matrix of the Kalman filter prediction equation when processing the t-th frame image, P t|t-1It represents the prediction covariance matrix when the prediction equation of the Kalman filter processes the t-1th frame image. The subscript T represents the vector transpose operation. R represents the measurement noise covariance matrix in the target detection algorithm. I represents the unit matrix with the same dimension as the prediction covariance matrix. If the prediction covariance matrix is a 4x4 matrix, then I is a 4X4 unit matrix with elements on the main diagonal being 1 and other elements being 0.
[0089] For example, in a video tracking system, a moving object (such as a car) is tracked. It can contain information such as the position (horizontal coordinate, vertical coordinate) and speed (horizontal speed, vertical speed) of the car at the t-1 frame. These information are the most accurate estimates obtained after being updated by the Kalman filter algorithm. Taking the above car tracking as an example, if the position of the car at the t-1 frame is (x t-1 ,y t-1 ), the speed is (v x,t-1 ,v y,t-1 ), then according to the motion model (such as the uniform motion model), the approximate position and speed of the car at the frame can be predicted, and this predicted value is
[0090] S5, based on the target information of each tracking target, a preset three-dimensional convolutional neural network is used to classify the behavior of each tracking target, and the behavior category of each tracking target is obtained.
[0091] Among them, Isolation Forest is an unsupervised learning algorithm for anomaly detection, based on the assumption that abnormal data points are more likely to be isolated in a data set (i.e., separated from other data points); the algorithm divides the data by constructing a random binary tree, and normal data points usually require more divisions to be isolated, while abnormal data points can be isolated with fewer divisions; its anomaly detection process is: When using Isolation Forest to detect anomalies:
[0092] 1. Construct an isolation forest, which includes randomly selecting features and randomly selecting split points to construct multiple binary trees.
[0093] 2. For each data point x, calculate its path length h(x) on each tree in the isolation forest.
[0094] 3. Then calculate the expected value of the path length E(h(x)), which is the average of the path lengths of the data point x on all trees.
[0095] 4. Use formula Calculate the anomaly score s(x).
[0096] If s(x) is close to 1, then data point x is likely to be an outlier; if s(x) is close to 0, then data point x is likely to be a normal point. For example, suppose in a behavior pattern dataset, there is a data point x whose path length h(x) is very short, which means it is easily isolated in the isolation forest. By calculating s(x), if a high value (close to 1) is obtained, then this data point x will be judged as an outlier and may represent some abnormal behavior pattern.
[0097] Optionally, the above anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target, specifically:
[0098] in,
[0099] Where s(x) represents the anomaly score of the data point x of the tracking target, h(x) represents the path length of the data point x in the isolation forest detection algorithm, E(h(x)) represents the expected value of the path length h(x), C(n) represents the normalization constant, which is related to the number of samples n, and H(i) represents the harmonic number, H(i)=ln(i)+0.5772156649.
[0100] S6, based on the behavior category of each tracking target, the anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target, and the warning level of each tracking target is obtained according to the anomaly score combined with the corresponding behavior category. The warning levels include low risk, medium risk and high risk.
[0101] Among them, when dividing the warning level according to the abnormality score, the warning level can be divided into low risk, medium risk and high risk. For example, if the range of the abnormality score is 0-100, then 0-30 can be divided into low risk, 30-60 can be divided into medium risk, and 60-100 can be divided into high risk; and after confirming the warning level, the alarm can be automatically triggered and the relevant systems can be linked. For example, low risk can notify security personnel by SMS, medium risk can trigger an alarm and send an alarm to the management center, and high risk can be linked to video recording and send real-time images to the security platform.
[0102] Furthermore, in order to correct the anomaly detection model, the real feedback of each tracking target can be obtained, and the model parameters of the anomaly detection model can be updated according to the real feedback, specifically:
[0103]
[0104] In the formula, θ t+1 represents the model parameters of the anomaly detection model at the t+1th frame image, θ trepresents the model parameters of the anomaly detection model at the t-th frame image, η represents the learning rate of the anomaly detection model, which is used to control the step size of each parameter update. If it is too large, it may cause excessive parameter updates or even failure to converge; if it is too small, the parameter updates will be very slow, resulting in a long training time; for example, when training a linear regression model, a suitable learning rate can be 0.01; The loss function L of the anomaly detection model is expressed as t The gradient indicates the direction in which the loss function grows fastest under the current parameter value. If you want to minimize the loss function during training, you need to update the parameters in the opposite direction of the gradient. For example, in a binary logistic regression model, the loss function can be the cross entropy loss, and the weights are updated by calculating the gradient of the cross entropy loss with respect to the model weights.
[0105] In incremental learning, the model is continuously updated as new data arrives. The formula indicates that each time new data is received (at time step t), the model updates its parameters based on the gradient of the loss function calculated from the current data to better adapt to the new data; this approach enables the model to continuously learn and improve without the need to retrain the entire model; for example, in an online recommendation system, as new user behavior data is continuously generated, the system can use incremental learning to update the recommendation model to provide more accurate recommendations.
[0106] Optionally, the above method further includes:
[0107] The audio data and environmental sensor data in the monitoring environment are obtained, and audio features and sensor features are obtained according to the audio data and environmental sensor data; wherein the audio features of the audio data are extracted using a convolutional neural network, specifically:
[0108] Where M(f, t) represents the audio feature, X k represents the frequency domain signal of audio data, H(f, k) represents the Mel filter, and k represents the index variable used to traverse the elements involved in the calculation. It starts from 1 and increases by 1 each time until it reaches K, where K represents the number of items involved in the summation. This formula may be used in the field of audio processing to extract audio features with more semantic and perceptual meanings, which can be further applied to various audio-related tasks and applications to improve the understanding and processing capabilities of audio content.
[0109] The Transformer model is used to align audio features and sensor features with video features in real-time video data through timestamps to obtain a feature matrix including multiple feature types, and early warning analysis is performed based on the feature matrix.
[0110] Among them, in addition to video data, audio data (such as scream detection) and environmental sensor data (such as access control sensors) can also be combined to further optimize the early warning mechanism; Convolutional Neural Networks (CNN) can be used to process audio spectrograms, and the environmental sensor (such as access control sensor) data and the video analysis results are aligned through timestamps to form a unified feature matrix. When forming the feature matrix, a multimodal Transformer can be used to process heterogeneous data: Output = Transformer (Q, K, V), Q, K, V are the query, key and value of video, audio and sensor features respectively; thereby improving the robustness of anomaly detection, especially in scenarios where visual signals are lost, and multimodal data provides clues in more dimensions to reduce the false alarm rate.
[0111] Optionally, the above method further includes:
[0112] Obtain the feature map of each layer of the deep neural network-based anomaly detection model, and generate a heat map in the monitoring environment based on the weight of the feature map to explain the model decision process.
[0113] Among them, the weight of the above feature map is specifically:
[0114] In the formula, α k represents the weight of the kth channel of the feature map, Z represents the normalization processing constant, represents the activation value of the kth channel of the feature map, i and j represent the spatial position of the feature map, and y c Represents the anomaly score of behavior category c obtained by the anomaly detection model.
[0115] The above heat map used to explain the model decision process is as follows:
[0116] In the formula, represents the heat map, ReLU is the rectified linear unit, A k represents the activation value of the kth channel, α k Represents the weight of the kth channel of the feature map.
[0117] Specifically, since complex models of security systems (such as anomaly detection models based on deep neural networks) are often regarded as "black boxes" and difficult to explain, the introduction of a visual interpretation module can display the key areas and features of anomaly detection results in real time, thereby improving user trust and understanding.
[0118] Furthermore, in the anomaly detection model, feature maps are gradually generated through convolutional layers and activation functions; Convolutional layer: The input image is convolved through multiple convolutional kernels (filters), and each convolutional kernel generates a feature map of a channel. For example, a convolutional layer has 64 convolutional kernels, so 64 channel feature maps will be generated. Activation function: The output of the convolutional layer is usually transformed nonlinearly through an activation function (such as ReLU) to obtain an activated feature map.
[0119] Among them, the gradient of the network is calculated to generate a heat map of the area of interest, which is used to explain the decision-making process of the model; specifically, based on the output heat map, the system can mark suspicious areas in the monitoring screen in real time. For example, in behavior analysis, the key action area of the behavior subject can be marked; in object detection, abnormal areas such as abandoned objects can be marked; its interpretability: provides a visual way to explain why the model makes a certain prediction. By observing the heat map, you can understand the image area that the model pays attention to when making predictions; in general, by calculating the gradient of the feature map and generating a heat map, it helps us understand the decision-making process of the model and realize the marking of abnormal areas in monitoring and other applications. According to the above solution, users can intuitively understand the basis for detection and improve their trust in the judgment of the system. At the same time, the heat map can be used as security training material to assist in subsequent model optimization.
[0120] Optionally, based on anomaly detection, the behavior trajectory can be analyzed through a time series prediction model to predict possible abnormal behaviors in the future and provide more advanced warnings for security personnel; for example, the time series model is a Transformer model, and the input sequence of the model can be X = [x1, x2, ..., x t ], the dependency between behavioral features is calculated through the model’s attention mechanism, which can be expressed as is the query,key and value matrix of video, audio,sensor features, d k is the feature dimension, thereby outputting future behavior predictions:
[0121] Among them, the above scheme can be used to predict whether the crowd will gather further and determine whether items may be abandoned; by predicting future behavior, security personnel can intervene in advance to reduce the probability of accidents and enhance the system's proactive early warning capabilities.
[0122] Optionally, for each model proposed in this embodiment, some deep learning computing tasks can be deployed to the camera end through model compression and edge computing optimization to reduce dependence on the central server; for example: remove neurons and connections that have little impact on the results, and remove them according to the weight parameters of the neurons by setting a removal threshold, that is, if the weight parameters of the neurons are less than the removal threshold, they are removed; floating-point operations can also be replaced with fixed-point operations to reduce computational complexity; thereby significantly reducing server load and improving system response speed.
[0123] Embodiment 2: The embodiment of the present application provides an AI-based hospital security early warning analysis system, which is applied to an AI-based hospital security early warning analysis method in any one of Embodiment 1, such as Figure 2 As shown, including:
[0124] The data acquisition and preprocessing module is used to decode the acquired real-time video data into continuous image frames and perform noise reduction on the continuous image frames through a Gaussian filter.
[0125] The state information recognition module is used to perform target detection processing on each frame of the continuous image frames after noise reduction using a preset target detection algorithm, and obtain the state information of the corresponding tracking target when there is a tracking target in the image, and the state information includes position information and speed information.
[0126] The prediction matching module is used to obtain the predicted state of each tracking target in each frame image using the prediction equation of the Kalman filter, match the predicted state of the current frame image with the state information, and obtain the matching result of the corresponding tracking target.
[0127] The state update module is used to update the state of the tracking target for which the matching result is successful, based on the predicted state and state information of the current frame, and obtain the target information of the corresponding tracking target.
[0128] The behavior category recognition module is used to classify the behavior of each tracking target based on the target information of each tracking target through a preset three-dimensional convolutional neural network, and obtain the behavior category of each tracking target.
[0129] The anomaly detection and warning module is used to calculate the anomaly score of each tracking target based on the behavior category of each tracking target using the anomaly detection model based on the isolation forest detection algorithm, and obtain the warning level of each tracking target based on the anomaly score combined with the corresponding behavior category. The warning levels include low risk, medium risk and high risk.
[0130] Embodiment 3: This embodiment of the present application provides an electronic device, such as Figure 3As shown, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any method in Embodiment 1 is implemented.
[0131] Embodiment 4: The embodiment of the present application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute any method in Embodiment 1.
[0132] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A hospital security early warning analysis method based on AI, characterized in that: The specific steps include: Decoding the collected real-time video data into continuous image frames, and performing noise reduction processing on the continuous image frames through a Gaussian filter; Based on each frame of the continuous image frames after noise reduction processing, a preset target detection algorithm is used to perform target detection processing on the image, and when a tracking target exists in the image, state information of the corresponding tracking target is obtained, wherein the state information includes position information and speed information; The prediction state of each tracking target in each frame image is obtained by using the prediction equation of Kalman filtering, the predicted state of the current frame image is matched with the state information, and the matching result of the corresponding tracking target is obtained; For the tracking target whose matching result is a successful match, updating the state of the tracking target based on the predicted state and the state information of the current frame to obtain target information corresponding to the tracking target; Based on the target information of each tracking target, the behavior of each tracking target is classified through a preset three-dimensional convolutional neural network, and the behavior category of each tracking target is obtained; Based on the behavior category of each tracking target, the anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target, and the warning level of each tracking target is obtained according to the anomaly score combined with the corresponding behavior category, and the warning level includes low risk, medium risk and high risk.
2. According to claim 1, the AI-based hospital security early warning analysis method is characterized in that: The noise reduction process of the continuous image frames is performed by using a Gaussian filter, specifically: Wherein, the pixel point (x, y) in the image after denoising is expressed as G(x, y), and σ represents the standard deviation of the Gaussian distribution in the Gaussian filter.
3. According to claim 1, the AI-based hospital security early warning analysis method is characterized in that: The predicted state of each tracking target in each frame of image is specifically: In the formula, represents the predicted state of one of the tracked targets in the t-th frame image, x t-1 Indicates the predicted state of the tracked target in the t-1th frame image, u t represents the control quantity in the prediction equation of Kalman filter, A represents the state transfer matrix, and B represents the control input matrix of the prediction equation.
4. According to claim 1, the AI-based hospital security early warning analysis method is characterized in that: The updating of the state of the tracking target based on the predicted state and the state information of the current frame to obtain the target information of the corresponding tracking target is specifically: in: K t =P t|t-1 H T (HP t|t-1 H T +R) -1 ;P t|t =(I-K t H)P t|t-1 In the formula, represents the target information of the corresponding tracking target obtained after the state is updated in the t-th frame image, represents the state information of the tracked target in the t-th frame image, and the predicted value of the t-th frame based on the information of the t-1 frame, K t represents the Kalman gain, z t represents the state information of the tracked target in the t-th frame image, H represents the measurement matrix in the target detection algorithm, and P t|t It represents the prediction covariance matrix of the Kalman filter prediction equation when processing the t-th frame image, P t|t-1 It represents the prediction covariance matrix when the prediction equation of the Kalman filter processes the t-1th frame image. The subscript T represents the vector transpose operation. R represents the measurement noise covariance matrix in the target detection algorithm. I represents the unit matrix with the same dimension as the prediction covariance matrix.
5. The AI-based hospital security early warning analysis method according to claim 1 is characterized in that: The anomaly detection model based on the isolation forest detection algorithm is used to calculate the anomaly score of each tracking target, specifically: Where s(x) represents the anomaly score of the data point x of the tracking target, h(x) represents the path length of the data point x in the isolation forest detection algorithm, E(h(x)) represents the expected value of the path length h(x), C(n) represents the normalization constant, which is related to the number of samples n, and H(i) represents the harmonic number, H(i)=ln(i)+0.5772156649.
6. The AI-based hospital security early warning analysis method according to claim 1 is characterized in that: The method further includes: obtaining true feedback of each tracking target, and updating the model parameters of the anomaly detection model according to the true feedback, specifically: In the formula, θ t+1 represents the model parameters of the anomaly detection model at the t+1th frame image, θ t represents the model parameters of the anomaly detection model at the t-th frame image, η represents the learning rate of the anomaly detection model, The loss function L of the anomaly detection model is expressed as t gradient.
7. The AI-based hospital security early warning analysis method according to claim 1 is characterized in that: The method further comprises: Acquire audio data and environmental sensor data in the monitoring environment, and obtain audio features and sensor features based on the audio data and the environmental sensor data; wherein the audio features of the audio data are extracted using a convolutional neural network, specifically: Where M(f, t) represents the audio feature, X k represents the frequency domain signal of the audio data, H(f, k) represents the Mel filter, k represents the index variable, and K represents the number of items involved in the summation; The Transformer model is used to align audio features and sensor features with video features in real-time video data through timestamps to obtain a feature matrix including multiple feature types, and early warning analysis is performed based on the feature matrix.
8. The AI-based hospital security early warning analysis method according to claim 1 is characterized in that: The method further comprises: The feature map of each layer of the network in the deep neural network-based anomaly detection model is obtained, and a heat map in the monitoring environment is generated according to the weight of the feature map to explain the model decision process.
9. The AI-based hospital security early warning analysis method according to claim 8 is characterized in that: The weight of the feature map is specifically: In the formula, α k represents the weight of the kth channel of the feature map, Z represents the normalization processing constant, represents the activation value of the kth channel of the feature map, i and j represent the spatial position of the feature map, and y c represents the anomaly score of behavior category c obtained by the anomaly detection model; Heatmaps used to explain the model's decision-making process, specifically: In the formula, represents the heat map, ReLU is the rectified linear unit, A k represents the activation value of the kth channel, α k Represents the weight of the kth channel of the feature map.
10. An AI-based hospital security early warning analysis system, applied to an AI-based hospital security early warning analysis method according to any one of claims 1 to 9, characterized in that: include: A data acquisition and preprocessing module, used for decoding the acquired real-time video data into continuous image frames, and performing noise reduction processing on the continuous image frames through a Gaussian filter; A state information recognition module is used to perform target detection processing on each frame of the continuous image frames after noise reduction using a preset target detection algorithm, and obtain state information of the corresponding tracking target when there is a tracking target in the image, wherein the state information includes position information and speed information; A prediction matching module, used to obtain the predicted state of each tracking target in each frame image by using the prediction equation of Kalman filtering, match the predicted state of the current frame image with the state information, and obtain a matching result of the corresponding tracking target; A state updating module, for the tracking target whose matching result is a successful match, updates the state of the tracking target based on the predicted state of the current frame and the state information, and obtains target information corresponding to the tracking target; A behavior category recognition module is used to classify the behavior of each tracking target through a preset three-dimensional convolutional neural network based on the target information of each tracking target, and obtain the behavior category of each tracking target; The anomaly detection and warning module is used to calculate the anomaly score of each tracking target based on the behavior category of each tracking target using an anomaly detection model based on the isolation forest detection algorithm, and obtain the warning level of each tracking target according to the anomaly score combined with the corresponding behavior category, wherein the warning level includes low risk, medium risk and high risk.
Citation Information
Cited By
Security monitoring system for intelligent building
CN120186305A
Intelligent construction site safety monitoring and abnormal behavior detection method
CN120526485A