Safety monitoring method and system based on image recognition

Through multi-directional cameras and directional motion deconvolution technology, combined with deep learning algorithms, efficient and accurate security monitoring of bank branches is achieved, solving the problems of dynamic threat adaptability and environmental impact in traditional methods, and improving the security and response efficiency of bank branches.

CN120656123AInactive Publication Date: 2025-09-16SHENZHEN ZIJIN FULCRUM TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510793716.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional bank branch security monitoring methods rely on static rules and manual settings, are unable to adapt to dynamically changing security threats, and are easily affected by environmental changes and lighting conditions, resulting in insufficient detection accuracy and the possibility of missed or false alarms.

Method used

Multi-directional cameras are used to capture bank branch videos, and directional motion deconvolution technology is used to remove motion blur. Combined with multi-target depth detection and image frame segmentation, personalized trajectory feature mining and visual recognition are performed to build a dynamic behavior prediction model and realize intelligent emergency response decision-making.

Benefits of technology

It improves video clarity and target recognition accuracy, reduces misjudgments, and can identify potential threats and abnormal behaviors in real time, automatically issue warnings and execute emergency responses, thereby improving the security and response efficiency of bank branches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656123A_ABST
    Figure CN120656123A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image recognition, in particular to a safety monitoring method and system based on image recognition. The method comprises the following steps: acquiring monitoring videos of a bank from multiple angles, and performing directional motion convolution elimination so as to construct a dynamic fuzzy elimination monitoring video; performing multi-target depth detection and image frame segmentation on the dynamic fuzzy elimination monitoring video, and extracting a plurality of personnel real-time image frames; performing continuous frame target tracking according to the plurality of person real-time image frames, and performing personalized trajectory feature mining to generate a plurality of person personalized trajectory features; and carrying out visual identification and potential threat object judgment on the plurality of person real-time image frames by holding articles one by one, and extracting suspicious person image frames. According to the invention, full-automatic and real-time response safety behavior analysis is realized, and the safety risk of a bank is reduced to the greatest extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a security monitoring method and system based on image recognition. Background Art

[0002] With the rapid development of information technology, the security of bank branches has gradually become a focus of public attention. Traditional bank security measures rely primarily on manual patrols, physical isolation, and conventional surveillance camera systems. However, with the continuous development of banking services, particularly driven by financial technology, bank branch service functions have become increasingly diversified, customer numbers have continued to grow, and transaction activities have become more frequent and complex. Traditional security methods are gradually exposing their limitations. In particular, as bank branches face increasingly complex security threats and potential risks, simple manual monitoring and mechanized security measures are no longer able to meet the requirements for efficient, accurate, and real-time response.

[0003] Against this backdrop, image recognition-based security monitoring technology has become an increasingly important tool for enhancing bank branch security. Compared to traditional security monitoring technologies, image recognition uses intelligent algorithms to analyze video footage, automatically detecting and identifying unusual events and behaviors, providing a more accurate and real-time security monitoring solution. Particularly in bank branch environments, image recognition can efficiently identify suspicious individuals, potential safety hazards, and unusual behavior patterns, helping security managers make timely decisions and respond accordingly.

[0004] However, despite the significant theoretical and technical advantages of image recognition technology, traditional image recognition-based bank branch security monitoring methods still face numerous challenges. First, traditional security monitoring methods often rely on static rules and models, which are unable to adapt to the dynamically changing security threats in bank branch environments. Second, existing monitoring systems are mostly manually configured, lacking intelligence and adaptability, and are unable to monitor and warn of complex security risks in real time. Furthermore, traditional monitoring methods often rely on simple motion detection and facial recognition, which are easily limited by factors such as environmental changes, lighting conditions, and image quality. This can lead to insufficient security detection accuracy and even missed or false positives. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention proposes a security monitoring method and system based on image recognition to solve at least one of the above technical problems.

[0006] To achieve the above object, the present invention provides a security monitoring method based on image recognition, comprising the following steps: Step S1: Obtain monitoring videos of the bank from multiple angles and perform directional motion convolution elimination to construct a monitoring video with motion blur eliminated; Step S2: Perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video to extract multiple real-time image frames of people; Step S3: Continuous frame target tracking is performed based on multiple real-time image frames of personnel, and personalized trajectory feature mining is performed to generate multiple personalized trajectory features of personnel; Step S4: Perform visual recognition of handheld items and potential threat objects on multiple real-time image frames of people one by one, and extract image frames of suspicious people; Step S5: performing deep dynamic behavior prediction on the suspicious person image frame based on multiple personalized trajectory features, and constructing a dynamic behavior trajectory sequence of the suspicious person; Step S6: Analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0007] This invention uses multi-directional cameras to capture dynamic video of bank branches from all angles, avoiding the blind spots and blind spots of traditional single-angle cameras. Video quality is significantly improved, especially in low-light conditions or when objects are moving at high speeds. Directed motion deconvolution technology effectively removes motion blur, ensuring video clarity. By eliminating motion blur, image recognition is enhanced, preventing object detection errors or false alarms caused by motion blur, and ensuring accurate identification of moving objects in rapidly changing environments. Motion blur removal effectively improves the clarity of real-time video, ensuring sufficient detail in surveillance images even in fast-action scenes (such as people entering and exiting quickly or during emergencies), enhancing the real-time responsiveness of the surveillance system. Using multi-target deep detection algorithms (such as YOLO and Faster R-CNN), multiple individuals can be accurately identified based on clear video and their individual image frames can be delineated, ensuring the precise location of each target. Because motion blur is eliminated, detected targets are clearer, enabling faster and more accurate multi-target recognition and image frame segmentation, reducing the need for manual intervention. In complex banking environments, people often interact or overlap. Deep detection combined with precise image frame segmentation ensures individual identification in these scenarios, avoiding overlapping detection and misidentification. By continuously tracking each person's trajectory, behavioral characteristics can be extracted, such as the temporal patterns of their entry, stay, and exit, and their movement paths. These personalized trajectory features provide the basis for subsequent behavioral analysis and anomaly detection. Object tracking algorithms (such as Kalman filtering or deep tracking) ensure the continuity of individuals between video frames, avoiding tracking failures caused by brief occlusions or rapid movement. By mining personalized trajectory features, the system can automatically identify each person's regular behavior patterns, providing comparative data for subsequent anomaly detection, improving the system's accuracy and intelligence. Visual recognition technology accurately identifies objects (such as bags, boxes, electronic devices, and weapons) held by individuals in each image frame, even if they are small or partially obscured, effectively preventing missed threats. Object recognition technology enables the system to identify potential threats, such as firearms and sharp tools, in real time. When a suspicious object is detected, the individual is immediately flagged as a suspect, enhancing the early warning capabilities of security monitoring. This technology helps security personnel quickly identify individuals in possession of dangerous goods, ensuring efficient and timely handling. By using dynamic behavior prediction based on historical trajectory data, the system can proactively identify unusual movements of suspicious individuals before their behavior patterns change. The system can also predict whether a person has violent tendencies or is likely to approach a bank counter to engage in illegal activities, thus providing early warning.By leveraging deep learning models (such as LSTM and RNN) combined with personalized trajectory features, the system can identify subtle behavioral changes that might not normally be apparent, providing early warning of high-risk behavior and minimizing potential risks. The construction of dynamic behavioral trajectory sequences enables the system to infer and analyze the real-time behavior of each suspicious individual, accurately predicting their subsequent actions and enhancing security. By analyzing dynamic behavioral trajectory sequences, the system can identify potentially dangerous actions, such as rapidly approaching staff or approaching the counter. Such abnormal behavior triggers an alarm, enabling security personnel to take timely action. In conjunction with pre-set emergency response strategies, the system not only issues an alarm but also automatically initiates security measures, such as locking access control, initiating video playback, and notifying security personnel, enabling rapid and effective response to security threats. The intelligent emergency response system automatically determines the threat level and implements emergency response plans, reducing security personnel's decision-making time and the probability of error in emergency situations, thereby improving response efficiency. Through real-time data analysis, the system can make immediate decisions and respond to suspicious behavior, minimizing bank security risks and providing decision support to security personnel.

[0008] In this specification, a security monitoring system based on image recognition is provided, which is used to execute the security monitoring method based on image recognition as described above, including: The video optimization module is used to obtain monitoring videos of the bank from multiple angles and perform directional motion deconvolution to construct monitoring videos with motion blur eliminated; Image segmentation module, used to perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video, and extract multiple real-time image frames of people; The target tracking module is used to track targets in consecutive frames based on multiple real-time image frames of people, and to mine personalized trajectory features, thereby generating multiple personalized trajectory features of people; The object visual recognition module is used to visually identify the objects held by each person and identify potential threats in multiple real-time image frames, and extract image frames of suspicious persons; The behavior prediction module is used to perform deep dynamic behavior prediction on the suspicious person's image frame based on multiple personalized trajectory features and construct a dynamic behavior trajectory sequence of the suspicious person; The emergency response decision module is used to analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0009] This invention uses directional motion deconvolution to effectively remove blur caused by rapid motion, which is crucial in dynamic environments like bank branches. Improved video clarity enables more accurate subsequent object detection and behavior analysis. Optimized surveillance video reduces misjudgments due to blur and ensures the fidelity of image details. This allows security personnel to clearly see the details of each individual, increasing the reliability of the surveillance video. Motion blur removal not only improves overall video quality but also enhances the ability to identify fast-moving objects, ensuring the system can accurately capture their behavior even in rapidly moving objects or in poorly lit environments. Deep learning technologies (such as YOLO and Faster R-CNN) can simultaneously identify multiple objects, ensuring real-time monitoring of the dynamic behavior of multiple individuals in bank branches without missing any potential threats. Image frame segmentation effectively prevents misidentification and missed detections in object detection. Each individual's image frame is independent and clear, enhancing the accuracy of subsequent processing. The system can quickly identify all objects in the image and update the image frames in real time, making the entire monitoring process efficient and highly real-time. Bank branches often involve complex personnel flows, and image segmentation technology can effectively improve the response speed of real-time monitoring. Continuous-frame target tracking technology ensures stable tracking of targets across multiple video frames. Even with occlusion or changes in movement, the system maintains consistent target identification, improving tracking accuracy. Each person's trajectory features unique characteristics, such as frequent activity areas, dwell time, and movement speed. Mining these personalized trajectory features provides unique data support for subsequent behavioral analysis, helping to identify potential anomalous behavior. Accurate trajectory tracking lays the foundation for subsequent abnormal behavior detection. The system can identify anomalies in a person's behavior patterns by comparing them with historical trajectory data and issue timely alerts. Using deep learning algorithms to visually identify objects held by individuals, the system can accurately identify potential threats, such as knives, guns, and packages, which is crucial for bank branch security. Identifying and assessing threats from held objects allows for the timely detection of potential danger sources and early warnings. By flagging suspicious individuals, security personnel can take preventative measures and mitigate security risks. The handheld object recognition module overcomes the inability of traditional surveillance methods to identify hidden threats, ensuring that any potential danger is promptly captured. By dynamically predicting human behavior, the system can anticipate suspicious individuals' movements, such as entering restricted areas or engaging in violent behavior, before actual threats occur, providing early warnings and mitigating risks. Dynamic behavior prediction significantly reduces security personnel's response delays. By accurately sequencing behavioral trajectories, security personnel can respond quickly and ensure the safety of bank branches. The behavior prediction module provides data support for optimizing security strategies, enabling the development of more refined security measures, such as adjusting key monitoring areas and strengthening patrols.The system rapidly generates emergency response decisions based on analysis of unusual and dangerous behavior. For example, when threatening behavior is detected, the system automatically triggers an alert and initiates appropriate response procedures based on the threat level (such as locking exits and notifying security). The intelligent nature of the emergency response decision module ensures standardized and automated responses, reduces errors caused by human intervention, and improves the efficiency and accuracy of incident response. Based on behavioral predictions and hazard analysis, the system automatically constructs emergency response strategies tailored to different threat types. This intelligent strategy enables fully automated, real-time response, enabling security teams to quickly make decisions and handle emergencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A schematic flow chart of the steps of a security monitoring method based on image recognition according to the present invention; Figure 2 Detailed implementation flow chart of step S1; Figure 3 Detailed implementation flow chart of step S2; Figure 4 Schematic diagram of the detailed implementation steps of step S3. DETAILED DESCRIPTION

[0011] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0012] This application provides a security monitoring method and system based on image recognition. The execution entities of the security monitoring method and system based on image recognition include, but are not limited to, the following: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. equipped with the system, which can be regarded as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of an audio and image management system, an information management system, and a cloud data management system.

[0013] See also Figures 1 to 4 The present invention provides a security monitoring method based on image recognition, which includes the following steps: Step S1: Obtain monitoring videos of the bank from multiple angles and perform directional motion convolution elimination to construct a monitoring video with motion blur eliminated; Step S2: Perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video to extract multiple real-time image frames of people; Step S3: Continuous frame target tracking is performed based on multiple real-time image frames of personnel, and personalized trajectory feature mining is performed to generate multiple personalized trajectory features of personnel; Step S4: Perform visual recognition of handheld items and potential threat objects on multiple real-time image frames of people one by one, and extract image frames of suspicious people; Step S5: performing deep dynamic behavior prediction on the suspicious person image frame based on multiple personalized trajectory features, and constructing a dynamic behavior trajectory sequence of the suspicious person; Step S6: Analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0014] This invention uses multi-directional cameras to capture dynamic video of bank branches from all angles, avoiding the blind spots and blind spots of traditional single-angle cameras. Video quality is significantly improved, especially in low-light conditions or when objects are moving at high speeds. Directed motion deconvolution technology effectively removes motion blur, ensuring video clarity. By eliminating motion blur, image recognition is enhanced, preventing object detection errors or false alarms caused by motion blur, and ensuring accurate identification of moving objects in rapidly changing environments. Motion blur removal effectively improves the clarity of real-time video, ensuring sufficient detail in surveillance images even in fast-action scenes (such as people entering and exiting quickly or during emergencies), enhancing the real-time responsiveness of the surveillance system. Using multi-target deep detection algorithms (such as YOLO and Faster R-CNN), multiple individuals can be accurately identified based on clear video and their individual image frames can be delineated, ensuring the precise location of each target. Because motion blur is eliminated, detected targets are clearer, enabling faster and more accurate multi-target recognition and image frame segmentation, reducing the need for manual intervention. In complex banking environments, people often interact or overlap. Deep detection combined with precise image frame segmentation ensures individual identification in these scenarios, avoiding overlapping detection and misidentification. By continuously tracking each person's trajectory, behavioral characteristics can be extracted, such as the temporal patterns of their entry, stay, and exit, and their movement paths. These personalized trajectory features provide the basis for subsequent behavioral analysis and anomaly detection. Object tracking algorithms (such as Kalman filtering or deep tracking) ensure the continuity of individuals between video frames, avoiding tracking failures caused by brief occlusions or rapid movement. By mining personalized trajectory features, the system can automatically identify each person's regular behavior patterns, providing comparative data for subsequent anomaly detection, improving the system's accuracy and intelligence. Visual recognition technology accurately identifies objects (such as bags, boxes, electronic devices, and weapons) held by individuals in each image frame, even if they are small or partially obscured, effectively preventing missed threats. Object recognition technology enables the system to identify potential threats, such as firearms and sharp tools, in real time. When a suspicious object is detected, the individual is immediately flagged as a suspect, enhancing the early warning capabilities of security monitoring. This technology helps security personnel quickly identify individuals in possession of dangerous goods, ensuring efficient and timely handling. By using dynamic behavior prediction based on historical trajectory data, the system can proactively identify unusual movements of suspicious individuals before their behavior patterns change. The system can also predict whether a person has violent tendencies or is likely to approach a bank counter to engage in illegal activities, thus providing early warning.By leveraging deep learning models (such as LSTM and RNN) combined with personalized trajectory features, the system can identify subtle behavioral changes that might not normally be apparent, providing early warning of high-risk behavior and minimizing potential risks. The construction of dynamic behavioral trajectory sequences enables the system to infer and analyze the real-time behavior of each suspicious individual, accurately predicting their subsequent actions and enhancing security. By analyzing dynamic behavioral trajectory sequences, the system can identify potentially dangerous actions, such as rapidly approaching staff or approaching the counter. Such abnormal behavior triggers an alarm, enabling security personnel to take timely action. In conjunction with pre-set emergency response strategies, the system not only issues an alarm but also automatically initiates security measures, such as locking access control, initiating video playback, and notifying security personnel, enabling rapid and effective response to security threats. The intelligent emergency response system automatically determines the threat level and implements emergency response plans, reducing security personnel's decision-making time and the probability of error in emergency situations, thereby improving response efficiency. Through real-time data analysis, the system can make immediate decisions and respond to suspicious behavior, minimizing bank security risks and providing decision support to security personnel.

[0015] In the embodiment of the present invention, see Figure 1 , is a schematic flow chart of the steps of a security monitoring method based on image recognition of the present invention. In this example, the steps of the security monitoring method based on image recognition include: Step S1: Obtain monitoring videos of the bank from multiple angles and perform directional motion convolution elimination to construct a monitoring video with motion blur eliminated; In this embodiment, after obtaining relevant authorization, multiple cameras are installed at key locations within a bank branch (such as entrances, counters, and service areas) to ensure coverage of key areas from every angle. The cameras should have high resolution (e.g., 1080p or higher) and good low-light performance to ensure clear video capture under all lighting conditions. The camera parameters are set to: 1920x1080 resolution, 30fps frame rate, and a 90-degree viewing angle to ensure coverage of the entire monitoring area. All cameras are activated and begin recording monitoring video. The recording period is set to 24 hours to ensure that all dynamic activities are captured. Video data is stored in real time on a secure server to prevent data loss. Each camera is set to generate approximately 10GB of video data per hour, using the H.264 encoding format to balance video quality and storage space. Recorded videos are categorized and tagged to ensure that each video file clearly identifies its source (e.g., camera location, time period, etc.), facilitating subsequent processing and analysis. The file name format is "camera1_20230401_0001.mp4" to facilitate accurate identification in subsequent steps. Perform motion blur analysis on the captured surveillance video. First, identify the motion trajectory of the moving objects in the video. Optical flow or frame difference methods can be used to calculate the direction and speed of each moving object. Set the optical flow window size to 15x15 pixels. Analyze the pixel changes of the moving object in each frame to extract its speed and direction. Based on the analyzed motion trajectory, design an appropriate convolution kernel. The direction of the convolution kernel should align with the direction of the moving object's motion for effective deblurring. Commonly used convolution kernels are one-dimensional Gaussian filters or directional filters. Set the convolution kernel size to 5x5, aligning its direction with the average direction of the object's motion. Adjust the kernel strength based on the speed of the object to accommodate varying degrees of blur. Apply the designed convolution kernel to each frame to perform directional motion deconvolution. This convolution operation reduces motion blur in the image and restores details of the moving object. The processed video should display a clearer image of the moving object. During processing, the stride of the convolution operation is set to 1 pixel, and the processing time per frame should be controlled within 100 milliseconds to ensure real-time performance. The processed motion blur removal monitoring video is quality-assessed using metrics such as PSNR (peak signal-to-noise ratio) and SSIM (structural similarity index) to ensure that the video quality meets the expected standards. If the evaluation results are unsatisfactory, the convolution kernel parameters should be adjusted and the processing should be repeated. The PSNR value should be set to above 30dB and the SSIM value should be set to above 0.85 to ensure the video's clarity and detail restoration. Each frame processed by the directional motion deconvolution is synthesized into a new motion blur removal monitoring video to ensure a smooth and natural video. The frame rate per second must be maintained during synthesis to avoid video playback lags.Maintain a 30fps video frame rate, ensuring 30 frames per second. The resulting video should be smooth and free of frame skipping. Store the resulting motion blur removal monitoring video on a secure server and tag it for easy access and analysis. Ensure the video's storage format and encoding method are consistent with the original video for ease of processing. The storage format should be MP4, with a naming format like "dynamic_blur_removed_20230401.mp4." After preparing the motion blur removal monitoring video, prepare it for subsequent behavioral analysis and target detection. A quick preview can be performed using video analysis software to confirm that the video quality meets analysis requirements. Confirm the clarity and recognizability of dynamic targets in the video to ensure smooth subsequent behavioral detection and analysis.

[0016] Step S2: Perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video to extract multiple real-time image frames of people; In this example, a deep learning model suitable for real-time multi-target detection, such as YOLOv5, Faster R-CNN, or SSD, was selected. These models can handle simultaneous multi-target detection in complex scenes and achieve a good balance between accuracy and speed. The YOLOv5 model was selected due to its fast inference speed and high detection accuracy, making it suitable for use in motion blur reduction surveillance video. If the existing model performs poorly in a specific scenario, the model can be fine-tuned using a labeled dataset containing bank branch locations. This can improve the model's robustness through data augmentation techniques (such as image flipping, rotation, and scaling). Fine-tuning is performed using a training set of 5,000 labeled images to ensure the model can accurately identify people and objects unique to bank branches. The trained model is loaded and the relevant parameters are configured to ensure the model is suitable for real-time video stream processing. A detection confidence threshold (e.g., 0.5) is set to filter out low-confidence detection results and reduce false detections. The model input size is set to 640x640 pixels to balance detection accuracy and processing speed. Read images frame by frame from motion-blur-removed surveillance video to perform depth detection on each frame. Use an efficient video reading library to ensure the reading speed matches the video frame rate (e.g., 30 fps). Set the processing time for each frame to no more than 33 milliseconds to achieve real-time detection and analysis. Preprocess each frame, including image resizing, normalization, and color space conversion. Ensure that the input data meets the model requirements to improve detection performance. Resize the image to 640x640 pixels and normalize pixel values ​​to the range [0, 1] to accommodate the YOLOv5 model input format. Input each preprocessed frame into the depth detection model to perform multi-object detection. The model identifies all dynamic objects (e.g., people) in each frame and returns detection results, including bounding box coordinates and confidence scores. If three people are detected in a frame, the result returned will be "Person 1: Bounding box (100, 150, 200, 300), confidence: 0.92." The detection results output by the model are post-processed, including non-maximum suppression (NMS) to eliminate duplicate detections and retain detection results with high confidence. Set the NMS threshold (such as 0.4) to ensure the accuracy of the detection results. If the overlap of the detected person bounding boxes exceeds 0.4, only the bounding box with the highest confidence is retained. The bounding box information of the person detected in each frame is stored in a data structure, including the person ID, bounding box coordinates, and confidence level. At the same time, the bounding box is drawn on the image for visualization for subsequent analysis. Use a red box to draw the bounding box of each person and mark the corresponding confidence level in the box to form an intuitive detection result. Perform image frame segmentation on each person based on the detected bounding box coordinates. Extract the area within each bounding box as a separate image for subsequent processing and analysis.If the bounding box for person 1 is (100, 150, 200, 300), crop that region from the original image to create a new image named "person_1.jpg." Store each extracted person image frame sequentially in a designated folder for easy access and analysis. Ensure consistent naming for efficient management. Set the storage path to " / extracted_frames / " and the file name format to "person_ID.jpg," such as "person_1.jpg" and "person_2.jpg."

[0017] Step S3: Continuous frame target tracking is performed based on multiple real-time image frames of personnel, and personalized trajectory feature mining is performed to generate multiple personalized trajectory features of personnel; In this embodiment, algorithms suitable for real-time target tracking are selected, such as Kalman filtering, CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability), or SORT (Simple Online and Realtime Tracking). These algorithms can effectively handle fast-moving and partially occluded targets. CSRT is chosen because it maintains high accuracy and effectively handles changes in target appearance when tracking targets in dynamic scenes. The tracker is initialized on each detected person image frame, ensuring that each target has a separate tracking instance. For each target, its initial position and ID are recorded for subsequent tracking. Each target ID is assigned a value of 1, 2, 3, and so on. During initialization, the bounding box coordinates of each target (e.g., (100, 150, 200, 300)) are recorded for subsequent tracking. Images are read frame by frame from the motion-blur-removed surveillance video to track the target for each frame. The read speed is kept consistent with the video frame rate (e.g., 30 fps) to achieve real-time processing. The processing time per frame is set to no more than 33 milliseconds to keep up with the video playback speed. For each frame, a tracking algorithm updates the position of each object. Based on the initial bounding box and motion model, the new position of each object is calculated and its trajectory is tracked. During processing, if the bounding box of an object in frame n is updated to (110, 160, 210, 310), the position information for that frame is recorded. A trajectory data structure is created for each object, recording the coordinates and timestamp of each object in each frame of the video. This forms a complete trajectory sequence for subsequent analysis. The trajectory data format is "Person ID: 1, Trajectory: [(t1, x1, y1), (t2, x2, y2),...]", where t is the timestamp and x and y are the corresponding coordinates. In the event of occlusion, the tracking algorithm should be able to handle temporary disappearance of the object. If the object is not detected in consecutive frames, a threshold is used to determine whether it reappears and re-identification is performed. The occlusion detection threshold is set to 3 frames. If an object is not detected in three consecutive frames, it is considered to be possibly occluded and the target is further tested for reappearance. Determine the content of personalized trajectory features, including entry and exit frequency, dwell time, movement speed, and behavioral patterns. These features should reflect the unique behavior of each target. Set the feature extraction window to 5 minutes and calculate the number of entries and exits and dwell time within this time period. Based on the recorded trajectory data, calculate the personalized features of each target. Analyze each target's dwell time within the monitoring area, frequently entered and exited areas, and overall movement speed.If a target entered an area three times in the past five minutes, spending 10 seconds, 20 seconds, and 5 seconds in each area, respectively, its characteristics could be recorded as "Area A: 3 times, average dwell time: 11.67 seconds." The calculated personalized trajectory characteristics are aggregated and stored in a data structure for subsequent analysis and decision support. Ensure that each target's characteristic information is clearly visible. The recording format is "Personnel ID: 1, Personalized Characteristics: Dwell Time: 11.67 seconds, Entry and Exit Frequency: 0.6 times / minute." Visualize each target's trajectory data to show the target's movement path within the monitored area. This can be achieved by drawing a trajectory map, helping analysts quickly understand the target's behavior. Use different colors to identify the trajectory of different targets, and plot the target's movement path to facilitate subsequent behavioral analysis.

[0018] Step S4: Perform visual recognition of handheld items and potential threat objects on multiple real-time image frames of people one by one, and extract image frames of suspicious people; In this example, a deep learning model suitable for handheld object recognition, such as YOLOv5 or Faster R-CNN, is selected. These models can detect and classify objects in images in real time and are suitable for identifying a variety of handheld items, including dangerous items. The YOLOv5 model is used because it offers fast inference speed and high accuracy when processing real-time video streams, adapting to the dynamic environment of bank branches. If the existing model's recognition performance is insufficient, a specialized handheld object dataset can be used to fine-tune the model. This dataset should include images of various handheld objects, annotated and categorized, to improve the model's recognition capabilities. Training is performed using 3,000 annotated images of handheld objects to ensure the model can accurately recognize common handheld objects, such as bags, knives, and bottles. The trained model is loaded and the relevant parameters are configured. A detection threshold (e.g., 0.5) is set to filter out low-confidence recognition results. This helps reduce false positives and ensures that only high-confidence detection results are considered. The input image size is set to 640x640 pixels to ensure the model's efficiency and accuracy when processing video streams. Read the image frames of a person frame by frame from the motion-blur-removed surveillance video to perform visual identification of handheld objects in each frame. Ensure that the reading speed matches the video frame rate (e.g., 30 fps). Set the processing time per frame to no more than 33 milliseconds to achieve real-time detection and maintain video smoothness. Preprocess each frame, including image resizing, normalization, and color space conversion, to ensure the input data meets the model requirements. The image is resized to 640x640 pixels and pixel values ​​are normalized to the range [0, 1] to accommodate the YOLOv5 model input format. The preprocessed image is fed into the deep learning model to perform visual identification of handheld objects. The model detects all handheld objects in each frame and returns bounding box coordinates and confidence scores. If the object detected as a "knife" in a frame is detected as a handheld object, its bounding box coordinates are (150, 200, 250, 300) and the confidence score is 0.92. The criteria for defining potential threat objects usually include offensive or dangerous items such as knives, guns, and explosives. These items should be compared with the preset threat object library through identification results. "Knife", "gun", and "explosives" are marked as high-risk items to facilitate subsequent judgment and processing. The recognition results of the model are compared with the threat object standards to determine whether the identified object is a potential threat object. If the detected object is in the threat object library, it is marked as a potential threat. If the identification result is "knife", the object is marked as a potential threat and the relevant information is recorded for subsequent processing. The relevant information of the potential threat object is recorded in the data structure, including the person ID, item category, threat level, and bounding box coordinates. Ensure that the threat object information of each suspicious person is clearly visible.The record format is "Person ID: 1, Potential Threat Item: Knife, Threat Level: High, Bounding Box: (150, 200, 250, 300)." Once a potential threat object is detected, the suspicious person is marked and their image frame is visualized. This is highlighted with a red frame on the monitoring interface, and the object category and confidence level information are added within the frame. If person ID 1 is holding a knife, a red border is drawn around their image frame, and the box is labeled "Knife (Confidence: 0.92)." The marked suspicious person's image frame is extracted for further analysis and processing. The suspicious person's image frame is cropped and stored for subsequent monitoring and investigation. If the suspicious person's bounding box is (150, 200, 250, 300), this area is extracted as a new image file named "suspicious person_1.jpg."

[0019] Step S5: performing deep dynamic behavior prediction on the suspicious person image frame based on multiple personalized trajectory features, and constructing a dynamic behavior trajectory sequence of the suspicious person; In this embodiment, based on previous analysis and records, personalized trajectory features are extracted for each suspicious individual. These features should include entry and exit frequency, dwell time, movement speed, and behavioral patterns, providing basic data for subsequent dynamic behavior prediction. The feature extraction window is set to 5 minutes, and the number of entries and exits and dwell time within this time period are calculated. If a person dwells in a specific area for 20 seconds and has an entry and exit frequency of 0.5 times / minute over the past 5 minutes, these data are recorded. The extracted personalized trajectory features are stored in a data structure for subsequent use. Ensure that the characteristic information of each suspicious individual is clearly visible to facilitate dynamic behavior prediction. The record format is "Personnel ID: 1, Personalized Features: Dwell time: 20 seconds, Entry and Exit Frequency: 0.5 times / minute, Movement Speed: 1.2 m / s." Select a deep learning model suitable for dynamic behavior prediction, such as an LSTM (Long Short-Term Memory) network or a GRU (Gated Recurrent Unit). These models can process time series data and are suitable for analyzing personalized trajectory features and predicting future behavior. The LSTM model was chosen because it excels at processing long sequences of data and can capture the temporal dependencies of dynamic behavior. The model was trained using existing trajectory feature data to enhance its ability to predict suspicious individual behavior. Data augmentation techniques (such as time series perturbation) can be used to expand the training set and improve the model's generalization. Training was performed on 1,000 trajectory data sets, with a training period of 50 epochs. The cross-entropy loss function was used to optimize model performance. After training, the model was evaluated on a validation set to ensure its predictive accuracy and reliability. Metrics such as mean squared error (MSE) or accuracy can be used for evaluation. The MSE on the validation set was set to below 0.05 to ensure the model's effectiveness in practical applications. The personalized trajectory features were organized into a time series format and prepared for input into the deep learning model for prediction. Ensure that the data format meets the model's input requirements. This typically involves normalizing the feature data. The feature data was normalized to the range [0, 1] to accommodate the LSTM model input. The organized data was then input into the trained LSTM model to perform dynamic behavior prediction. The model outputs a sequence of future behavioral trajectories of suspicious individuals, predicting their future locations and likely behavior types. The model might predict a suspicious individual's trajectory over the next five seconds as [(x1', y1'), (x2', y2'), (x3', y3')], representing their future movement path. The predicted dynamic trajectory sequence is recorded in a data structure for subsequent analysis and emergency response. Ensure that the predicted trajectory information for each suspicious individual is clearly visible. The record format is "Person ID: 1, Predicted trajectory: [(x1', y1'), (x2', y2'), (x3', y3')], Timestamp: [t1, t2, t3]."Visualizing the predicted dynamic behavior trajectory sequence to display the expected movement path of suspicious individuals within the monitoring area can help security personnel quickly understand the suspicious individual's future behavior. Predicted trajectories are identified with different colors and superimposed on the monitoring screen for real-time monitoring. Based on the predicted dynamic behavior trajectory, combined with real-time monitoring data, intelligent decision-making support is provided for security management. If the prediction results indicate that a suspicious individual may enter a high-risk area, security personnel should be notified in a timely manner to intervene. If a suspicious individual is predicted to enter the cash storage area within the next 5 seconds, the system should automatically generate an early warning to alert security personnel.

[0020] Step S6: Analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0021] In this embodiment, first, criteria for abnormal and dangerous behavior are defined. These typically include rapid running, frequent looking back, sudden stops, and unusual movement routes. These behaviors should be summarized and summarized through analysis of historical data and expert experience. "Rapid running" is defined as moving more than 5 meters in 1 second, while "frequent looking back" is defined as looking back more than 3 times in 10 seconds. Based on the dynamic trajectory sequence of a suspicious individual, their movement patterns are analyzed in real time and compared with the defined abnormal behavior criteria. This can be achieved by setting up a real-time monitoring algorithm. Thresholds for movement speed and direction changes are used to detect whether a suspicious individual exhibits abnormal acceleration or sharp turns within a short period of time. Each suspicious individual's dynamic behavior is monitored using the set criteria, and the occurrence of abnormal movements is recorded. If a suspicious individual's behavior is detected to meet the abnormal criteria, it is marked as abnormal and dangerous. If a target's movement speed exceeds a set threshold (e.g., 3 meters / second) during the monitoring period and is accompanied by frequent looking back, it is marked as "dangerous behavior." Based on the detected abnormal behavior, an emergency response mechanism is designed. This mechanism should include multiple measures such as immediate monitoring of the suspicious individual, rapid dispatch of security personnel, and activation of the alarm system. When unusual behavior is detected, the system automatically triggers an alarm and records the incident information, ensuring security personnel can arrive at the scene quickly. Once unusual movement is detected, the system should immediately initiate real-time monitoring of the suspicious individual and continuously track their movements. Simultaneously, the system notifies nearby security personnel via the communication system to quickly intervene. If a suspicious individual is detected as "running" within 5 seconds, two security personnel are automatically notified to arrive at the scene, ensuring a timely response to potential threats. Detailed information on each response is recorded in a database, including the event time, suspicious individual ID, abnormal behavior detected, response measures, and security personnel arrival time, for subsequent analysis and improvement. The record format is "Event ID: 1, Time: t1, Person ID: 1, Abnormal Behavior: Running, Response Time: t2." Historical abnormal behavior data and corresponding emergency response results are collected and analyzed to build intelligent decision-making models. Machine learning algorithms (such as random forests and support vector machines) can be used to identify effective response strategies. Analyze the response results of the past 100 abnormal events to evaluate the effectiveness of different response measures and optimize future emergency response strategies. Regularly evaluate the effectiveness of your emergency response strategy. Optimize existing strategies by comparing response times, outcomes, and feedback from security personnel. Ensure that response measures are continuously adapted to emerging risks and threats. Conduct quarterly strategy reviews to gather feedback from security personnel and adjust strategies to enhance overall emergency response capabilities.

[0022] In this embodiment, refer to Figure 2 , is a flowchart of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Obtain monitoring videos from multiple angles of the bank based on multi-directional cameras at bank branches; Performing a network point ambient light calculation on the monitoring video to identify the ambient light intensity of multiple areas; Identify dark areas according to the ambient light intensity and extract the dark areas; Performing gamma adaptive correction on the dark area to obtain brightness optimized monitoring videos at multiple angles; Perform global time synchronization on brightness optimization monitoring videos from multiple angles to obtain time-synchronized monitoring videos; Directional motion convolution is performed on the time-synchronized monitoring video to construct a dynamic blur-removed monitoring video.

[0023] In this embodiment, after obtaining relevant authorization, multiple high-resolution cameras are installed inside and around bank branches to ensure coverage of all key areas, including entrances, counters, and ATMs. The cameras should have night vision capabilities to provide clear video surveillance in low-light environments. 1080p resolution cameras are selected, set to 30 frames per second (fps) for real-time recording, ensuring smooth video and clear details. Video data from each camera is collected in real time by the monitoring system and stored on a secure server. This ensures the integrity and traceability of the video data and facilitates subsequent analysis and processing. The video storage format is set to H.264 to conserve storage space while maintaining high image quality. Each camera is expected to generate approximately 500MB of data per day. Ambient light intensity is calculated for each monitored video. Image processing algorithms are used to extract brightness information from video frames and calculate the average light intensity for each area. The brightness value is extracted from each video frame and converted to lux using a light intensity formula. The average light intensity for a particular area is calculated to be 200 lux. Based on this calculated light intensity, dark areas are identified. Define a threshold for dark areas (e.g., below 300 lux) and classify each area. If the illumination intensity in an area is 250 lux and the threshold is set to 300 lux, the area is marked as dark. Extract the identified dark areas from the surveillance video and generate corresponding image segments for subsequent image processing and correction. Use an image segmentation algorithm (such as a threshold-based method) to extract the dark areas in each frame, forming a new image set. For dark areas, establish an adaptive gamma correction model to increase the brightness of these areas. The gamma correction formula is: Output brightness = Input brightness ^ γ, where γ is the correction factor (typically between 1.0 and 2.5). For dark areas, select a γ value of 2.2 for brightness correction. Apply the gamma correction model to each extracted dark area, adjusting the brightness pixel by pixel to ensure that the corrected image brightness is at the expected level. If the original brightness of a pixel is 50 cd / m², after gamma correction, the new brightness value can be calculated as 50^2.2, and the output brightness may reach 100 cd / m². The corrected dark areas are merged back into the original surveillance video to generate a brightness-optimized surveillance video, ensuring overall video consistency and clarity. If the image after the dark area correction is a new video frame, it replaces the original video frame to maintain timeline continuity. Timestamps are extracted from video streams from different cameras, and the time difference between each video frame is analyzed for global time synchronization. A base frame (such as the first frame of the first camera) is set as a reference, and the time deviation of other cameras relative to this base frame is calculated. A time synchronization algorithm (such as linear interpolation) is used to adjust the playback speed or delay of videos from different cameras to ensure that all frames at the same time point are displayed synchronously.If the second camera's time delay is 0.5 seconds, the first 0.5 seconds of its video stream are delayed to ensure synchronization with the first camera. The time-synchronized video streams are merged to generate a global time-synchronized monitoring video, ensuring temporal consistency and coherence. The frame count of the final video should match that of the baseline video, ensuring temporal coordination across all video segments. Motion blur analysis is performed on the time-synchronized monitoring video to identify blur caused by motion. Image processing techniques such as edge detection and motion estimation can be used. Optical flow is used to detect areas of significant motion in the video and identify frames with high blur. A directed motion deconvolution algorithm is applied to repair the identified blurry areas. This algorithm restores clarity by performing deconvolution on the image. The size and orientation of the convolution kernel are set to match the direction of motion, processing the blurred image and restoring details. The clean frames are then merged back into the time-synchronized monitoring video to generate the final motion-deblurred monitoring video, ensuring optimal video quality. In the final video, details in each frame are clearly visible, and motion trajectories are smooth, improving overall monitoring effectiveness.

[0024] In this embodiment, the specific steps of performing global time synchronization on the brightness optimization monitoring videos of multiple angles to obtain the time-synchronized monitoring videos are as follows: Marking a same scene feature point in the brightness optimization monitoring videos of the multiple angles; Performing panoramic spatial registration processing on the monitoring videos of the multiple angles based on the same scene feature points to construct a panoramic monitoring video; Perform perspective fusion boundary recognition on panoramic monitoring videos and extract perspective fusion boundary lines; Smooth transition optimization is performed on the perspective fusion boundary line to obtain a panoramic monitoring video with smoothed and optimized boundaries; Calculating the acquisition frequency timestamp of the multi-directional camera; Global time synchronization is performed on the boundary smoothing optimized panoramic monitoring video according to the timestamp to obtain a time-synchronized monitoring video.

[0025] In this embodiment, a feature point extraction algorithm (such as SIFT, SURF, or ORB) is used to identify key feature points within a scene in brightness-optimized monitoring videos from multiple angles. These feature points should exhibit high repeatability and stability to facilitate matching across different viewpoints. Assuming 500 feature points are extracted from a given scene, it is important to ensure that these feature points can be effectively identified and matched across images captured by different cameras. The extracted feature points are labeled and matched within each video frame. A feature matching algorithm (such as FLANN or BFMatcher) is used to ensure that identical feature points are found across different video frames. For example, if 200 feature points are found in the first-view video and 180 feature points are found in the second-view video, the matching algorithm identifies 150 of these feature points as identical. The coordinates of the matched feature points and their corresponding timestamps are recorded to form a database containing feature point location information for subsequent panoramic spatial registration. The recording format can be "feature point ID: coordinates (x, y), video 1 timestamp, video 2 timestamp" to correlate feature points from different viewpoints. Select an appropriate panoramic spatial registration algorithm (such as RANSAC or ICP) and use the marked feature points for spatial registration. This process aims to align video data from different viewpoints to form a consistent panoramic view. Set the inlier threshold of the RANSAC algorithm to 3 pixels to ensure registration accuracy. Perform panoramic spatial registration based on feature points from the same scene. By calculating the transformation matrix between feature points, all video data is mapped into a unified coordinate system. If the matching error of a feature point is greater than 3 pixels during the registration process, it is excluded from the registration calculation to ensure the accuracy of the registration result. Once spatial registration is complete, merge all video data according to the unified coordinate system to generate a seamless panoramic monitoring video. Ensure that image overlap and information loss are avoided during the merging process. The resulting panoramic video resolution is set to 3840x2160 (4K) to ensure clarity and detail. Boundary identification is performed on the panoramic monitoring video, and edge detection algorithms (such as Canny edge detection) are used to extract view fusion boundaries. These boundaries represent the transition areas between different viewpoints and require processing to reduce visual abruptness. Set the high threshold of Canny edge detection to 100 and the low threshold to 50 to ensure the clarity of the boundary lines. Record the extracted boundary lines and analyze their location, length, shape, and other characteristics to provide data support for subsequent smooth transition optimization. Record the starting and ending coordinates of each boundary line, as well as its proportion in the panoramic video. Select an appropriate smoothing algorithm (such as Gaussian smoothing or Bézier curve smoothing) to optimize the perspective fusion boundary lines to achieve a natural transition. Ensure that the brightness changes in the boundary area are smooth and do not affect the overall effect of the panoramic video. Set the standard deviation of Gaussian smoothing to 1.5 to ensure a natural smooth transition effect.A smoothing algorithm is applied to each boundary line, reducing sudden brightness changes at the boundary and creating a more natural transition between different viewpoints. If a boundary line exhibits significant brightness variations, smoothing is performed to adjust the brightness of the boundary area to match the surrounding environment. The smoothed and optimized boundary lines are then reintegrated into the panoramic surveillance video, ensuring that the optimization process does not affect the overall video coherence and clarity. The resulting optimized video uses a 3840x2160 resolution to maintain high quality when displayed on a large screen. The acquisition frequency of multiple cameras is calculated, and timestamps are extracted from each video stream. The differences between timestamps are analyzed for global time synchronization. Each camera's acquisition frequency is set to 30 fps, resulting in 30 timestamps generated per second. A global time synchronization algorithm (such as linear interpolation or clock synchronization) is used to adjust the timing of each video stream to ensure that all videos are played synchronously at the same time. If the timestamp deviation of a camera is 0.1 seconds, the time of its video stream is delayed or advanced by 0.1 seconds to achieve synchronization. The synchronized panoramic surveillance videos are merged to generate the final time-synchronized surveillance video, ensuring temporal consistency and coherence. The number of frames in the final video should be consistent with the baseline video to ensure that all video clips are coordinated in time and improve the overall monitoring effect.

[0026] In this embodiment, the specific steps of performing directional motion convolution elimination on the time-synchronized monitoring video to construct a dynamic blur elimination monitoring video are: Perform time-sequential frame decomposition on the time-synchronized monitoring video and extract all image frames; Calculating a time difference between adjacent image frames of the image frame; Calculating the total inter-frame delay according to the time difference; Calculating a delay average of all inter-frame delays; Adaptively adjusting the frame rate according to the delay average value to obtain an adaptive frame rate; Performing global frame delay optimization on the time-synchronized monitoring video based on the adaptive frame rate to obtain a global delay-optimized monitoring video; Perform dynamic target motion blur analysis on the global delay optimization monitoring video and extract dynamic target motion blur data; Motion vector estimation is performed based on the motion blur data of the dynamic target to obtain the motion trajectory and speed of the dynamic target; Directional motion convolution is performed based on the motion trajectory and speed, thereby constructing a dynamic blur elimination monitoring video.

[0027] In this embodiment, all image frames are extracted from a time-synchronized surveillance video using a video processing tool or library (such as OpenCV). Each extracted frame ensures that it reflects key moments in the video for subsequent analysis. The extraction frequency is set to 30 frames per second. For a 60-second video, a total of 1,800 frames are extracted. The resolution of each frame is maintained at 1920x1080 to ensure image quality. The extracted frames are stored in a structured folder in chronological order, using names such as "frame_001.jpg" and "frame_002.jpg," for easy subsequent processing. The storage path is set to " / video_frames / " for quick access and processing. As each frame is extracted, its timestamp is recorded. This can be obtained from the video's metadata to ensure accurate time information for each frame. Assuming the timestamp of frame 1 is 00:00:01.000 and the timestamp of frame 2 is 00:00:01.033, the time difference between the two frames is 33 milliseconds. Calculate the time difference between all adjacent image frames, using simple subtraction to determine the time difference between each pair of frames. Store the results in an array or list for subsequent analysis. If there are 10 frames with timestamps of 0.000s, 0.033s, 0.067s, 0.100s, 0.133s, 0.167s, 0.200s, 0.233s, 0.267s, and 0.300s, the time differences between the adjacent frames are 33ms, 34ms, 33ms, and so on. Statistically calculate the total delay between all frames. This statistical result will provide the basis for subsequent average delay calculations. If the array of time differences for 10 frames is [33, 34, 33, 33, 34, 33, 34, 33, 34] milliseconds, the total delay is 330 milliseconds. Calculate the average inter-frame delay and divide the total delay by the number of frames to obtain the average delay value. This value will guide subsequent adaptive frame rate adjustments. If the total delay is 330 milliseconds and the number of frames is 9, the average delay is 330ms / 9 = 36.67ms. Based on the calculated average delay, establish an adaptive frame rate adjustment model. This model should account for the impact of inter-frame delay on video playback smoothness. Set the base frame rate to 30fps. If the average delay is 36.67ms, the new frame rate can be calculated as follows: New frame rate = 1000ms / (Average delay + 33.33ms) to ensure a reasonable display time for each frame. Adjust the video playback frame rate based on the adaptive model. This can be achieved by modifying video playback parameters or resampling frames to ensure smooth video playback. If the new frame rate is 28fps, adjust the video to 28 frames per second to maintain a good viewing experience despite delays. Design a global frame delay optimization strategy to optimize the time-synchronized monitoring video based on the adaptive frame rate.The focus is on minimizing the impact of latency on video quality. The latency optimization goal is to reduce overall latency to less than 30ms to ensure an unimpaired user experience. Latency optimization strategies are applied to each frame, making necessary adjustments, which may include reordering, dropping, or inserting frames, to ensure overall video smoothness. Frames with latency exceeding 30ms are processed and their display time appropriately reduced to control overall latency. Motion blur analysis of dynamic targets is performed on the global latency optimization monitoring video. Image processing methods (such as motion estimation and optical flow) are used to identify and extract motion blur data for dynamic targets. The optical flow calculation window size is set to 5x5 pixels to ensure detailed analysis of the blur level of moving targets. During the analysis, the speed and direction of motion of the dynamic targets are extracted, which provides the basis for subsequent motion vector estimation. If a blurred trajectory of a dynamic target is detected, its speed is estimated to be 2 meters per second, with a southeast direction. A motion vector estimation model is developed based on the extracted dynamic target motion blur data. This model should accurately reflect the target's trajectory and speed. A simple linear regression model is used to predict the target's trajectory based on the blurred data. Using a motion vector estimation model, calculate the trajectory and velocity of dynamic targets. Record the estimated results for subsequent processing. If the displacement of a dynamic target between 10 frames is 5 meters, its velocity is estimated to be 5 meters / (10 frames / 30 fps) ≈ 15 meters / second. Select an appropriate directional motion deconvolution algorithm to deblur dynamic targets. Common methods include deconvolution and motion compensation. Setting the deconvolution kernel direction to align with the direction of motion of the dynamic target for more effective deblur removal. Apply the motion deconvolution algorithm to dynamic targets in the global delay-optimized monitoring video, processing each frame to restore clarity and detail. If a frame shows a high degree of dynamic target blur, deconvolution is applied to that frame to restore clarity. The deblurred frames are recombined to generate a motion-deblurred monitoring video, ensuring optimal video quality. The final video should display clear dynamic target trajectories and details, improving overall monitoring effectiveness.

[0028] In this embodiment, refer to Figure 3 , is a flowchart of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Perform multi-target depth detection on the surveillance video with motion blur removal, marking all personnel within the network site; Locate the spatial position of all personnel in the network one by one and obtain the real-time spatial coordinates of each person; Calculating the time when the person entered the network point to obtain initial time information of entering the network point; The dynamic blur elimination monitoring video is segmented into image frames according to the real-time spatial coordinates of each person, and timestamps are marked based on the initial entry time information to extract multiple real-time image frames of the person.

[0029] In this example, a deep learning algorithm (such as YOLO, SSD, or Faster R-CNN) is used for multi-target detection to identify and label all individuals within a bank branch. A model suitable for real-time detection is selected to ensure both speed and accuracy. The YOLOv5 model is chosen, offering a good balance between speed and accuracy, making it suitable for real-time video stream processing. If existing models do not meet requirements, the model can be fine-tuned using an annotated dataset containing bank branch scenes to improve its recognition capabilities in specific environments. Enhancement techniques (such as image rotation and scaling) are used during training to enhance model robustness. 2,000 annotated images containing individuals are collected and trained to ensure performance in real-world environments. The trained deep learning detection model is applied to motion blur removal surveillance video. The video stream is processed frame by frame, detecting and labeling individuals in each frame. Detection results include the bounding box location of each individual and their confidence score. If five individuals are detected in a frame, the bounding box coordinates (such as the top-left and bottom-right corners) and confidence score for each individual are recorded to ensure accurate identification. Based on each detected person's bounding box, calculate their real-time spatial coordinates in the video. This step typically involves converting the pixel coordinates of the video frame into coordinates in the actual scene, which may require the camera's intrinsic and extrinsic parameters. Assuming a camera focal length of 1000 pixels, use the camera model to convert the pixel coordinates in the bounding box (e.g., the upper left corner (300, 400)) into real-world coordinates. If the actual scene is set to 3 meters high, the corresponding spatial coordinates are calculated. For each detected person, obtain their spatial coordinates one by one and store this information in a data structure for later use. Record each person's ID, bounding box coordinates, and spatial coordinates. The record format can be "Person ID: 1, Bounding Box: (300, 400, 350, 450), Spatial Coordinates: (1.5m, 2.0m, 0.0m)." For each detected person, record the time they entered the network point. This time can be obtained from the video frame timestamp to ensure the accuracy of each timestamp. If a person is detected in the fifth frame of the video, and the timestamp of that frame is 00:00:05.000, the person's entry time is recorded as 5 seconds. Each person's initial entry time information is stored along with their spatial coordinates for subsequent analysis and processing, ensuring the integrity of each record. The record format is "Person ID: 1, Entry Time: 00:00:05.000, Spatial Coordinates: (1.5m, 2.0m, 0.0m)." Based on the bounding box coordinates of each detected person, the corresponding image frame is extracted from the motion blur removal monitoring video and image frame segmentation is performed. This can be achieved through simple image cropping. If the bounding box of a person is (300, 400, 350, 450), this region is extracted from the video frame to generate a new image of 50x50 pixels.Embed timestamp information in each extracted image frame for easy tracking and analysis. This can be achieved by adding timestamp text to a corner of the image. Add the text "00:00:05.000" in the upper left corner of the image frame to mark the time information of the image frame. Store each extracted image frame and its corresponding timestamp in a designated folder to facilitate subsequent detection and analysis. You can use serialization to save each image frame. Set the storage path to " / extracted_frames / " and the file name format to "person ID_timestamp.jpg", such as "1_00:00:05.000.jpg".

[0030] In this embodiment, refer to Figure 4 , is a flowchart of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Perform continuous frame target tracking based on multiple real-time image frames of personnel and extract the temporal displacement trajectory of each image frame; Calculate the frequency of each person entering and exiting the network point according to the time series displacement trajectory; Performing a time distribution analysis on the frequency of entering and exiting the network point to obtain the time distribution of the frequency of entering and exiting each person; Calculate the location dwell time based on the time series displacement trajectory to obtain the dwell time of each person at different locations; Based on the entry and exit frequency time distribution of each person and the length of time each person stays at different locations, personalized trajectory features are mined to generate multiple personalized trajectory features of each person.

[0031] In this embodiment, an appropriate target tracking algorithm (such as Kalman filtering, Mean Shift, or CSRT) is selected to track the real-time image frame of each person. The algorithm should be able to maintain target tracking across consecutive frames in the video, ensuring accurate identification even when the person is moving rapidly or obscured. The Discriminative Correlation Filter with Channel and Spatial Reliability (CSRT) algorithm is selected because it performs well with fast-moving targets and those with changing appearances. In the motion blur reduction surveillance video, each person is tracked frame by frame. The position of each image frame in each frame is recorded to facilitate subsequent extraction of the temporal displacement trajectory. In the first frame of the video, the bounding box of person A is (300, 400, 350, 450), and in the second frame it changes to (310, 410, 360, 460). The coordinates of these two positions can be recorded. For each tracked target, its displacement trajectory throughout the video is recorded. The target's bounding box coordinates in each frame are stored as a time series data for subsequent analysis. If the coordinates of person A in the video are [(300, 400), (310, 410), (320, 420)], the resulting temporal displacement trajectory is [(300, 400), (310,410), (320, 420)]. Define the criteria for "entering and exiting the network point." This can typically be determined based on the change in a person's spatial coordinates. If their coordinates move from the outside area into the network point, it counts as an entry; if their coordinates move from the outside area into the network point, it counts as an exit. Set the entry point boundary to x < 2.0 m (within 2 meters of the entrance). If a person's spatial coordinates change from x = 2.5 m to x = 1.5 m, it counts as an entry. Based on the tracked temporal displacement trajectory, count the number of times each target enters and exits the network point during the observation period. A simple counting method can be used to record each entry and exit event. If person A is recorded entering three times and exiting twice during the monitoring period, their entry and exit frequency is 3 / monitoring time (e.g., 30 minutes). The frequency of entry and exit of each person is recorded in a data structure, including the person ID and frequency data, for subsequent analysis. The record format is "Person ID: 1, entry and exit frequency: 0.1 times / minute". To define the calculation method for time distribution, you can choose to divide the monitoring time into multiple time periods (such as every hour, every half hour) for frequency statistics to analyze the changes in entry and exit frequency in different time periods. Divide the monitoring time into 6 time periods (one period every 10 minutes) to observe the changes in entry and exit frequency. Count the entry and exit frequency of each person in each time period and store the results as time distribution data for subsequent analysis. If the entry and exit frequency of person A is 0.2 times in the first 10 minutes, you can record "Time period 1: 0.2 times". The dwell time is defined as the duration of a person's stay at a specific location.This can be calculated using the detected time-series displacement trajectory. If person A spends time at location (1.5m, 2.0m) from 5 seconds to 15 seconds, then their dwell time is 10 seconds. For each target, use the time-series displacement trajectory to calculate their dwell time at each location, recording the start and end times of each dwell. If person A's time at location (1.5m, 2.0m) is recorded as [5s, 15s], then their dwell time is 15s - 5s = 10s. Record each target's dwell time at different locations in a data structure for subsequent analysis. The record format is "Person ID: 1, Location (1.5m, 2.0m) Dwell Time: 10 seconds." Based on the temporal distribution of each target's entry and exit frequency and the dwell time at different locations, a personalized trajectory feature mining model is established. This model should comprehensively account for the influence of entry and exit frequency and dwell time. Parameters for the feature mining model are set, including entry and exit frequency, dwell time, and movement speed. Calculate personalized trajectory characteristics for each target and extract their unique behavioral patterns. This can be achieved through methods such as cluster analysis and association rule analysis. If person A frequently enters and exits a specific area and spends a long time in a specific area, it can be inferred that they are frequent customers. Record each person's personalized trajectory characteristics in a data structure for subsequent analysis and application. The record format is "Person ID: 1, Personalized characteristics: Frequent entry and exit, long stay."

[0032] In this embodiment, step S4 includes the following steps: Perform visual recognition of handheld items for each person in real-time image frames of multiple people, and mark the handheld item image frames; Performing image magnification processing on the handheld object image frame and performing secondary extraction to obtain the handheld object image frame; Perform deep object classification and recognition on the handheld object image frame to obtain object classification information; Identify potential threat objects based on item classification information to obtain potential threat object data; Based on the data of potential threat objects, the corresponding personnel are located and marked, and the image frames of suspicious personnel are extracted.

[0033] In this embodiment, deep learning models (such as YOLO, Faster R-CNN, or RetinaNet) are used for visual recognition of handheld objects. These models should be specifically trained for handheld objects to ensure accurate recognition of a wide range of objects. The YOLOv5 model is used, which offers high recognition speed and accuracy, making it suitable for real-time monitoring. For each detected person, the object recognition model processes their image frame in real time, identifies the object being held, and marks the object's bounding box. This process generates the bounding box coordinates and confidence score for each object. If a person's handheld object is identified as a "bag," its bounding box coordinates (e.g., (400, 500, 450, 550)) and confidence score (e.g., 0.95) are recorded. The recognition results for each person's handheld object are recorded in a data structure, including information such as the person ID, object category, bounding box coordinates, and confidence score, for subsequent analysis. The record format is "Person ID: 1, Item: Bag, Bounding Box: (400, 500, 450, 550), Confidence: 0.95." Based on the recognition results, the image frame of the handheld item is extracted. This can be achieved through image cropping, which extracts the area of ​​the handheld item from the original video frame. If the bounding box of the handheld item is (400, 500, 450, 550), this portion is cropped from the original image to generate a new image. The extracted image frame of the handheld item is upscaled to enhance the visibility and detail of the item. This can be achieved through interpolation algorithms (such as bilinear interpolation or bicubic interpolation) to ensure that the image quality remains clear after upscaling. The image of the handheld item is upscaled to 1.5 times its original size to ensure clear details and suitable for subsequent object classification and recognition. Further feature extraction is performed on the upscaled image to ensure that key information is preserved. Edge detection or feature point detection algorithms (such as SIFT or ORB) can be used to extract features. Use the Canny edge detection algorithm to extract the object's outline for subsequent classification and identification. Select an appropriate deep learning classification model (such as Inception, ResNet, or MobileNet) for object classification. This model must be specifically trained for object classification to ensure it can recognize a wide range of object types. Use the ResNet50 model, which performs well in object classification tasks and is suitable for processing magnified images of handheld objects. Input the extracted and magnified images of handheld objects into the object classification model for deep object classification and identification. Record the object classification information and its confidence level. If the recognition result is "knife" with a confidence level of 0.92, record this information for subsequent analysis. Record the classification results for each handheld object in a data structure, including information such as person ID, item category, bounding box coordinates, and confidence level, to facilitate subsequent threat assessment. The record format is "Person ID: 1, Item Category: Knife, Confidence: 0.92."The criteria for defining potential threat objects typically include offensive or dangerous items (such as weapons and knives). A threat object library can be established based on object classification information. Threat object categories include "knife," "gun," and "explosives." The classification information is compared with the threat object criteria to determine whether a potential threat object exists. If the identified object category is in the threat object library, it is marked as a potential threat. If the identification result is "knife," the object is marked as a potential threat and the relevant information is recorded. The relevant information about the potential threat object is recorded in a data structure, including the person ID, item category, threat level, and bounding box coordinates, for subsequent processing. The record format is "Person ID: 1, Potential Threat Item: Knife, Threat Level: High." Based on the potential threat object data, the relevant person is located and their image frame is marked. This is achieved by comparing the person ID with the potential threat object record information. If the person with person ID 1 is holding a potential threat item, a red bounding box is drawn in the original image to mark it. The marked image frame of the suspicious person is extracted for further analysis and processing. The image frame of the suspicious person is cropped and stored. If the suspicious person's bounding box is (400, 500, 450, 550), extract that area into a new image file named "suspicious_person_1.jpg." Record the suspicious person's related information in a data structure, ensuring that each suspicious person's image frame and the threat object information are clearly visible for subsequent security management and monitoring.

[0034] In this embodiment, the specific steps of step S5 are: Perform dynamic behavior evolution analysis on the suspicious person's image frame to extract the dynamic behavior characteristics of the suspicious person; Conduct multi-period action pattern mining based on the dynamic behavior characteristics of suspicious persons to generate action patterns in multiple time periods; Extracting the personalized trajectory features of the suspicious person based on multiple personalized trajectory features of the persons; Deep dynamic behavior prediction is performed on the personalized trajectory features and action patterns in multiple time periods to construct a dynamic behavior trajectory sequence of suspicious persons.

[0035] In this embodiment, a deep learning model (such as LSTM, GRU, or 3D convolutional neural network) is used to analyze the dynamic behavior of suspicious individuals. These models can process time-series data and extract motion features of individuals in surveillance videos. The LSTM (Long Short-Term Memory) model is selected because it is suitable for capturing dynamic time-series information and effectively analyzing individuals' motion trajectories and behavioral evolution. Based on the image frame and time-series trajectory data of the suspicious individual, the motion trajectory of each suspicious individual is extracted, including information such as their coordinates, velocity, and acceleration for each frame in the video. This data should include timestamps and spatial coordinates. If the trajectory data of a suspicious individual during surveillance is [(x1, y1, t1), (x2, y2, t2), ...], each coordinate point and its corresponding time are recorded. The extracted dynamic behavior information is converted into a feature vector and input into the LSTM model for training. The feature vector should contain dynamic features such as velocity, acceleration, and directional change to facilitate subsequent analysis. The generated feature vector format is [velocity, acceleration, directional change] and can be used for model training and behavioral analysis. Divide the monitoring time into multiple time periods (e.g., every 5 or 10 minutes) to analyze the suspicious individual's movement patterns during these time periods. Each time period should contain corresponding dynamic behavior data. Divide the one-hour monitoring data into 12 5-minute time periods so that each period can be analyzed independently. Use a clustering algorithm (such as K-means or DBSCAN) to perform cluster analysis on the dynamic behavior characteristics within each time period to identify common movement patterns of the suspicious individual. These patterns can reflect the suspicious individual's behavior characteristics during a specific time period. If the suspicious individual's behavior characteristics are clustered into categories such as "wandering" or "rapid movement" within a certain time period, record these movement patterns. Based on the previous personalized trajectory feature analysis, extract the suspicious individual's personalized behavioral features during the monitoring period. These features should reflect the individual's unique behavior patterns, such as frequency of entry and exit and duration of stay. Record the suspicious individual's dwell time and number of entries and exits in specific areas to construct a personalized trajectory feature. Analyze the suspicious individual's trajectory data to extract their personalized trajectory features. This can be achieved by calculating the duration of stay and frequency of activity in various areas. If a suspicious individual stays in a specific area for 20 seconds and enters the area three times during the monitoring period, their characteristics can be represented as "Dwell time: 20 seconds, Entry and exit frequency: 0.5 times / minute." The extracted personalized trajectory features are recorded in a data structure for subsequent behavior prediction and analysis. Ensure that the characteristics of each suspicious individual are clearly visible. The record format is "Personnel ID: 1, Personalized trajectory features: Stay 20 seconds, Entry and exit frequency: 0.5 times / minute." A deep learning model (such as LSTM or GRU) is used for dynamic behavior prediction, combining the suspicious individual's personalized trajectory features with their movement patterns over multiple time periods to predict their future trajectory.The LSTM model is used, which is suitable for processing time series data and can effectively capture the temporal dependencies of dynamic behavior. The individual trajectory features and time-segment motion patterns of suspicious individuals are integrated into the input data and prepared for input into the deep learning model for training and prediction. The input format is [personal trajectory features, motion pattern 1, motion pattern 2], ensuring that the data structure meets the model requirements. After model training is complete, the model is used to predict the future behavior of the suspicious individual, generating a dynamic trajectory sequence. This sequence should include predicted coordinates, timestamps, and possible behavior types. The predicted result is [(x1', y1', t1'), (x2', y2', t2'), ...], representing the expected movement trajectory of the suspicious individual in the future. The predicted dynamic trajectory sequence is recorded in a data structure for subsequent security monitoring and decision support. Ensure that the predicted trajectory information for each suspicious individual is clearly visible. The record format is "Person ID: 1, Predicted trajectory: [(x1', y1', t1'), (x2', y2', t2')]."

[0036] In this embodiment, the specific steps of step S6 are: Obtaining a preset bank branch security zone; performing behavioral trajectory intersection detection on a suspicious person's dynamic behavioral trajectory sequence based on the bank branch security zone; and when an intersection is identified, highlighting the suspicious person and visually marking him / her, and generating a first-level warning signal; Analyze the abnormal and dangerous actions of the dynamic behavior trajectory sequence of suspicious persons. When abnormal and dangerous actions are detected, a secondary warning signal is generated; Identify that the first-level warning signal and the second-level warning signal appear simultaneously on the same person, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0037] In this embodiment, the bank branch's security zones are first determined. These typically include the entrance, the perimeters surrounding the entrances and exits, the counter area, and high-risk areas (such as cash storage areas). These zones should be calibrated within the monitoring system to ensure accurate identification of these specific areas. The boundary coordinates of the security zones are set to [(x1, y1), (x2, y2), (x3, y3), (x4, y4)], and the corresponding polygonal areas are drawn within the video surveillance system. The defined security zones are stored in a database for subsequent retrieval and use. The data structure should include information such as the security zone ID, name, and boundary coordinates. The record format is "Security Zone ID: 1, Name: Entrance, Boundary Coordinates: [(x1, y1), (x2, y2), ...]." Sequential data on the dynamic behavior trajectories of suspicious individuals is collected, ensuring that the data includes the coordinates and timestamp information for each trajectory point. The trajectory data should include the suspicious individual's movement path within the monitoring area. The trajectory data format is "Personnel ID: 1, Trajectory: [(x1, y1, t1), (x2, y2, t2), ...]." Intersection detection is performed on the dynamic trajectory sequence of a suspicious individual within the preset warning zone. Geometric calculation methods (such as line segment intersection detection) can be used to determine whether a trajectory point enters the warning zone. If the trajectory point (x, y) is within the warning zone boundary, the individual's trajectory is considered to intersect the warning zone. Once a suspicious individual's dynamic trajectory is detected to intersect the warning zone, the individual is immediately highlighted visually, for example, with a red frame on the monitoring screen, and a Level 1 warning signal is generated. The warning information is recorded in the format of "Personnel ID: 1, Warning Level: Level 1, Timestamp: t1." The criteria for defining abnormal and dangerous behavior typically include fast running, frequent turning back, and constant position changes, which are inconsistent with normal behavior. These actions should be identified by analyzing the dynamic behavior characteristics of the suspicious individual. The standard for "fast running" is set as moving more than 5 meters in 1 second or a detected speed exceeding 3 meters per second. Suspicious individuals' dynamic trajectory sequences are analyzed for unusual and dangerous movements. Rule-based detection methods or deep learning models (such as motion recognition networks) are used to identify abnormal behavior. If a suspicious individual is detected moving at a speed exceeding 3 meters per second within a short period of time, a Level 2 warning signal is generated. Detected unusual and dangerous movements and related information are recorded in a data structure, including the individual's ID, action type, warning level, and timestamp. The record format is "Personnel ID: 1, Warning Level: Level 2, Action Type: Running Fast, Timestamp: t2." Simultaneous occurrences of Level 1 and Level 2 warning signals are identified, allowing for rapid emergency response for the same suspicious individual. The emergency response strategy should include notifying security personnel, activating surveillance video, and triggering an alarm. If both Level 1 and Level 2 warnings are triggered for Person 1, the system should immediately notify security personnel to proceed to the area.When the system detects the simultaneous warning signals, it initiates the pre-set emergency response process, including recording the event information, accessing surveillance footage, and dispatching security personnel to the scene. The response event is recorded as "Time: t3, Personnel ID: 1, Response Result: Security has arrived, assessing the situation."

[0038] In this embodiment, a security monitoring system based on image recognition is provided, which is used to execute the security monitoring method based on image recognition as described above, including: The video optimization module is used to obtain monitoring videos of the bank from multiple angles and perform directional motion deconvolution to construct monitoring videos with motion blur eliminated; Image segmentation module, used to perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video, and extract multiple real-time image frames of people; The target tracking module is used to track targets in consecutive frames based on multiple real-time image frames of people, and to mine personalized trajectory features, thereby generating multiple personalized trajectory features of people; The object visual recognition module is used to visually identify the objects held by each person and identify potential threats in multiple real-time image frames, and extract image frames of suspicious persons; The behavior prediction module is used to perform deep dynamic behavior prediction on the suspicious person's image frame based on multiple personalized trajectory features and construct a dynamic behavior trajectory sequence of the suspicious person; The emergency response decision module is used to analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

[0039] This invention uses directional motion deconvolution to effectively remove blur caused by rapid motion, which is crucial in dynamic environments like bank branches. Improved video clarity enables more accurate subsequent object detection and behavior analysis. Optimized surveillance video reduces misjudgments due to blur and ensures the fidelity of image details. This allows security personnel to clearly see the details of each individual, increasing the reliability of the surveillance video. Motion blur removal not only improves overall video quality but also enhances the ability to identify fast-moving objects, ensuring the system can accurately capture their behavior even in rapidly moving objects or in poorly lit environments. Deep learning technologies (such as YOLO and Faster R-CNN) can simultaneously identify multiple objects, ensuring real-time monitoring of the dynamic behavior of multiple individuals in bank branches without missing any potential threats. Image frame segmentation effectively prevents misidentification and missed detections in object detection. Each individual's image frame is independent and clear, enhancing the accuracy of subsequent processing. The system can quickly identify all objects in the image and update the image frames in real time, making the entire monitoring process efficient and highly real-time. Bank branches often involve complex personnel flows, and image segmentation technology can effectively improve the response speed of real-time monitoring. Continuous-frame target tracking technology ensures stable tracking of targets across multiple video frames. Even with occlusion or changes in movement, the system maintains consistent target identification, improving tracking accuracy. Each person's trajectory features unique characteristics, such as frequent activity areas, dwell time, and movement speed. Mining these personalized trajectory features provides unique data support for subsequent behavioral analysis, helping to identify potential anomalous behavior. Accurate trajectory tracking lays the foundation for subsequent abnormal behavior detection. The system can identify anomalies in a person's behavior patterns by comparing them with historical trajectory data and issue timely alerts. Using deep learning algorithms to visually identify objects held by individuals, the system can accurately identify potential threats, such as knives, guns, and packages, which is crucial for bank branch security. Identifying and assessing threats from held objects allows for the timely detection of potential danger sources and early warnings. By flagging suspicious individuals, security personnel can take preventative measures and mitigate security risks. The handheld object recognition module overcomes the inability of traditional surveillance methods to identify hidden threats, ensuring that any potential danger is promptly captured. By dynamically predicting human behavior, the system can anticipate suspicious individuals' movements, such as entering restricted areas or engaging in violent behavior, before actual threats occur, providing early warnings and mitigating risks. Dynamic behavior prediction significantly reduces security personnel's response delays. By accurately sequencing behavioral trajectories, security personnel can respond quickly and ensure the safety of bank branches. The behavior prediction module provides data support for optimizing security strategies, enabling the development of more refined security measures, such as adjusting key monitoring areas and strengthening patrols.The system rapidly generates emergency response decisions based on analysis of unusual and dangerous behavior. For example, when threatening behavior is detected, the system automatically triggers an alert and initiates appropriate response procedures based on the threat level (such as locking exits and notifying security). The intelligent nature of the emergency response decision module ensures standardized and automated responses, reduces errors caused by human intervention, and improves the efficiency and accuracy of incident response. Based on behavioral predictions and hazard analysis, the system automatically constructs emergency response strategies tailored to different threat types. This intelligent strategy enables fully automated, real-time response, enabling security teams to quickly make decisions and handle emergencies.

[0040] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0041] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A security monitoring method based on image recognition, characterized in that: The following steps are involved: Step S1: Obtain monitoring videos of the bank from multiple angles and perform directional motion convolution elimination to construct a dynamic blur elimination monitoring video; Step S2: Perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video to extract multiple real-time image frames of people; Step S3: Continuous frame target tracking is performed based on multiple real-time image frames of personnel, and personalized trajectory feature mining is performed to generate multiple personalized trajectory features of personnel; Step S4: Perform visual recognition of handheld items and potential threat objects on multiple real-time image frames of people one by one, and extract image frames of suspicious people; Step S5: performing deep dynamic behavior prediction on the suspicious person image frame based on multiple personalized trajectory features, and constructing a dynamic behavior trajectory sequence of the suspicious person; Step S6: Analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

2. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S1 are: Obtain monitoring videos from multiple angles of the bank based on multi-directional cameras at bank branches; Performing a network point ambient light calculation on the monitoring video to identify the ambient light intensity of multiple areas; Identify dark areas according to the ambient light intensity and extract the dark areas; Performing gamma adaptive correction on the dark area to obtain brightness optimized monitoring videos at multiple angles; Perform global time synchronization on brightness optimization monitoring videos from multiple angles to obtain time-synchronized monitoring videos; Directional motion convolution is performed on the time-synchronized monitoring video to construct a dynamic blur-removed monitoring video.

3. The image recognition-based security monitoring method according to claim 2, characterized in that: The specific steps of performing global time synchronization on the brightness optimization monitoring videos of multiple angles to obtain time-synchronized monitoring videos are as follows: Marking a same scene feature point in the brightness optimization monitoring videos of the multiple angles; Performing panoramic spatial registration processing on the monitoring videos of the multiple angles based on the same scene feature points to construct a panoramic monitoring video; Perform perspective fusion boundary recognition on panoramic monitoring videos and extract perspective fusion boundary lines; Smooth transition optimization is performed on the perspective fusion boundary line to obtain a panoramic monitoring video with smoothed and optimized boundaries; Calculating the acquisition frequency timestamp of the multi-directional camera; Global time synchronization is performed on the boundary smoothing optimized panoramic monitoring video according to the timestamp to obtain a time-synchronized monitoring video.

4. The image recognition-based security monitoring method according to claim 2, characterized in that: The specific steps of performing directional motion convolution elimination on the time-synchronized monitoring video to construct the dynamic blur elimination monitoring video are as follows: Perform time-sequential frame decomposition on the time-synchronized monitoring video and extract all image frames; Calculating a time difference between adjacent image frames of the image frame; Calculating the total inter-frame delay according to the time difference; Calculating a delay average of all inter-frame delays; Adaptively adjusting the frame rate according to the delay average value to obtain an adaptive frame rate; Performing global frame delay optimization on the time-synchronized monitoring video based on the adaptive frame rate to obtain a global delay-optimized monitoring video; Perform dynamic target motion blur analysis on the global delay optimization monitoring video and extract dynamic target motion blur data; Motion vector estimation is performed based on the motion blur data of the dynamic target to obtain the motion trajectory and speed of the dynamic target; Directional motion convolution is performed based on the motion trajectory and speed, thereby constructing a dynamic blur elimination monitoring video.

5. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S2 are: Perform multi-target depth detection on the surveillance video with motion blur removal, marking all personnel within the network site; Locate the spatial position of all personnel in the network one by one and obtain the real-time spatial coordinates of each person; Calculating the time when the person entered the network point to obtain initial time information of entering the network point; The dynamic blur elimination monitoring video is segmented into image frames according to the real-time spatial coordinates of each person, and timestamps are marked based on the initial entry time information to extract multiple real-time image frames of the person.

6. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S3 are: Perform continuous frame target tracking based on multiple real-time image frames of personnel and extract the temporal displacement trajectory of each image frame; Calculate the frequency of each person entering and exiting the network point according to the time series displacement trajectory; Performing a time distribution analysis on the frequency of entering and exiting the network point to obtain the time distribution of the frequency of entering and exiting each person; Calculate the location dwell time based on the time series displacement trajectory to obtain the dwell time of each person at different locations; Based on the entry and exit frequency time distribution of each person and the length of time each person stays at different locations, personalized trajectory features are mined to generate multiple personalized trajectory features of each person.

7. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S4 are: Perform visual recognition of handheld items for each person in real-time image frames of multiple people, and mark the handheld item image frames; Performing image magnification processing on the handheld object image frame and performing secondary extraction to obtain the handheld object image frame; Perform deep object classification and recognition on the handheld object image frame to obtain object classification information; Identify potential threat objects based on item classification information to obtain potential threat object data; Based on the data of potential threat objects, the corresponding personnel are located and marked, and the image frames of suspicious personnel are extracted.

8. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S5 are: Perform dynamic behavior evolution analysis on the suspicious person's image frame to extract the dynamic behavior characteristics of the suspicious person; Conduct multi-period action pattern mining based on the dynamic behavior characteristics of suspicious persons to generate action patterns in multiple time periods; Extracting the personalized trajectory features of the suspicious person based on multiple personalized trajectory features of the persons; Deep dynamic behavior prediction is performed on the personalized trajectory features and action patterns in multiple time periods to construct a dynamic behavior trajectory sequence of suspicious persons.

9. The image recognition-based security monitoring method according to claim 1, characterized in that: The specific steps of step S6 are: Obtaining a preset bank branch security zone; performing behavioral trajectory intersection detection on a suspicious person's dynamic behavioral trajectory sequence based on the bank branch security zone; and when an intersection is identified, highlighting the suspicious person and visually marking him / her, and generating a first-level warning signal; Analyze the abnormal and dangerous actions of the dynamic behavior trajectory sequence of suspicious persons. When abnormal and dangerous actions are detected, a secondary warning signal is generated; Identify that the first-level warning signal and the second-level warning signal appear simultaneously on the same person, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

10. A security monitoring system based on image recognition, characterized in that: The method for performing the image recognition-based security monitoring method according to claim 1 comprises: The video optimization module is used to obtain monitoring videos of the bank from multiple angles and perform directional motion deconvolution to construct monitoring videos with motion blur eliminated; Image segmentation module, used to perform multi-target depth detection and image frame segmentation on the dynamic blur elimination monitoring video, and extract multiple real-time image frames of people; The target tracking module is used to track targets in consecutive frames based on multiple real-time image frames of people, and to mine personalized trajectory features, thereby generating multiple personalized trajectory features of people; The object visual recognition module is used to visually identify the objects held by each person and identify potential threats in multiple real-time image frames, and extract image frames of suspicious persons; The behavior prediction module is used to perform deep dynamic behavior prediction on the suspicious person's image frame based on multiple personalized trajectory features and construct a dynamic behavior trajectory sequence of the suspicious person; The emergency response decision module is used to analyze abnormal and dangerous actions based on the dynamic behavior trajectory sequence of suspicious persons, make rapid emergency response decisions, and build an intelligent emergency response analysis strategy.

Citation Information

Cited By

  • Monitoring method and system for preventing external force damage, electronic equipment and storage medium

    CN121191091A

  • Suspicious person monitoring method and system for bank security

    CN121747202A