Security and protection system and security and protection method based on multi-mode perception and photoelectric linkage

By using a multimodal sensing and photoelectric linkage security system, which utilizes panoramic cameras, detail cameras, and thermal imaging modules, combined with advanced image processing and target tracking algorithms, the system solves the problems of blind spots and response delays in existing security systems in complex environments, and achieves efficient target recognition and tracking.

CN120977056APending Publication Date: 2025-11-18中国人民武装警察部队杭州支队
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510906613.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing security systems suffer from large blind spots, poor environmental adaptability, and high response delays in complex environments, making it difficult to meet the needs of large-scale and timely security.

Method used

A security system based on multimodal perception and photoelectric linkage is adopted, including a panoramic camera, a detail camera, a thermal imaging module and a PTZ pan-tilt unit. It combines LAB color space conversion, CLAHE algorithm and Brown-Conrady distortion correction, and uses YOLOv7 network detection, DeepSORT algorithm tracking and LSTM network prediction to achieve second-level target localization and tracking.

Benefits of technology

It eliminates blind spots in monitoring, improves target recognition accuracy and response speed, reduces the workload of manual monitoring, builds a closed-loop security system that adapts to complex environments and reduces the rate of missed detections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977056A_ABST
    Figure CN120977056A_ABST
Patent Text Reader

Abstract

The invention relates to a security system and a security method based on multi-modal sensing and photoelectric linkage, the security system comprises an environment sensing module, the environment sensing module comprises a panoramic camera, a detail camera and a thermal imaging module, the detail camera is provided with a PTZ holder, and the PTZ holder is also provided with a laser module; the system further comprises an on-duty terminal, the on-duty terminal comprises a display module, a sound-light alarm module, a controller and a communication module, the controller is in communication connection with the environment sensing module, the PTZ holder and the laser module through the communication module, and the display module and the sound-light alarm module are in communication connection with the controller. According to the invention, monitoring blind areas in a complex environment are eliminated, the target identification precision is improved, second-level automatic tracking and accurate positioning are realized, the emergency disposal time is shortened, the manual monitoring load is reduced, security vulnerabilities caused by attention fatigue are avoided, and a monitoring-tracking-evidence obtaining-disposal whole-process closed-loop security system is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of security and defense technology, and in particular to a security and defense system based on multi-modal perception and photoelectric linkage and a security and defense method. BACKGROUND

[0002] The current military security field mainly relies on two technical modes of artificial sentry combined with fixed monitoring and traditional electronic fence system. The artificial sentry mode realizes monitoring through human patrol combined with fixed cameras, but is limited by the fact that the visual angle of the fixed camera is usually less than 90°, and there are significant monitoring blind spots in complex terrain, and the average delay of artificial response is more than 8 seconds, and the long-term sentry is prone to cause the detection rate to exceed 15% due to attention fatigue. The traditional electronic fence system realizes perimeter alarm based on infrared or vibration sensing, although it can trigger sound and light alarms, but it can only complete single-point early warning and cannot actively track and locate the intrusion target, and the time from discovering the target to artificial disposal is more than 10 seconds.

[0003] Both types of technology have the common problem of insufficient environmental adaptability: under night, rain, fog and other low-light or adverse weather conditions, the accuracy of the traditional visual recognition scheme drops by more than 40%, and the false positive rate increases significantly; at the same time, the monitoring mode of artificial intervention needs to continuously invest in human cost, and it is difficult to meet the security and defense needs of a large range and high timeliness. The defects of existing technologies in terms of automation tracking, multi-scene adaptability, response timeliness and human cost control, etc. urgently need to build a more intelligent and efficient security and defense system. SUMMARY

[0004] In order to solve the problems of existing security and defense systems relying on manpower, having large monitoring blind spots, weak environmental adaptability, and high response delay, a security and defense system and method based on multi-modal perception and photoelectric linkage are proposed, which eliminates monitoring blind spots in complex environments, improves target recognition accuracy, realizes second-level automatic tracking and accurate positioning, shortens emergency disposal time, reduces artificial monitoring load, avoids security and defense loopholes caused by attention fatigue, and builds a monitoring-tracking-evidence-disposal full-process closed-loop security and defense system.

[0005] To achieve the above purpose, the first aspect of the present application proposes a security and defense system based on multi-modal perception and photoelectric linkage, which comprises an environment perception module, the environment perception module comprises a panoramic camera, a detail camera and a thermal imaging module, the panoramic camera is fixedly arranged on both sides of a detection area, the detection ranges of two panoramic cameras cover the detection area, the detail camera is provided with a PTZ holder, and a laser module is further arranged on the PTZ holder;

[0006] Further comprising a on-duty terminal, the on-duty terminal comprises a display module, an audible and light alarm module, a controller and a communication module, the controller is in communication connection with the environment perception module, the PTZ holder and the laser module through the communication module, the display module, the audible and light alarm module and the controller are in communication connection.

[0007] The second aspect of the application provides a security and protection system and method based on multi-modal perception and photoelectric linkage, comprising:

[0008] Step 1: Draw a polygonal electronic fence in a three-dimensional geographic information system, start a double panoramic camera and a thermal imaging module to collect images, convert the images to a LAB color space and perform adaptive light compensation using a CLAHE algorithm, and correct the images in real time through a Brown-Conrady distortion model;

[0009] Step 2: Scale the corrected images to a resolution of 640x640, input them into a YOLOv7 network for target detection, use Mosaic data enhancement preprocessing, extract features through an E-ELAN backbone network, fuse three scale feature maps using FPN, output detection boxes with a confidence score >0.6, and filter targets based on biological morphological parameters;

[0010] Step 3: Use DeepSORT algorithm for the detected target, predict an 8-dimensional state vector through Kalman filtering, combine it with a 128-dimensional appearance feature extracted by ResNet50 for trajectory association, input the motion state of 10 consecutive frames into a three-layer LSTM network, and predict the trajectory for the next 1.2 seconds by stacking the spatio-temporal attention mechanism;

[0011] Step 4: Trigger a three-level response according to the target behavior, including:

[0012] Primary response: When the target intrudes into the electronic fence, start the 5Hz stroboscopic searchlight and 110dB directional sound wave;

[0013] Intermediate response: When the target stays for more than 5 seconds, continuously track it through the PTZ holder and record it through the detail camera;

[0014] Emergency response: After manual confirmation, turn on the laser module to suppress the target;

[0015] Step 5: Superimpose the AR light track on the three-dimensional map of the on-duty terminal to display the target position, encrypt the trajectory data and store it in eMMC.

[0016] Further, step 1 includes:

[0017] Step 1.1 Adaptive light compensation: Divide the image into 8x8 grids, and clip the gray histogram in each grid so that the number of pixels at each gray level does not exceed a threshold T:

[0018]

[0019] where N is the total number of pixels in the grid, and β is the limiting factor;

[0020] The clipping threshold is expressed as formula (2):

[0021]

[0022] where βl is the clipping factor, taking 2 or 3, and then the excess part is evenly distributed to all gray levels;

[0023] Step 1.2 Correct the image: Set the adaptive light compensation image coordinates as (x, y), and the corrected coordinates as (x corrected ,y cprrected );

[0024] where the correction includes radial distortion as shown in formula (3):

[0025]

[0026] where r = x 2 +y 2 ;

[0027] Also includes tangential distortion as shown in formula (4):

[0028]

[0029] Dual panoramic cameras cover the detection area, eliminating the monitoring blind area of fixed cameras, cooperating with three-dimensional electronic fences to realize global situation awareness, and combining LAB color space conversion with thermal imaging modules to break through the problem of low recognition rate at night / rainy weather. Through multi-modal data, the environmental adaptability is improved; CLAHE algorithm and Brown-Conrady distortion correction solve the problems of uneven brightness and geometric distortion, and improve the subsequent target detection accuracy.

[0030] Further, the input YOLOv7 network for target detection in step 2 includes:

[0031] Bounding box regression: as described in formulas (5-8):

[0032] b x =σ(t x )+c x (5);

[0033] b y =σ(t y )+c y (6);

[0034]

[0035] wherein (c x ,c y ) is the coordinate of the upper left corner of the current grid, (p w ,p h ) is the width and height of the preset anchor frame,

[0036] (t x ,t y ,t w ,t h ) is the offset of the network prediction, and sigma is a sigmoid function that compresses the prediction value to the range of (0, 1) to ensure that the center point is within the current grid; the final target frame center coordinates are (b x ,b y ), and the width and height are (b w ,b h );

[0037] The confidence screening is shown in formula (9):

[0038]

[0039] wherein B det and B pred represent the areas of the two boundary frames, respectively;

[0040] NMS step: for each class, sort by confidence from high to low, then iterate one by one, and keep the detection frame with confidence > 0.6, and remove redundant frames by non-maximum suppression (IoU threshold 0.5).

[0041] Mosaic data enhancement plus E-ELAN backbone network improves the small target detection capability in complex background, FPN three-scale feature fusion, and gives consideration to target positioning accuracy and semantic information, and outputs the detection frame with confidence > 0.6 to meet the high reliability requirement of the security scene.

[0042] Further, the biological form parameters in step 2 include:

[0043] The multi-spectrum biological recognition engine is shown in formula (10):

[0044] F (l) = ReLU(W (l) *F (l-1) +b (l) ) (10);

[0045] wherein W (l) is a convolution kernel (64 3x3 convolution kernels), * represents a convolution operation, and the ReLU activation function is ReLU(x) = max(0, x);

[0046] Multiple different scale max pooling is performed on the input feature map, such as 4x4, 8x8, 16x16, and then all the pooling results are flattened and spliced as shown in formula (11):

[0047] F spp = [MaxPool(F,4x4), MaxPool(F,8x8), MaxPool(F,16x16)] (11);

[0048] Wherein, the value of the output position (i, j) of each max pooling operation is the maximum value in the corresponding window in the input;

[0049] Physical size calculation is triggered when the target body length is greater than 1.2m and the shoulder height is greater than 0.6m.

[0050] Combined with the multispectral recognition engine, animal or small object interference is effectively excluded, and the detection accuracy is improved.

[0051] Further, the step 3 includes:

[0052] Step 3.1 motion modeling:

[0053] Take the state vector X = [x, y, w, h, v x ,v y ,v w ,v h ] T , wherein (x, y) is the center position, (w, h) is the boundary box width and height, (v x ,v y ) is the center point speed, and (v w ,v h ) is the width and height change rate;

[0054] State transition equation: x k|k-1 = Fx k-1 (12);

[0055] Wherein, the state transition matrix F is:

[0056]

[0057] Delta t is the time interval (the time between two frames);

[0058] Covariance update: P k|k-1 = FP k-1 F T + Q (13);

[0059] Wherein, P is the state covariance matrix, and Q is the process noise covariance matrix.

[0060] Step 3.2 Calculate the similarity between detection box and predicted trajectory space by Mahalanobis distance as shown in equation (14):

[0061]

[0062] where: y j is the state of the jth detection box, usually: [x, y, w, h] T i is the predicted state of the ith trajectory, H is the observation matrix that maps the state vector to the observation space, H = [I4, 0 4x4 ], that is, only take the position and size, S i = HP i H T + R, R is the observation noise covariance, when the Mahalanobis distance is greater than 0.8, it is considered that the association is invalid.

[0063] Step 3.3 Data association:

[0064] Cosine distance calculation formula d (2) (i,j) = 1-f i T f j (15);

[0065] where f i (128 dimensions, ||f i ||2 = 1), since the feature vector has been normalized, the cosine distance is 1 minus the dot product of the two vectors;

[0066] Step 3.4 LSTM trajectory prediction model:

[0067] Position coordinates are normalized and linearly scaled to [-1, 1], and velocity vectors are unitized;

[0068] Let the input vector be x i , the hidden state at the last time be h t-i , and the memory cell state be C t-1

[0069] Forget gate, input gate, and output gate calculation:

[0070]

[0071] Candidate memory cell state:

[0072] Update the memory cell state:

[0073] Output hidden state: h t = o t ⊙ tanh(C t )(19).​

[0074] where σ is a sigmoid function, and is an element-wise multiplication;

[0075] Step 3.5 future trajectory prediction:

[0076] Input: LSTM last time hidden state h t ;

[0077] Output: future 6 frames 1.2 seconds of trajectory coordinate sequence

[0078]

[0079] where W y and b y are the weight matrix and bias vector of the full connection layer.

[0080] Kalman filter 8-dimensional state vector prediction, accurate modeling of target motion state, solve the problem of high delay of traditional system positioning, ResNet50 appearance feature plus cosine distance correlation, improve the accuracy of cross-frame target matching, avoid missing detection caused by target loss, three-layer LSTM superposition of space-time attention, 1.2 seconds in advance to predict trajectory, provide decision basis for PTZ head compensation.

[0081] Further, the intermediate response of step 4 includes:

[0082] Convert the predicted trajectory of step 3 into PTZ head spherical coordinates (θ, ψ):

[0083] Let the image pixel coordinates be (u, v), the camera intrinsic parameters be focal length (f x , f y ), and the principal point (c x , c y )

[0084]

[0085] Convert to spherical coordinates:

[0086]

[0087] Drive the PTZ head by PID control:

[0088] Let e(t) be the current time angle deviation;

[0089] PID control quantity:

[0090] where K p = 0.8, K i = 0.05, and K d = 0.2.

[0091] Supplement the trajectory offset for preview compensation:

[0092] Suppose the current calculated target angle is θ, φ, and the LSTM predicted trajectory offset is Δθ, Δφ

[0093] The compensated target angle is:

[0094]

[0095] Wherein, α is the preview gain coefficient.

[0096] By accurately mapping the coordinates through the camera internal parameter, the positioning accuracy is improved; combined with PID control, the dynamic stability is enhanced; by using LSTM preview compensation, the tracking real-time performance is improved; multiple modules are cooperated to enhance the robustness, and the PTZ holder tracking is realized.

[0097] Through the above technical scheme, the beneficial effects of the present application are:

[0098] 1. The present application adopts a multi-modal perception scheme of double panoramic cameras combined with a thermal imaging module, the two panoramic cameras cross cover the detection area, cooperate with the infrared perception ability of the thermal imaging, solve the monitoring problems in the visual angle blind area of the traditional single camera and the low light, haze and other harsh environments. The PTZ holder links the detail camera and the laser module, can focus and track the target detected by the panoramic camera in real time, forms a three-dimensional monitoring network of "global scanning and local close-up".

[0099] 2. The method of the present application improves the image definition through LAB color space conversion, CLAHE adaptive light compensation and distortion correction technology, provides high-quality data source for subsequent detection, adopts YOLOv7 detection network combined with biological form parameter filtering to reduce invalid alarms triggered by non-targets such as animals and sundries, and improves the effectiveness of monitoring data. The DeepSORT algorithm is used in combination with Kalman filtering and ResNet50 appearance feature trajectory association, which is accurate and high, the LSTM network is used to stack the spatio-temporal attention mechanism to predict the future 1.2 second trajectory, and the PID preview control of the PTZ holder is used to realize the continuous and stable tracking of the dynamic target. Even if the target moves quickly, the system can still adjust the holder angle in advance through preview compensation to ensure that the monitoring picture is not lost. A monitoring-tracking-evidence-disposal whole-process closed-loop security system is constructed. BRIEF DESCRIPTION OF DRAWINGS

[0100] Figure 1 It is a structural schematic view of the security system based on multi-modal perception and photoelectric linkage of the present application;

[0101] Figure 2 It is a step flow chart of the security method based on multi-modal perception and photoelectric linkage of the present application. DETAILED DESCRIPTION

[0102] Embodiment 1

[0103] As Figure 1 shown, a security system based on multi-modal perception and photoelectric linkage includes an environment perception module, the environment perception module includes a panoramic camera, a detail camera and a thermal imaging module, the panoramic camera is fixedly arranged on both sides of a detection area, the detection ranges of the two panoramic cameras cover the detection area, the detail camera is provided with a PTZ holder, and a laser module is further arranged on the PTZ holder;

[0104] Further comprising a on-duty terminal, the on-duty terminal includes a display module, an audible and visual alarm module, a controller and a communication module, the controller is in communication connection with the environment perception module, the PTZ holder and the laser module through the communication module, and the display module, the audible and visual alarm module and the controller are in communication connection.

[0105] In this embodiment, the panoramic imaging module adopts 2 Hikvision DS-2CD63 series 190° super wide-angle fisheye cameras, the horizontal field of view is increased by 210% compared with the human eye, and it supports 0.0005 lux starlight level low illumination imaging. The detail camera adopts Dahua DH-IPC-HFW5849 series 30 times optical zoom detail camera (resolution 2560x1440@60fps), supports H.265 encoding and digital image stabilization function. The thermal imaging module adopts multi-spectral perception module integrated FLIR Lepton 3.5 thermal imaging module, realizes visible light / infrared dual-channel data acquisition. The PTZ holder selects Pelco Spectra HD PTZ (horizontal 360° continuous rotation, pitch -90°~+90°), the PTZ holder is provided with a harmonic reduction stepper motor (positioning accuracy ±0.1°), and an absolute value encoder feedback angle in real time. The laser module adopts Osram SPL LL90_3 high-power LED laser module wavelength 450nm blue light (rain and mist scattering rate reduced by 40%), micro-lens array homogenization technology, spot uniformity >90%; peak light intensity 200,000 lux (clear weather illumination ≥2000 meters). The controller adopts NXP i.MX8M Plus quad-core Cortex-A53@1.8GHz, integrated 2.3TOPS NPU accelerates AI inference. The communication module adopts Quectel RM500Q 5G module, supports 3 kilometers low delay transmission (end-to-end delay <50ms). The display module adopts 10.1 inch IPS LCD (1920x1200), HDMI 2.0 interface real-time rendering 32 times zoom picture+AR light rail guide. The sound and light alarm module adopts digital power amplifier to drive piezoelectric ceramic loudspeaker, outputs 110dB directional sound (effective distance 50 meters) setting data storage module selects 128GB eMMC 5.1, circularly stores 30 days H.265 encoding track data. The system uses SM2 national secret algorithm for encryption to prevent signal hijacking. The system sets double redundant power supply system, including main circuit: AC24V→DC12V step-down module (TI TPS5430, conversion efficiency 95%), backup circuit: super capacitor group (100F / 5.5V), cold start switching time ≤1ms.

[0106] Example 2

[0107] Based on the security system based on multi-modal perception and photoelectric linkage in example 1, the security method is illustrated in this example:

[0108] As Figure 2 shown, the security method comprises:

[0109] Step 1: Draw a polygon electronic fence in the three-dimensional geographic information system, start the dual panoramic camera and thermal imaging module to collect images, convert the images to LAB color space and perform adaptive light compensation using CLAHE algorithm, and correct the images in real time through the Brown-Conrady distortion model;

[0110] Step 2: Scale the corrected images to 640x640 resolution, input the YOLOv7 network for target detection, use Mosaic data enhancement preprocessing, extract features through the E-ELAN backbone network, fuse three scale feature maps using FPN, output detection boxes with confidence>0.6, and filter targets based on biological morphological parameters;

[0111] Step 3: Use DeepSORT algorithm for detected targets, predict 8-dimensional state vectors through Kalman filtering, combine 128-dimensional appearance features extracted by ResNet50 for trajectory association, input continuous 10-frame motion states into a three-layer LSTM network, and predict future 1.2-second trajectories by stacking spatio-temporal attention mechanisms;

[0112] Step 4: Trigger three-level responses according to target behavior, including:

[0113] Primary response: When the target intrudes into the electronic fence, start the searchlight 5Hz stroboscopic and 110dB directional sound wave;

[0114] Intermediate response: When the target stays for more than 5 seconds, continuously track through the PTZ cloud platform and record through the detail camera;

[0115] Emergency response: After manual confirmation, turn on the laser module for laser suppression;

[0116] Step 5: Superimpose AR light track on the three-dimensional map of the on-duty terminal to display the target position, encrypt the track data and store it to eMMC.

[0117] Step 1 includes:

[0118] Step 1.1 Adaptive light compensation: Divide the image into 8x8 grids, and clip the gray histogram in each grid so that the number of pixels at each gray level does not exceed the threshold T:

[0119]

[0120] Where N is the total number of pixels in the grid, and β is the limiting factor;

[0121] The clipping threshold is represented by formula (2):

[0122]

[0123] Wherein, beta1 is a clipping factor, taking 2 or 3, and then the excess part is evenly distributed to all gray levels;

[0124] Step 1.2 corrects the image: the adaptive light compensation image is set to normalized image coordinates (x, y), and the corrected coordinates are (x corrected ,y cprrected );

[0125] Wherein, the correction includes radial distortion as shown in formula (3):

[0126]

[0127] Wherein, r=x 2 +y 2 ;

[0128] Also includes tangential distortion as shown in formula (4):

[0129]

[0130] The reverse mapping grid is constructed, and the non-distortion image is generated by bilinear interpolation, and the distortion rate is controlled at 0.21%.

[0131] Step 2 specifically includes:

[0132] First, set up a multispectral biometric recognition engine to process the image:

[0133] Dual-channel data alignment: match visible light and thermal imaging coordinates through affine transformation;

[0134] Five-layer convolutional neural network architecture:

[0135] Conv1: 64 3x3 convolution kernels extract contour features, and ReLU activation;

[0136] SPP layer: fusion of 4x4 / 8x8 / 16x16 multi-scale features;

[0137] Fully connected layer outputs 7-class biological probability (wolf, boar, etc.);

[0138] The biological form parameters of step 2 include:

[0139] The multispectral biometric recognition engine is shown in formula (10):

[0140] F (l) =ReLU(W (l) *F (l-1) +b (l) )(10);

[0141] Wherein, W (l)For the convolution kernel (64 3x3 convolution kernels), * represents the convolution operation, and the ReLU activation function is ReLU(x) = max(0, x).

[0142] The input feature map is subjected to maximum pooling of multiple different scales, such as 4x4, 8x8, and 16x16, and then all the pooling results are flattened and spliced as shown in formula (11):

[0143] F spp =[MaxPool(F,4x4), MaxPool(F,8x8), MaxPool(F,16x16)](11);

[0144] Wherein, the value of the output position (i, j) of each maximum pooling operation is the maximum value in the corresponding window in the input;

[0145] Physical size calculation is triggered when the target body length is greater than 1.2 m and the shoulder height is greater than 0.6 m.

[0146] Step 2.1 Pretreatment stage: The input frame is scaled to 640x640 resolution, and the pixel value is normalized to the [0, 1] interval;

[0147] Mosaic data enhancement is used to improve the robustness of small target detection.

[0148] Step 2.2 Feature extraction and fusion:

[0149] Backbone network: E-ELAN extended architecture, optimized and accelerated inference through group convolution and gradient path;

[0150] Feature pyramid (FPN) fusion of 80x80 / 40x40 / 20x20 three-scale feature maps.

[0151] Step 2.3 Detection head output:

[0152] Input YOLOv7 network for target detection:

[0153] Bounding box regression: as described in formulas (5-8):

[0154] b x =σ(t x )+c x (5);

[0155] b y =σ(t y )+c y (6);

[0156]

[0157] Wherein, (c x ,cy ) is the coordinate of the upper left corner of the current grid, (p w ,p h ) is the width and height of the preset anchor frame,

[0158] (t x ,t y ,t w ,t h ) is the offset of the network prediction, and sigma is a sigmoid function that compresses the prediction value to the range of (0, 1) to ensure that the center point is within the current grid; the final target frame center coordinates are (b x ,b y ), and the width and height are (b w ,b h );

[0159] The confidence screening is shown in formula (9):

[0160]

[0161] Wherein, B det and B pred represent the areas of the two boundary frames, respectively;

[0162] NMS step: for each class, sort by confidence from high to low, then iterate one by one, keep the detection frame with confidence > 0.6, and remove redundant frames by non-maximum suppression (IoU threshold 0.5).

[0163] The step 3 comprises:

[0164] Step 3.1 motion modeling:

[0165] Take the state vector X = [x, y, w, h, v x ,v y ,v w ,v h ] T , wherein (x, y) is the center position, (w, h) is the width and height of the boundary frame, (v x ,v y ) is the center point velocity, and (v w ,v h ) is the width and height change rate;

[0166] State transition equation: x k|k-1 =Fx k-1 (12);

[0167] Wherein, the state transition matrix F is:

[0168]

[0169] Δt is the time interval (the time between two frames);

[0170] Covariance update: P k|k-1 = FP k-1 F T + Q (13);

[0171] where P is the state covariance matrix, and Q is the process noise covariance matrix.

[0172] Step 3.2 Calculate the similarity between the detection frame and the predicted trajectory space by Mahalanobis distance as shown in formula (14):

[0173]

[0174] where: y j is the state of the jth detection frame, usually: [x, y, w, h] T , x i is the predicted state of the ith trajectory, H is the observation matrix, which maps the state vector to the observation space, H = [I4, 0 4x4 ], that is, only the position and size are taken, S i = HP i H T + R, R is the observation noise covariance, when the Mahalanobis distance is greater than 0.8, it is considered that the association is invalid.

[0175] Step 3.3 Data association:

[0176] The cosine distance calculation formula d (2) (i,j) = 1-f i T f j (15);

[0177] where f i (128 dimensions, ||f i ||2 = 1), since the feature vector has been normalized, the cosine distance is 1 minus the dot product of the two vectors;

[0178] Step 3.4 LSTM trajectory prediction model:

[0179] The position coordinates are normalized and linearly scaled to [-1, 1], and the velocity vector is unitized;

[0180] Let the input vector be x i , the hidden state at the last time be h t-i , and the memory cell state be C t-1

[0181] Forget gate, input gate, and output gate calculation:

[0182]

[0183] Candidate memory cell state: Updated memory cell state: Output hidden state: h t = o t tanh(C t )(19);

[0184] where σ is the sigmoid function and ⊙ is the element-wise multiplication;

[0185] Step 3.5 Future trajectory prediction:

[0186] Input: LSTM last time hidden state h t ;

[0187] Output: Future 6-frame 1.2-second trajectory coordinate sequence

[0188]

[0189] where W y and b y are the weight matrix and bias vector of the fully connected layer.

[0190] The intermediate response described in step 4 includes:

[0191] Convert the predicted trajectory of step 3 to PTZ dome coordinates (θ, ψ):

[0192] Let the image pixel coordinates be (u, v), the camera intrinsic parameters be focal length (f x , f y ), and the principal point (c x , c y )

[0193]

[0194] Convert to spherical coordinates:

[0195]

[0196] Drive the PTZ dome through PID control:

[0197] Let e(t) be the current time angle deviation;

[0198] PID control quantity:

[0199] where K p = 0.8, K i = 0.05, and K d = 0.2;

[0200] Supplement the target angle with the trajectory offset predicted by the LSTM:

[0201] Where θ,φ is the current calculated target angle, and Δθ,Δφ is the trajectory offset predicted by the LSTM.

[0202] The compensated target angle:

[0203]

[0204] Where α is the preview gain coefficient.

[0205] Embodiment 3

[0206] Based on the multi-modal perception and photoelectric linkage-based security system in Embodiment 1 and the security method in Embodiment 2, in order to facilitate understanding, the prison guard tower night duty scene is used for illustration in this embodiment:

[0207] Test environment: Rainy and foggy weather (visibility < 50 m), target moving speed 5 m / s.

[0208] System initialization:

[0209] Draw a polygon electronic fence (accuracy ± 0.5 m) in the three-dimensional geographic information system; start the dual panoramic camera scanning, and generate a distortion-free panoramic view through the adaptive light compensation algorithm (ISO 100-12800 automatic adjustment).

[0210] Target identification and tracking:

[0211] When the personnel break into the warning area, YOLOv7 detects the target (confidence > 95%), DeepSORT assigns a unique ID and extracts the motion features;

[0212] The LSTM trajectory prediction module calculates the target future position $P_{t+1.2s}$ = $f(V_t,a_t,θ_t)$, and outputs to the controller, which controls the PTZ cloud through the PID algorithm.

[0213] Photoelectric linkage disposal:

[0214] The searchlight locks the target within 0.5 seconds, starts the strong light tracking (illuminance ≥ 100,000 lux), synchronously triggers the directional sound wave alarm (frequency band 2-5 kHz, sound pressure 110 dB), and displays the AR superimposed picture (target position + threat level) on the display module of the on-duty terminal.

[0215] Emergency escalation:

[0216] If the target accelerates towards the guard tower, the optical suppression is started after the emergency response manual confirmation, the laser module is switched to the 5Hz stroboscopic mode, and the detail camera 32x zoom lens automatically records the evidence.

[0217] Results: The time from recognition to locking is 1.8s, and the tracking loss rate is less than 0.1%. It shows that the application has good use effect.

[0218] The above-described embodiments are only preferred embodiments of the present application and are not intended to limit the scope of the application. Any equivalent changes or modifications made in accordance with the structure, features and principles described in the patent scope of the present application should be included in the patent scope of the present application.

Claims

1. A security system based on multi-modal perception and photoelectric linkage, characterized in that, The environment perception module comprises a panoramic camera, a detail camera and a thermal imaging module, the panoramic camera is fixedly arranged on both sides of the detection area, the detection ranges of the two panoramic cameras cover the detection area, and the detail camera is provided with a PTZ holder, and a laser module is further arranged on the PTZ holder; The on-duty terminal comprises a display module, an audible and visual alarm module, a controller and a communication module, the controller is in communication connection with the environment perception module, the PTZ holder and the laser module through the communication module, and the display module and the audible and visual alarm module are in communication connection with the controller.

2. The security method of the security system based on multi-modal perception and photoelectric linkage according to claim 1 comprises: Step 1: a polygonal electronic fence is demarcated in a three-dimensional geographic information system, double panoramic cameras and a thermal imaging module are started to collect images, the images are subjected to adaptive light compensation by using LAB color space conversion and CLAHE algorithm, and the images are corrected in real time through a Brown-Conrady distortion model; Step 2: the corrected images are zoomed to 640*640 resolution, input into a YOLOv7 network for target detection, preprocessed by using Mosaic data enhancement, features are extracted through an E-ELAN backbone network, three scale feature maps are fused by using FPN, detection boxes with a confidence degree greater than 0.6 are output, and targets are filtered based on biological morphological parameters; Step 3: a DeepSORT algorithm is used for the detected targets, an 8-dimensional state vector is predicted through Kalman filtering, 128-dimensional appearance features extracted by using ResNet50 are combined for trajectory association, 10 continuous motion states are input into a three-layer LSTM network, and a future 1.2-second trajectory is predicted by stacking a space-time attention mechanism; Step 4: a three-level response is triggered according to target behaviors, including: Primary response: when the target invades the electronic fence, a searchlight is started to flash at a frequency of 5Hz and a directional sound wave is started at a volume of 110dB; Intermediate response: when the target stays for more than 5 seconds, the target is continuously tracked through the PTZ holder and recorded through the detail camera; Emergency response: after manual confirmation, the laser module is started to suppress the target; Step 5: the target position is displayed on the three-dimensional map of the on-duty terminal superimposed with AR light tracks, the trajectory data is encrypted and returned, and stored in an eMMC.

3. The security method based on multi-modal perception and photoelectric linkage according to claim 2, characterized in that, Step 1 comprises: Step 1.1 adaptive light compensation: the image is divided into 8*8 grids, and the gray histogram in each grid is clipped so that the number of pixels of each gray level does not exceed a threshold T: Wherein, N is the total number of pixels in the grid, and β is a limiting factor; The clipping threshold is represented by formula (2): Wherein, βl is a clipping factor, which is 2 or 3, and then the excess part is evenly distributed to all gray levels; Step 1.2 Correcting the image: the image after adaptive light compensation is set to normalized image coordinates (x, y), and the corrected coordinates are (x corrected ,y cprrected ) Wherein, the correction includes radial distortion as shown in formula (3): where r = x 2 + y 2 ; And tangential distortion as shown in formula (4):

4. The security method based on multi-modal perception and photoelectric linkage according to claim 2, characterized in that, Step 2 of inputting the YOLOv7 network for target detection comprises: Boundary box regression: as shown in formulas (5-8): b x = σ(t x )+ c x (5); b y = σ(t y )+ c y (6) wherein (c x ,c y ) is the coordinate of the upper left corner of the current grid, (p w ,p h ) is the width and height of the preset anchor frame, (t x ,t y ,t w ,t h ) is the offset of the network prediction, and σ is a sigmoid function that compresses the prediction value to the range of (0, 1) to ensure that the center point is within the current grid; the final target frame center coordinates are (b x ,b y ), and the width and height are (b w ,b h ). Confidence screening as shown in formula (9): wherein B det and B pred represent the area of two bounding boxes, respectively; NMS step: For each class, sort the bounding boxes by confidence score in descending order, then iterate through them and keep the ones with confidence score > 0.6, and remove redundant boxes by non-maximum suppression.

5. The security method based on multi-modal perception and photoelectric linkage according to claim 2, characterized in that, The biological form parameters in step 2 include: The multi-spectrum biometric recognition engine is shown in formula (10): F (l) = ReLU(W (l) * F (l-1) + b (l) )(10); where W (l) is a convolution kernel (64 3x3 convolution kernels), * denotes a convolution operation, and the ReLU activation function is ReLU(x) = max(0, x). Multiple maximum pooling operations with different scales are performed on the input feature map, such as 4x4, 8x8, and 16x16, and then all the pooling results are flattened and spliced as shown in formula (11): F spp = [MaxPool(F, 4x4), MaxPool(F, 8x8), MaxPool(F, 16x16)] (11) In each maximum pooling operation, the value of the output position (i, j) is the maximum value in the corresponding window in the input. Physical size calculation is triggered when the target body length is greater than 1.2 m and the shoulder height is greater than 0.6 m.

6. The security system and security method based on multi-modal perception and photoelectric linkage according to claim 2, characterized in that, The step 3 includes: Step 3.1 motion modeling: Take the state vector X = [x, y, w, h, v x ,v y ,v w ,v h ] T , where (x, y) is the center position, (w, h) is the bounding box width and height, (v x ,v y ) is the center point speed, (v w ,v h ) is the width and height change rate; State transition equation: x k|k-1 = Fx k-1 (12); wherein the state transition matrix F is: Δt is the time interval; Covariance update: P k|k-1 = FP k-1 F T + Q (13). wherein P is the state covariance matrix and Q is the process noise covariance matrix. Step 3.2 Calculate the similarity between the bounding box and the predicted trajectory space by Mahalanobis distance as shown in formula (14): where: y j is the state of the jth bounding box, usually: [x, y, w, h] T , x i is the predicted state of the ith track, H is the observation matrix that maps the state vector to the observation space, H = [I4, 0 4x4 ], that is, only the position and size are taken, S i = HP i H T + R, R is the observation noise covariance, when the Mahalanobis distance is greater than 0.8, it is considered that the association is invalid. Step 3.3 data association: Cosine distance calculation formula d (2) (i,j) = 1 - f i T f j (15); where f i (128 dimensions, ||f i ||2 = 1), since the feature vectors are normalized, the cosine distance is simply 1 minus the dot product of the two vectors; Step 3.4 LSTM trajectory prediction model: Position coordinates are normalized and linearly scaled to [-1, 1], and velocity vectors are unitized. Let the input vector be x i the hidden state at the previous time is h t-i the memory cell state is C t-1 Forget gate, input gate, and output gate calculation: Candidate memory cell states: Updating memory cell states: Output hidden state: h t = o t ⊙ tanh(C t )(19); wherein σ is the sigmoid function and is the element-wise multiplication. Step 3.5 future trajectory prediction: Input: LSTM last time step hidden state h t ; Output: 6-frame 1.2-second trajectory coordinate sequence where W y and b y are the weight matrix and bias vector of the fully connected layer.

7. The security method based on multi-modal perception and photoelectric linkage according to claim 2, characterized in that, The intermediate response in step 4 includes: Convert the predicted trajectory in step 3 to PTZ dome coordinates (θ, ψ): Let image pixel coordinates be (u, v), camera intrinsic parameters be focal length (f x ,f y ), and principal point (c x ,c y ) Convert to spherical coordinates: Drive the PTZ dome by PID control: Let e(t) be the current angular deviation. PID control amount: where K p = 0.8, K i = 0.05, K d = 0.2; Superimpose the trajectory offset for preview compensation: Let the current calculated target angle be θ, φ, and the LSTM predicted trajectory offset be Δθ, Δφ. The compensated target angle is: wherein α is the preview gain coefficient.