Deep learning perimeter security system and method

Through deep learning technology, combined with multimodal fusion, target morphology completion, trajectory prediction and self-supervised learning, the detection accuracy of the perimeter security system in complex wild environments is solved, and the target detection and tracking effect with high accuracy and stability is achieved.

CN120220057AInactive Publication Date: 2025-06-27CENTURY ZHONGKE (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510286616.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing perimeter security system is susceptible to occlusion in complex environments in the wild, and the detection accuracy decreases when the target information is missing. It is difficult for traditional optimization solutions to fundamentally solve the problem of monitoring blind spots.

Method used

Deep learning technology is adopted to collect environmental data through monitoring modules, and optical, thermal and infrared features are extracted using multimodal fusion method to perform object detection and morphological completion. Combining target segmentation, semantic segmentation and timing analysis methods, morphological completion and target classification of the obstructed parts are carried out. At the same time, LSTM trajectory prediction and reinforcement learning technology are used for target tracking, and false positive screening is performed by combining self-supervised learning and comparative learning methods.

Benefits of technology

In complex environments in the wild, the accuracy and stability of target detection are improved, monitoring blind spots are reduced, the system's continuous monitoring capabilities are enhanced, and the false alarm rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220057A_ABST
    Figure CN120220057A_ABST
Patent Text Reader

Abstract

The invention discloses a perimeter security system and method based on deep learning, and relates to the technical field of security, a multi-modal fusion technology is introduced in a target detection stage, and sufficient key point information can be obtained even if a target is only partially exposed through combined extraction of optical, thermal and infrared features; in the target recognition stage, the target form is speculated by adopting a target complementation technology and combining with a generative adversarial network, and the shielded part is complemented through a semantic segmentation method, so that the system completes recognition without depending on a complete target contour, and the classification stability is improved; in the target tracking stage, LSTM trajectory prediction and reinforcement learning are combined, so that the system can speculate the next motion position of the target based on a historical trajectory, and the target cannot be lost due to transient shielding; meanwhile, under the condition of long-time shielding, the system adopts a Markov decision process to optimize the probability of occurrence of the target, and the recognition rate when the target appears again is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security technology, in particular to a perimeter security system and method based on deep learning. Background Art

[0002] The perimeter security in the wild is different from that in open areas such as urban buildings and factory areas. There are a large number of obstacles such as grass, shrubs, and rocks in the wild environment, and the detection target will be blocked by natural barriers, resulting in ineffective monitoring.

[0003] Most of the existing security systems use fixed cameras combined with thermal imaging, infrared pair-shooting, or vibration optical fibers for intrusion detection. Among them, thermal imaging devices rely on the temperature difference between the target and the background for recognition. However, grass and low trees can partially block thermal radiation, making the camera only able to capture scattered temperature features. These fragmented information is difficult to form a complete target shape in traditional detection algorithms and is easily misjudged as environmental noise without threat. Visible light cameras are also difficult to play a role in a blocked environment. When the contour of the detected object is missing, it is difficult for the target detection algorithm based on edge features to recognize, which easily leads to missed reports.

[0004] To address this problem, some security systems introduce multi-spectral imaging technology to enhance the target detection ability through the fusion of visible light and infrared data. However, the resolution of thermal imaging is limited and cannot delicately depict the target contour. Visible light cameras are greatly affected by light, and their recognition effect decreases at night or in bad weather. It still cannot effectively solve the problem of recognition failure caused by missing target information in a blocked environment. Therefore, there is an urgent need for a perimeter security solution based on deep learning to solve such problems. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] The present invention provides a perimeter security system and method based on deep learning to solve the problems that the existing perimeter security system is vulnerable to occlusion in the complex wild environment, the detection accuracy decreases when target information is missing, and the traditional optimization scheme is difficult to fundamentally solve the problem of monitoring blind spots.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, an embodiment of the present invention provides a perimeter security system based on deep learning, which includes

[0009] a monitoring module, configured to collect environmental data of a monitoring area, extract features based on the environmental data, and perform target detection using the feature data;

[0010] a target processing module, configured to complement the shape of the target based on the target detection result and perform target classification;

[0011] The target processing module includes a target completion unit, a semantic speculation unit, and a target classification unit;

[0012] The target completion unit uses a target segmentation method to morphologically complete the occluded part of the target based on the target edge contour and skeleton key points;

[0013] The semantic speculation unit uses a semantic segmentation method to speculate on the morphology of the invisible part of the target according to the features of the detected visible part of the target;

[0014] The target classification unit uses a time series analysis method to classify the target in combination with the historical movement trajectory of the target;

[0015] The target tracking module tracks the target based on the target classification result and adjusts the camera angle;

[0016] The false alarm screening module screens the target detection results according to the target tracking data;

[0017] The false alarm screening module includes an environment screening unit;

[0018] The environment screening unit uses a time series modeling method combined with a self-supervised learning method to screen the environment data;

[0019] The system adjustment module is used to adjust the target detection parameters and the camera scheduling method according to the false alarm screening result.

[0020] As a preferred solution of the perimeter security system based on deep learning according to the present invention, wherein: the environment data includes visible light images, thermal imaging data, and infrared signals;

[0021] The monitoring module includes a feature extraction unit and a target detection unit;

[0022] The feature extraction unit is used to process the environment data by using a multimodal fusion method, extract optical features, thermal features, and infrared features, perform target detection based on the feature data, and obtain the visible part features of the target, including the target edge contour, skeleton key points, temperature distribution, and movement trajectory;

[0023] The target detection unit is used to perform a preliminary detection of the target according to the feature data and output the target detection result.

[0024] As a preferred solution of the perimeter security system based on deep learning according to the present invention, wherein: the target tracking module includes a trajectory prediction unit and a camera scheduling unit;

[0025] The trajectory prediction unit is used to speculate on the possible position of the target based on time series when the target is briefly occluded;

[0026] The camera scheduling unit adopts a cruise strategy and adjusts the camera angle according to the possible positions where the target may appear.

[0027] As a preferred solution of the perimeter security system based on deep learning according to the present invention, wherein: the system adjustment module includes a detection parameter adjustment unit and a scheduling optimization unit;

[0028] The detection parameter adjustment unit adjusts the target detection parameters according to the false alarm screening results;

[0029] The scheduling optimization unit is used to adjust the camera scheduling method.

[0030] In a second aspect, the present invention provides a perimeter security method based on deep learning, including,

[0031] Step S1, collect the environmental data of the monitoring area, extract features based on the environmental data, and perform target detection using the feature data;

[0032] During the feature extraction process of step S1:

[0033] Adopt a multi-modal fusion method to process the environmental data, extract optical features, thermal features and infrared features, perform target detection based on the feature data, and obtain the visible part features of the target, including the target edge contour, skeleton key points, temperature distribution and movement trajectory;

[0034] Step S2, according to the target detection result, complete the morphology of the target and perform target classification;

[0035] In step S2, adopt a target segmentation method, and based on the target edge contour and skeleton key points obtained in step S1, complete the morphology of the occluded part of the target;

[0036] In step S2, adopt a semantic segmentation method to infer the morphology of the invisible part of the target according to the visible part features of the detected target;

[0037] In step S2, adopt a time series analysis method to classify the target in combination with the historical movement trajectory of the target;

[0038] Step S3, based on the target classification result of step S2, perform target tracking on the target and adjust the camera angle;

[0039] In step S3, adopt a trajectory prediction method to infer the possible positions where the target may appear based on time series when the target is briefly occluded;

[0040] In step S3, adopt a cruise strategy and adjust the camera angle according to the possible positions where the target may appear;

[0041] Step S4: Based on the target tracking data in step S3, perform false alarm screening on the target detection results in step S1;

[0042] In step S4, a time series modeling method combined with a self-supervised learning method is used to screen the environmental data;

[0043] Step S5: According to the false alarm screening results in step S4, adjust the target detection parameters and the camera scheduling method.

[0044] As a preferred solution of the perimeter security method for deep learning described in the present invention, wherein: in step S2, the step of using the target segmentation method to morphologically complete the occluded target part based on the target edge contour and the skeleton key points obtained in step S1 is as follows,

[0045] Construct the input for target morphology completion. The input for target morphology completion consists of the target edge contour and the skeleton key points, both of which are obtained by the feature extraction unit in step S1. At the same time, the thermal imaging feature and the infrared feature are used to enhance the perception ability of the target structure;

[0046] Use the generative adversarial network GAN for target morphology completion, which includes two modules:

[0047] Generator G(x): Based on the UNet architecture, input the features of the visible part of the target and generate the complete target morphology,

[0048] and discriminator D(x): Based on the convolutional neural network CNN, discriminate the authenticity of the completed target and the real target;

[0049] The adversarial loss function of GAN is defined as:

[0050]

[0051] Among them, G(x) represents the generator, which takes the features of the partially visible target x as input and generates the complete target morphology. D(x) represents the discriminator, which is used to distinguish the real target x r and the generated target x g , D(G(x g )) represents the judgment of the discriminator on the generated target G(x g ), x r represents the real target data, which follows the real data distribution P real (x), x g represents the generated target data, which follows the generator distribution P gen (x), P real (x) represents the distribution of the real target, and P gen (x) represents the distribution of the target generated by the generator. E represents the expectation operator;

[0052] To enhance the realism of the completion target and avoid blurred completion results, L1 loss and structural similarity loss SSIM are used for optimization here, and the optimization formula is:

[0053] L rec = ||G(x) - x r ||1 + α·SSIM(G(x), x r ),

[0054] where L rec is the reconstruction loss, which measures the similarity between the generated target and the real target. ||G(x) - x r ||1 is the L1 loss, which is used to constrain the pixel-level error between the generated image and the real image. SSIM(G(x), x r ) represents the structural similarity index, and α is the weight coefficient of the SSIM loss, which controls the proportion of the L1 loss and the structural similarity loss.

[0055] As a preferred solution of the perimeter security method for deep learning described in the present invention, in step S2, the step of using the semantic segmentation method to infer the shape of the invisible part according to the detected features of the visible part of the target is as follows:

[0056] Use the Gaussian mixture model for background modeling, and its probability model is defined as:

[0057]

[0058] where P(B t |X t ) represents the probability of the background state B t under the target observation data X t at the current frame t. X t is the target observation data of the current frame, K is the number of mixture components of the Gaussian mixture model,

[0059] is the weight of the i-th Gaussian component, represents the i-th Gaussian distribution, with its mean being μ i and covariance being Σ i ;

[0060] Use the target boundary extraction method based on saliency detection to detect the possible occlusion areas of the target, and the extraction formula is:

[0061] S(x) = σ(W·f(x) + b),

[0062] Among them, S(x) is the significance prediction result, representing the importance distribution of the target in the image, F(x) is the feature transformation performed by the feature extraction network on the input data x, W is the learnable weight matrix, b is the bias term, and σ is the activation function.

[0063] As a preferred solution of the perimeter security method for deep learning described in the present invention, wherein: the core task of target trajectory prediction in step S3 is:

[0064] 1. During short-term occlusion, infer the possible positions where the target may appear based on historical data to ensure tracking continuity.

[0065] 2. When the occlusion time is long, introduce reinforcement learning and state estimation so that the target can still be effectively tracked even under long-term occlusion.

[0066] 3. When the target reappears, quickly associate the existing tracking ID to avoid misjudging it as a new target.

[0067] In step S3, the step of using the trajectory prediction method to infer the possible positions where the target may appear based on time series during short-term occlusion of the target is as follows:

[0068] Use the long short-term memory network LSTM to model the target motion trajectory, with the historical trajectory data X t as the input to predict the future target position. The prediction formula is:

[0069] h t = σ(W h h t-1 + W x X t + b h ),

[0070] wherein, h t represents the hidden state at the current time t, h t-1 represents the hidden state at the previous time t-1, X t represents the target trajectory data of the current frame, W h is the hidden layer weight matrix, W x is the input layer weight matrix, b h is the bias term, and σ is the activation function;

[0071] Perform future trajectory prediction based on the time series model. The prediction formula is:

[0072]

[0073] wherein, represents the possible position of the predicted target after t + k frames, f LSTM represents the LSTM prediction function, N represents the length of the historical trajectory window, Xt-N is the target trajectory data at time t-N.

[0074] As a preferred solution of the perimeter security method for deep learning according to the present invention, wherein: in step S3, the step of using a trajectory prediction method to infer the possible position of the target based on time series when the target is temporarily occluded further includes:

[0075] When the occlusion time is too long, a Markov decision process MDP is constructed, and the state transition relationship is defined as:

[0076] S t = F(S t-1 , A t-1 ) + ∈,

[0077] where S t represents the state of the target at time t, including position, speed and direction, S t-1 represents the state of the target at time t-1, F(·) is the state transition function, A t-1 is the action at the previous moment, and ∈ is the noise term;

[0078] A deep Q-network is used for decision optimization, and its Q-value update formula is:

[0079]

[0080] where Q(S t , A t ) represents the Q-value after taking action A t in state S t , r t is the current reward, γ is the discount factor, A′ is the possible action in the next step, and S t+1 is the predicted state at time t+1,

[0081] represents the maximum Q-value that can be obtained by taking the optimal action A′ at time t+1;

[0082] When the target reappears, quickly associate the existing tracking ID, and calculate the matching degree between the current frame target and the historical target using color histogram and feature point matching. The calculation formula is:

[0083]

[0084] where d sim represents the similarity score, H(x new ) and H(x old ) are the color histograms of the newly detected target and the historical target respectively, and ||·|| represents the L2 norm;

[0085] The target position is estimated using Kalman filtering, and the matching error is calculated. The calculation formula is as follows:

[0086]

[0087] Among them, is the estimated position of the current target, is the target estimated position of the previous frame, Z t is the observation value of the current frame, K t is the Kalman gain, H is the observation matrix, I is the identity matrix,

[0088] When the matching error is small, the target is re-associated; otherwise, a new target ID is created.

[0089] As a preferred solution of the perimeter security method based on deep learning described in the present invention, wherein: in step S4, the step of screening environmental data by using the time series modeling method in combination with the self-supervised learning method is as follows

[0090] In the false alarm screening process, the long short-term memory network LSTM is used for time series modeling to capture the changes of the target in time series. The update formula for defining the target state O t is as follows:

[0091] O t = f LSTM (O t-1 , D t , E t ),

[0092] Among them, O t represents the target state at the current time t, O t-1 represents the target state at the previous time t-1, D t represents the target detection result at the current time, E t represents the environmental data at the current time, f LSTM represents the non-linear mapping function of the LSTM network;

[0093] Calculate the anomaly score A t of the current target within multiple historical time steps, and judge whether the target is a false alarm. The calculation formula is as follows:

[0094]

[0095] Among them, A t represents the anomaly score at the current time t, T represents the length of the historical time window used to calculate the anomaly score, w i is the weight coefficient, O t-i represents the target state at time t-i, ||·|| 2 represents the Euclidean distance;

[0096] When A t exceeds the set threshold τ, the target is determined as a false alarm;

[0097] In step S4, in the absence of manual annotation, self-supervised learning automatically extracts false alarm features and optimizes them using contrastive learning to make the features of the same target similar at different times and in different environments, while the features of different targets are far apart. The contrastive loss function is defined as:

[0098]

[0099] where L contrast represents the loss function of contrastive learning, F(O t ) represents the feature representation of target O t , represents the features of target O t under different environmental conditions, that is, the positive sample, F(O - ) represents the representation vector of the negative sample O - after passing through the feature extraction network, represents calculating the feature similarity of all negative samples and summing them after exponentiation;

[0100] Based on the features optimized by self-supervised learning, the final decision formula for false alarm screening is defined as:

[0101] P false (O t ) = σ(W o ·F(O t ) + b o ),

[0102] where P false (O t ) represents the probability that target O t is determined as a false alarm, W o represents the weight matrix of the decision layer, b o represents the bias term of the decision layer, and σ represents the Sigmoid activation function;

[0103] When P false (O t ) exceeds the set threshold τ f , the determined target is a false alarm.

[0104] The beneficial effects of the present invention are as follows: In the present invention, a multi-modal fusion technology is introduced in the target detection stage. Through the joint extraction of optical, thermal, and infrared features, sufficient key point information can be obtained even if the target is only partially exposed. In the target recognition stage, a target completion technology is adopted. Combining with a generative adversarial network to speculate on the target shape and using a semantic segmentation method to complete the occluded part, the system can complete recognition without relying on the complete target contour, improving the stability of classification.

[0105] For moving targets, since shrubs and grass may cause temporary occlusion, traditional tracking methods are prone to losing the target, resulting in the system being unable to continue tracking after the target disappears briefly. In the present invention, LSTM trajectory prediction and reinforcement learning are combined in the target tracking stage, enabling the system to speculate on the next movement position of the target based on the historical trajectory, so that temporary occlusion will not cause the loss of the target. At the same time, in the case of long-term occlusion, the system adopts a Markov decision process to optimize the target appearance probability and improve the recognition rate when the target reappears. In addition, due to factors such as the movement of grass caused by the wind and accidental touches by small animals, false alarms may occur. The present invention uses a self-supervised learning combined with contrastive learning method, enabling the system to autonomously learn the characteristics of environmental noise, thereby reducing the false alarm rate and improving the false alarm screening ability.

[0106] In the present invention, in the camera scheduling stage, reinforcement learning is combined to optimize the monitoring strategy, enabling the camera to dynamically adjust the angle according to the target movement trajectory, reducing the monitoring blind area, and improving the continuous monitoring ability in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0108] Figure 1 It is a schematic framework diagram of the perimeter security system for deep learning of the present invention.

[0109] Figure 2 It is a schematic flowchart of the perimeter security method for deep learning of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0110] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.

[0111] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0112] Secondly, as used herein, an "embodiment" or "embodiments" refers to specific features, structures, or characteristics that may be included in at least one implementation of the present invention. The phrase "in an embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments.

[0113] Embodiment 1, referring to Figure 1 and Figure 2 , this embodiment provides a perimeter security system for deep learning, including:

[0114] A monitoring module, configured to collect environmental data of a monitored area, extract features based on the environmental data, and perform target detection using the feature data;

[0115] The environmental data includes visible light images, thermal imaging data, and infrared signals;

[0116] The monitoring module includes a feature extraction unit and a target detection unit;

[0117] The feature extraction unit is configured to process the environmental data using a multimodal fusion method, extract optical features, thermal features, and infrared features, perform target detection based on the feature data, and obtain visible part features of the target, including target edge contours, skeleton key points, temperature distribution, and motion trajectories;

[0118] The target detection unit is configured to perform preliminary detection on the target according to the feature data and output a target detection result;

[0119] A target processing module, which performs morphological completion on the target based on the target detection result and performs target classification;

[0120] The target processing module includes a target completion unit, a semantic inference unit, and a target classification unit;

[0121] The target completion unit uses a target segmentation method to perform morphological completion on the occluded part of the target based on the target edge contour and skeleton key points;

[0122] The semantic inference unit uses a semantic segmentation method to infer the morphology of the invisible part of the target according to the visible part features of the detected target;

[0123] The target classification unit uses a time series analysis method to classify the target in combination with the historical motion trajectory of the target;

[0124] The target tracking module tracks the target based on the target classification result and adjusts the camera angle;

[0125] The target tracking module includes a trajectory prediction unit and a camera scheduling unit;

[0126] The trajectory prediction unit is used to infer the possible position of the target based on time series when the target is briefly occluded;

[0127] The camera scheduling unit adopts a cruise strategy and adjusts the camera angle according to the possible position of the target;

[0128] The false alarm screening module screens the target detection results according to the target tracking data;

[0129] The false alarm screening module includes an environment screening unit;

[0130] The environment screening unit screens the environment data by using a time series modeling method combined with a self-supervised learning method;

[0131] The system adjustment module is used to adjust the target detection parameters and the camera scheduling method according to the false alarm screening results;

[0132] The system adjustment module includes a detection parameter adjustment unit and a scheduling optimization unit;

[0133] The detection parameter adjustment unit adjusts the target detection parameters according to the false alarm screening results;

[0134] The scheduling optimization unit is used to adjust the camera scheduling method.

[0135] This embodiment also provides a perimeter security method for the above-mentioned deep learning perimeter security system, including:

[0136] Step S1, collect the environmental data of the monitoring area, extract features based on the environmental data, and use the feature data for target detection;

[0137] During the feature extraction process of step S1:

[0138] Adopt a multi-modal fusion method to process the environmental data, extract optical features, thermal features and infrared features, perform target detection based on the feature data, and obtain the visible part features of the target, including the target edge contour, skeleton key points, temperature distribution and motion trajectory;

[0139] Step S2, according to the target detection result, complete the target morphology and classify the target;

[0140] In step S2, using the target segmentation method, based on the target edge contour and skeleton key points obtained in step S1, morphological completion is performed on the occluded target part;

[0141] In step S2, using the semantic segmentation method, based on the features of the visible part of the detected target, the morphology of its invisible part is inferred;

[0142] In step S2, using the time series analysis method, combined with the historical movement trajectory of the target, the target is classified;

[0143] In step S2, the steps of performing morphological completion on the occluded target part using the target segmentation method based on the target edge contour and skeleton key points obtained in step S1 are as follows:

[0144] Construct the input for target morphological completion. The input for target morphological completion consists of the target edge contour and skeleton key points, both of which are obtained by the feature extraction unit in step S1. At the same time, thermal imaging features and infrared features are used to enhance the perception ability of the target structure;

[0145] Use the generative adversarial network GAN for target morphological completion, which includes two modules:

[0146] Generator G(x): Based on the UNet architecture, input the features of the visible part of the target and generate the complete target morphology.

[0147] And discriminator D(x): Based on the convolutional neural network CNN, discriminate the authenticity of the completed target and the real target;

[0148] The adversarial loss function of GAN is defined as:

[0149]

[0150] Among them, G(x) represents the generator, which takes the features of the partially visible target x as input and generates the complete target morphology. D(x) represents the discriminator, which is used to distinguish the real target x r and the generated target x g , D(G(x g )) represents the judgment of the discriminator on the generated target G(x g ), x r represents the real target data, which follows the real data distribution P real (x), x g represents the generated target data, which follows the generator distribution P gen (x), P real (x) represents the distribution of the real target, and P gen (x) represents the distribution of the target generated by the generator. E represents the expectation operator;

[0151] To enhance the realism of the completion target and avoid blurred completion results, the L1 loss and the structural similarity loss SSIM are used for optimization here, and the optimization formula is:

[0152] L rec = ||G(x) - x r ||1 + α·SSIM(G(x), x r ),

[0153] where L rec is the reconstruction loss, which measures the similarity between the generated target and the real target.

[0154] ||G(x) - x r ||1 is the L1 loss, which is used to constrain the pixel-level error between the generated image and the real image.

[0155] SSIM(G(x), x r ) represents the structural similarity index, and α is the weight coefficient of the SSIM loss, which controls the proportion of the L1 loss and the structural similarity loss.

[0156] In step S2, the step of using the semantic segmentation method to infer the shape of the invisible part according to the detected features of the visible part of the target is as follows.

[0157] The Gaussian mixture model is used for background modeling, and its probability model is defined as:

[0158]

[0159] where P(B t |X t ) represents the probability of the background state B t under the target observation data X t at the current frame t, X t is the target observation data of the current frame, K is the number of mixture components of the Gaussian mixture model,

[0160] is the weight of the i-th Gaussian component, represents the i-th Gaussian distribution, whose mean is μ i , and the covariance is Σ i ;

[0161] The target boundary extraction method based on saliency detection is used to detect the possible occlusion areas of the target, and the extraction formula is:

[0162] S(x) = σ(W·f(x) + b),

[0163] Among them, S(x) is the significance prediction result, representing the importance distribution of the target in the image, f(x) is the feature transformation performed by the feature extraction network on the input data x, W is the learnable weight matrix, b is the bias term, and σ is the activation function;

[0164] Specifically, in step S2, the GAN generator predicts the target shape of the occluded part and uses the L1 loss and the structural similarity loss SSIM for optimization, making the completed shape close to the real target; secondly, the Gaussian mixture model GMM is used for background modeling to assist in inferring the possible occluded areas of the target, and at the same time, key features are extracted using saliency detection; it can effectively improve the target completion effect under conditions of partial occlusion, complex lighting, or temperature changes;

[0165] In step S3, based on the target classification result of step S2, target tracking is performed on the target, and the camera angle is adjusted;

[0166] In step S3, a trajectory prediction method is adopted. When the target is briefly occluded, the possible position where the target may appear is inferred based on time series;

[0167] In step S3, a cruising strategy is adopted to adjust the camera angle according to the possible position where the target may appear;

[0168] The core task of the target trajectory prediction in step S3 is:

[0169] 1. When briefly occluded, infer the possible position where the target may appear based on historical data to ensure tracking continuity,

[0170] 2. When the occlusion time is long, introduce reinforcement learning and state estimation to enable effective tracking of the target even under long-term occlusion conditions,

[0171] 3. When the target reappears, quickly associate the existing tracking ID to avoid misjudging it as a new target,

[0172] In step S3, the steps of inferring the possible position where the target may appear based on time series when the target is briefly occluded by adopting the trajectory prediction method are as follows:

[0173] The long short-term memory network LSTM is used to model the target motion trajectory. Using the historical trajectory data X t as the input, predict the future target position, and the prediction formula is:

[0174] h t =σ(W h h t-1 +W x X t +b h ),

[0175] Among them, h tRepresents the hidden state at the current moment t, h t-1 Represents the hidden state at the previous moment t-1, X t Represents the target trajectory data of the current frame, W h Is the hidden layer weight matrix, W x Is the input layer weight matrix, b h Is the bias term, σ is the activation function;

[0176] Based on the time series model for future trajectory prediction, the prediction formula is:

[0177]

[0178] Among them, Represents the possible position of the prediction target after t+k frames, f LSTM Represents the LSTM prediction function, N represents the length of the historical trajectory window, X t-N Is the target trajectory data at time t-N;

[0179] In step S3, when using the trajectory prediction method, the steps of inferring the possible position of the target based on time series when the target is temporarily occluded also include:

[0180] When the occlusion time is too long, a Markov decision process MDP is constructed, and the state transition relationship is defined as:

[0181] S t = F(S t-1 , A t-1 ) + ∈,

[0182] Among them, S t Represents the state of the target at time t, including position, speed and direction, S t-1 Represents the state of the target at time t-1, F(·) is the state transition function, A t-1 Is the action at the previous moment, ∈ is the noise term;

[0183] Use the deep Q network for decision optimization, and its Q value update formula is:

[0184]

[0185] Among them, Q(S t , A t ) represents the Q value after taking action A t in state S t , r t Is the current reward, γ is the discount factor, A' is the possible next action, S t+1 Is the predicted state at time t+1,

[0186] Represents the maximum Q value that can be obtained by taking the optimal action A′ at time t+1;

[0187] When the target reappears, quickly associate the existing tracking ID, and calculate the matching degree between the current frame target and the historical target using color histogram and feature point matching. The calculation formula is:

[0188]

[0189] where d sim represents the similarity score, H(x new ) and H(x old ) are the color histograms of the newly detected target and the historical target respectively, and ||·|| represents the L2 norm;

[0190] Use Kalman filter to estimate the target position and calculate the matching error. The calculation formula is:

[0191]

[0192] where is the estimated position of the current target, is the target estimated position of the previous frame, Z t is the observation value of the current frame, K t is the Kalman gain, H is the observation matrix, I is the identity matrix,

[0193] When the matching error is small, the target is re-associated, otherwise a new target ID is created;

[0194] Specifically, here the LSTM long short-term memory network is used for target trajectory modeling, predicting the future position of the target through historical trajectory data, and combining with Kalman filter for error optimization; the Markov decision process MDP is combined with the deep Q network DQN for reinforcement learning to enable the system to maintain tracking stability under long-term occlusion; in addition, when the target reappears, ID association is performed through color histogram matching and trajectory prediction error to avoid misjudgment as a new target; effectively improving the target tracking continuity of the system under occlusion, fast movement or environmental interference;

[0195] Step S4, based on the target tracking data in step S3, perform false alarm screening on the target detection results in step S1;

[0196] In step S4, a time series modeling method is combined with a self-supervised learning method to screen the environmental data;

[0197] In step S4, the steps of screening the environmental data by combining a time series modeling method with a self-supervised learning method are,

[0198] During the false alarm screening process, a long short-term memory network (LSTM) is used for temporal modeling to capture the temporal changes of the target, and the target state O is defined. t The update formula of

[0199] O t is: LSTM (O t-1 , D t , E t ),

[0200] where O t represents the target state at the current time step t, O t-1 represents the target state at the previous time step t - 1, D t represents the target detection result at the current time, and E t represents the environmental data at the current time, and f LSTM represents the non-linear mapping function of the LSTM network;

[0201] Calculate the anomaly score A t of the current target within multiple historical time steps, and determine whether the target is a false alarm. The calculation formula is:

[0202]

[0203] where A t represents the anomaly score at the current time step t, T represents the length of the historical time window used to calculate the anomaly score, w i is the weight coefficient, O t-i represents the target state at time t - i, and ||·|| 2 represents the Euclidean distance;

[0204] When A t exceeds the set threshold τ, the target is determined to be a false alarm;

[0205] In step S4, in the absence of manual annotation, self-supervised learning automatically extracts false alarm features and uses contrastive learning for optimization to make the features of the same target similar at different times and in different environments, while the features of different targets are far away. The contrastive loss function is defined as:

[0206]

[0207] where L contrast represents the loss function of contrastive learning, F(O t ) represents the feature representation of the target O t , represents the feature of the target O t under different environmental conditions, that is, the positive sample, and F(O - ) represents the negative sample O -The representation vector after the feature extraction network, It represents calculating the feature similarity for all negative samples and summing them after exponentiation;

[0208] Based on the features optimized by self-supervised learning, the final decision formula for false alarm screening is defined as:

[0209] P false (O t ) = σ(W o ·F(O t ) + b o ),

[0210] where P false (O t ) represents the probability that the target O t is determined to be a false alarm, W o represents the weight matrix of the decision layer, b o represents the bias term of the decision layer, and σ represents the Sigmoid activation function;

[0211] When P false (O t ) exceeds the set threshold τ f , the determined target is a false alarm;

[0212] Specifically, here LSTM is used to model the target state, calculate the anomaly score to judge whether the target is a false alarm; contrastive learning is used for self-supervised training to keep the features of the same target consistent under different environmental conditions, while false alarm targets are far from the correct targets; finally, the false alarm probability is calculated based on the optimized features and a decision is made through the threshold; it can adaptively learn false alarm patterns and reduce detection errors caused by environmental factors;

[0213] Step S5, according to the false alarm screening result of step S4, adjust the target detection parameters and the camera scheduling method.

[0214] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A deep learning perimeter security system, characterized by: include, The monitoring module is used to collect environmental data of the monitoring area, extract features based on the environmental data, and use the feature data to detect targets; The target processing module completes the target's morphology based on the target detection results and performs target classification; The target processing module includes a target completion unit, a semantic inference unit and a target classification unit; The target completion unit uses the target segmentation method to complete the morphology of the obscured target part based on the target edge contour and skeleton key points; The semantic inference unit uses semantic segmentation to infer the shape of the invisible part based on the features of the visible part of the detected target; The target classification unit uses a time series analysis method and combines the historical movement trajectory of the target to classify the target; The target tracking module tracks the target based on the target classification results and adjusts the camera angle; The false alarm screening module screens the false alarms of the target detection results according to the target tracking data; The false alarm screening module includes an environmental screening unit; The environmental screening unit uses a time series modeling method combined with a self-supervised learning method to screen environmental data; The system adjustment module is used to adjust the target detection parameters and the camera scheduling method according to the false alarm screening results.

2. A deep learning perimeter security system as claimed in claim 1, characterized in that: The environmental data includes visible light images, thermal imaging data and infrared signals; The monitoring module includes a feature extraction unit and a target detection unit; A feature extraction unit is used to process environmental data using a multimodal fusion method, extract optical features, thermal features, and infrared features, perform target detection based on feature data, and obtain features of the visible part of the target, including target edge contours, skeleton key points, temperature distribution, and motion trajectory; The target detection unit is used to perform preliminary detection on the target based on the feature data and output the target detection result.

3. A deep learning perimeter security system as claimed in claim 2, characterized in that: The target tracking module includes a trajectory prediction unit and a camera scheduling unit; The trajectory prediction unit is used to estimate the possible location of the target based on the time sequence when the target is temporarily occluded; The camera scheduling unit adopts a cruising strategy to adjust the camera angle according to the possible location of the target.

4. A deep learning perimeter security system as claimed in claim 3, characterized in that: The system adjustment module includes a detection parameter adjustment unit and a scheduling optimization unit; A detection parameter adjustment unit, which adjusts target detection parameters according to false alarm screening results; Scheduling optimization unit, used to adjust the camera scheduling method.

5. A deep learning perimeter security method, based on a deep learning perimeter security system according to any one of claims 1 to 4, characterized in that: include: Step S1, collecting environmental data of the monitoring area, performing feature extraction based on the environmental data, and performing target detection using the feature data; In the feature extraction process of step S1: The environmental data is processed by a multimodal fusion method to extract optical features, thermal features and infrared features, and target detection is performed based on the feature data to obtain the visible part features of the target, including the target edge contour, skeleton key points, temperature distribution and motion trajectory; Step S2, according to the target detection result, the target is morphologically completed and the target is classified; In step S2, a target segmentation method is used to complete the morphology of the obscured target part based on the target edge contour and skeleton key points obtained in step S1; In step S2, a semantic segmentation method is used to infer the shape of the invisible part of the target based on the features of the visible part of the detected target; In step S2, a time series analysis method is used to classify the target in combination with the historical movement trajectory of the target; Step S3, based on the target classification result of step S2, tracking the target and adjusting the camera angle; In step S3, a trajectory prediction method is used to infer the possible position of the target based on the time sequence when the target is temporarily occluded; In step S3, a cruising strategy is adopted to adjust the camera angle according to the possible location of the target; Step S4, based on the target tracking data of step S3, screening the target detection result of step S1 for false alarms; In step S4, the environmental data is screened by using a time series modeling method combined with a self-supervised learning method; Step S5, adjusting the target detection parameters and the camera scheduling method according to the false alarm screening result of step S4.

6. A perimeter security method based on deep learning as claimed in claim 5, characterized in that: In step S2, the step of using the target segmentation method to complete the morphology of the obscured target part based on the target edge contour and skeleton key points obtained in step S1 is as follows: Constructing the input of target morphology completion, the input of target morphology completion consists of the target edge contour and skeleton key points, both of which are obtained by the feature extraction unit in step S1, and at the same time, using thermal imaging features and infrared features to enhance the perception ability of the target structure; The generative adversarial network GAN is used to complete the target morphology, which includes two modules: Generator G(x): Based on the UNet architecture, it inputs the visible features of the target and generates the complete target shape. And discriminator D(x): based on convolutional neural network CNN, it discriminates the authenticity of the completed target and the real target; The adversarial loss function of GAN is defined as: Among them, G(x) represents the generator, which takes the partially visible target feature x as input to generate the complete target shape, and D(x) represents the discriminator, which is used to distinguish the real target x r and generate the target x g , D(G(x g )) represents the discriminator's generation of the target G(x g )’s judgment, x r Represents the real target data, which obeys the real data distribution P real (x), x g Represents the generated target data, which obeys the generator distribution P gen (x), P real (x) represents the distribution of the true target, P gen (x) represents the distribution of the target generated by the generator, and E represents the expectation operator; L1 loss and structural similarity loss SSIM are used for optimization, and the optimization formula is: L rec =||G(x)-x r ||1+α·SSIM(G(x),x r ), Among them, L rec is the reconstruction loss, which measures the similarity between the generated target and the real target. ||G(x)-x r ||1 is the L1 loss, which is used to constrain the pixel-level error between the generated image and the real image. SSIM(G(x),x r ) represents the structural similarity index, α is the weight coefficient of SSIM loss, which controls the ratio of L1 loss to structural similarity loss.

7. A perimeter security method based on deep learning as claimed in claim 6, characterized in that: In step S2, the step of using the semantic segmentation method to infer the shape of the invisible part of the target based on the features of the visible part of the detected target is as follows: The Gaussian mixture model is used for background modeling, and its probability model is defined as: Among them, P(B t |X t ) indicates that at the current frame t, the background state B t In the target observation data X t The probability of X t is the target observation data of the current frame, K is the number of mixed components of the Gaussian mixture model, is the weight of the i-th Gaussian component, represents the i-th Gaussian distribution, whose mean is μ i , the covariance is Σ i ; The target boundary extraction method based on saliency detection is used to detect the possible occlusion area of ​​the target. The extraction formula is: S(x)=σ(W·f(x)+b), Among them, S(x) is the saliency prediction result, representing the importance distribution of the target in the image, f(x) is the feature transformation of the input data x by the feature extraction network, W is the learnable weight matrix, b is the bias term, and σ is the activation function.

8. A perimeter security method based on deep learning as claimed in claim 7, characterized in that: In step S3, the step of using the trajectory prediction method to infer the possible position of the target based on the time sequence when the target is temporarily blocked is: The long short-term memory network LSTM is used to model the target motion trajectory, with the historical trajectory data X t As input, predict the future target position, the prediction formula is: h t =σ(W h h t-1 +W x X t +b h ), Among them, h t represents the hidden state at the current time t, h t-1 represents the hidden state at the previous moment t-1, X t Represents the target trajectory data of the current frame, W h is the hidden layer weight matrix, W x is the input layer weight matrix, b h is the bias term, σ is the activation function; Based on the time series model, the future trajectory is predicted, and the prediction formula is: in, represents the possible position of the predicted target after t+k frames, f LSTM represents the LSTM prediction function, N represents the historical trajectory window length, X t-N is the target trajectory data at time tN.

9. A perimeter security method based on deep learning as claimed in claim 8, characterized in that: In step S3, the step of using the trajectory prediction method to estimate the possible position of the target based on the time sequence when the target is temporarily blocked also includes: When the occlusion time is too long, a Markov decision process MDP is constructed and the state transition relationship is defined as: S t =F(S t-1 ,A t-1 )+∈, Among them, S t Represents the state of the target at time t, including position, speed and direction, S t-1 represents the state of the target at time t-1, F(·) is the state transition function, A t-1 is the action at the previous moment, ∈ is the noise term; A deep Q network is used for decision optimization, and its Q value update formula is: Among them, Q(S t ,A t ) represents the state S t Take action A t The Q value after t is the current reward, γ is the discount factor, A′ is the possible next action, S t+1 is the predicted state at time t+1, represents the maximum Q value that can be obtained by taking the optimal action A′ at time t+1; When the target reappears, quickly associate it with the existing tracking ID, and use the color histogram and feature point matching to calculate the matching degree between the current frame target and the historical target. The calculation formula is: Among them, d sim represents the similarity score, H(x new ) and H(x old ) are the color histograms of the new detected target and the historical target, respectively, and ||·|| represents the L2 norm; Kalman filtering is used to estimate the target position and calculate the matching error. The calculation formula is: in, is the estimated position of the current target, The estimated position of the target in the previous frame, Z t is the observation value of the current frame, K t is the Kalman gain, H is the observation matrix, I is the identity matrix, When the matching error is small, the target is reassociated, otherwise a new target ID is created.

10. A perimeter security method based on deep learning as claimed in claim 9, characterized in that: In step S4, the step of using the time series modeling method combined with the self-supervised learning method to screen the environmental data is: In the process of false alarm screening, the long short-term memory network LSTM is used for time series modeling to capture the changes in the target in time series and define the target state O t The update formula is: O t =f LSTM (O t-1 ,D t ,E t ), Among them, O t represents the target state at the current time t, O t-1 represents the target state at the previous time t-1, D t represents the target detection result at the current moment, E t represents the current environmental data, f LSTM Represents the nonlinear mapping function of the LSTM network; Calculate the anomaly score A of the current target in multiple historical time steps t , to determine whether the target is a false alarm, the calculation formula is: Among them, A t represents the anomaly score at the current time t, T represents the length of the historical time window used to calculate the anomaly score, and w i is the weight coefficient, O t-i represents the target state at time ti, ||·|| 2 represents the Euclidean distance; When A t When it exceeds the set threshold τ, the target is judged as a false alarm; In step S4, in the absence of manual annotation, self-supervised learning automatically extracts false positive features and uses contrastive learning for optimization, so that the features of the same target at different times and environments remain similar, while the features of different targets are far apart. The contrastive loss function is defined as: Among them, L contrast Denotes the loss function of contrastive learning, F(O t ) indicates the target O t The characteristic representation of Indicates the target O t Features under different environmental conditions, that is, positive samples, F(O - ) represents the negative sample O - The representation vector after the feature extraction network, Indicates that the feature similarity of all negative samples is calculated and summed after exponentialization; Based on the features optimized by self-supervised learning, the final decision formula for false positive screening is defined as: P false (The t )=σ(W o ·F(O t )+b o ), Among them, P false (O t ) indicates the target O t The probability of being judged as a false alarm, W o represents the weight matrix of the decision layer, b o represents the bias term of the decision layer, and σ represents the Sigmoid activation function; When P false (O t ) exceeds the set threshold τ f , the target is judged as a false alarm.

Citation Information

Cited By

  • Dangerous behavior feature identification method for Internet of Things security of smart community

    CN120997780A