Intelligent security video monitoring system for preventing animal attack behaviors

Through synchronous acquisition and shared feature extraction technology, the YOLOv7 detection head and dynamic threat assessment are improved, and the problem of switching delay and timestamp synchronization error of intelligent security systems in low light is solved, improving the detection accuracy of animal attack behavior and threat assessment accuracy.

CN120236243AActive Publication Date: 2025-07-01延安大学西安创新学院

Patent Information

Application Number
CN202510703512.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The existing intelligent security system reduces the fusion accuracy caused by switching delay and cross-modal timestamp synchronization error under low light conditions, ignores the complementarity of the frequency domain phase spectrum and high-frequency texture, and calculates redundancy and feature confusion.

Method used

By synchronously collecting thermal radiation image sequences and RGB image sequences, cross-modal video stream data packets are generated, cross-modal shared features are extracted using shared weight convolution layer, and multi-scale candidate boxes are improved for YOLOv7 detection head generation, and dynamic threat evaluation is combined with Kalman filtering and LSTM-TCN model to trigger the hierarchical response mechanism.

Benefits of technology

The time stamp alignment of cross-modal data in dynamic environments is achieved, which improves the accuracy of animal target detection and the accuracy of threat assessment, reduces missed detection under low light, and improves the system's all-weather monitoring capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236243A_ABST
    Figure CN120236243A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent security and protection video monitoring system for preventing animal attack behaviors, and relates to the technical field of intelligent security and protection, and the system comprises the steps: synchronously collecting a thermal radiation image sequence and an RGB image sequence, and generating a cross-modal video stream data packet; performing multispectral feature extraction on the cross-modal video stream data packet to generate a visible light enhancement feature and an infrared specific feature; cross-modal shared features are extracted from the visible light enhancement features and the infrared specific features through a shared weight convolution layer, and a mixed feature tensor is formed through splicing; horizontally dividing the mixed feature tensor into equal-width regions, calculating region pseudo anchor points, generating a multi-scale candidate frame by improving a YOLOv7 detection head, guiding similar aggregation and heterogeneous rejection loss by the pseudo anchor points, and outputting a high-confidence candidate frame set; and converting bounding box parameters of each candidate box in the high-confidence candidate box set into a state vector. According to the method, the alignment of the time stamps of the cross-modal data is ensured by dynamically adjusting the primary and secondary data streams.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent security, and in particular to an intelligent security video monitoring system for preventing animal attack behaviors. Background Art

[0002] In recent years, intelligent security systems have gradually adopted multi-modal sensing technologies in scenarios such as wildlife protection and border monitoring, improving all-weather monitoring capabilities by fusing visible light and infrared imaging. Visible light cameras obtain RGB images through CMOS sensors and perform real-time processing in combination with ISP; infrared cameras rely on microbolometers to capture thermal radiation information, and the two switch through independent data streams at different times to adapt to light changes. Through multi-spectral feature extraction methods, such as visible light enhancement technology based on frequency domain phase spectrum and infrared high-frequency feature extraction using wavelet transform, multi-scale object detection is achieved in combination with deep learning models. In addition, dynamic threat assessment usually adopts single-modal trajectory analysis or pose estimation, such as the combination scheme of OpenPose key point detection and LSTM time series modeling.

[0003] However, the existing technologies still have significant defects: existing systems rely on fixed thresholds or manually set light intensity to switch the primary and secondary data streams, resulting in switching delays when the signal-to-noise ratio of visible light drops sharply under low light, and the cross-modal timestamp synchronization error is not solved, affecting subsequent fusion accuracy. Traditional methods directly splice infrared and visible light features, ignoring the complementarity of frequency domain phase spectrum and high-frequency texture, and do not design a lightweight shared weight convolutional architecture, resulting in computational redundancy and feature confusion. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an intelligent security video monitoring system for preventing animal attack behaviors to solve the problem of switching delay when the signal-to-noise ratio of visible light drops sharply under low light in existing intelligent security systems in a dynamic environment.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: The present invention provides an intelligent security video monitoring system for preventing animal attack behaviors, which includes a synchronous acquisition module that synchronously acquires a thermal radiation image sequence and an RGB image sequence to generate a cross-modal video stream data packet; A feature extraction module that performs multi-spectral feature extraction on the cross-modal video stream data packet to generate visible light enhancement features and infrared specific features; A feature fusion module that extracts cross-modal shared features from the visible light enhancement features and the infrared specific features through a shared weight convolutional layer and splices them to form a mixed feature tensor; The detection and tracking module horizontally divides the hybrid feature tensor into equally wide regions, calculates regional pseudo-anchors, and generates multi-scale candidate boxes through an improved YOLOv7 detection head. The pseudo-anchors guide the aggregation of the same class and the repulsion loss of different classes, and output a set of candidate boxes with high confidence; The evaluation and analysis module converts the bounding box parameters of each candidate box in the set of candidate boxes with high confidence into state vectors, and combines Kalman filtering and cross-frame feature association to generate a list of animal targets; The hierarchical response module extracts the motion trajectory from the list of animal targets, combines OpenPose key point detection, and fuses historical behavior features to input into the LSTM-TCN hybrid model, outputs the attack probability and maps it to the dynamic threat level, and triggers the hierarchical response mechanism.

[0007] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the synchronous acquisition of the thermal radiation image sequence and the RGB image sequence includes the following steps, Detect the ambient light intensity through the light intensity sensor of the visible light camera. When it is higher than the average light intensity value, start RGB image acquisition. When it is lower than the average light intensity value, start infrared thermal imaging acquisition; The main control chip synchronously sends VSYNC vertical synchronization pulses, aligns the timestamps of the thermal radiation image sequence and the RGB image sequence through the FPGA timer, and compresses them into independent infrared video streams and visible light video streams; Dynamically switch the primary and secondary data streams according to the signal-to-noise ratio detection result of the visible light video stream, and generate cross-modal video stream data packets.

[0008] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the multi-spectral feature extraction includes the following steps, Separate the three-channel matrix of the RGB image, generate a complex frequency domain matrix through two-dimensional fast Fourier transform and Hanning window function, extract the phase spectrum components of each channel and superimpose them with the three-channel grayscale matrix to generate visible light enhanced features; Perform three-level discrete wavelet transform on the thermal radiation image, decompose it to the third-level diagonal high-frequency sub-band through Haar wavelet basis, and generate infrared specific features through Sobel gradient calculation and non-maximum suppression.

[0009] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the extraction of cross-modal shared features includes the following steps, Stitch the three-channel data of the visible light enhanced features and the single-channel data of the infrared specific features into a four-channel input matrix; Perform downsampling of large convolutional kernel features, max pooling compression, and enhancement of small convolutional kernel features in sequence through a shared-weight convolutional layer, and output cross-modal shared features.

[0010] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the steps for splicing to form the hybrid feature tensor include the following. Uniformly divide the spatial dimension of the cross-modal shared features into grid regions, perform global average pooling on each grid to generate local feature descriptors, and compress them into local feature matrices through a fully connected layer. Adjust the number of channels of the infrared specific features and the visible light enhanced features to be consistent with the spatial resolution of the local feature matrix. Splice the local feature matrix, the adjusted infrared specific features, and the visible light enhanced features along the channel dimension to form a hybrid feature tensor.

[0011] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the steps for generating multi-scale candidate boxes by improving the YOLOv7 detection head include the following. Divide the hybrid feature tensor into non-overlapping regions of equal width, extract the average values of the infrared and visible light feature vectors within the regions to generate a set of pseudo anchor points. Input the set of pseudo anchor points into the improved YOLOv7 detection head for multi-scale feature map convolution processing to generate multi-scale candidate boxes.

[0012] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the pseudo anchor-guided homogeneous aggregation and heterogeneous exclusion loss refers to performing homogeneous aggregation constraints on the infrared specific features of the candidate boxes to drive homogeneous features to gather towards the pseudo anchor points, and performing heterogeneous exclusion constraints on the left and right adjacent blocks of the block where the candidate boxes are located to suppress cross-block feature confusion.

[0013] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the steps for generating an animal target list include the following. Convert the bounding box parameters of each candidate box in the high-confidence candidate box set into state vectors, and associate the infrared pseudo anchor features and the visible light enhanced pseudo anchor features. Predict the current target position based on the state vector, and update the state vector through Mahalanobis distance matching to generate a predicted box after motion correction. Perform cross-frame association on the corrected predicted box with the infrared pseudo anchor features and the visible light enhanced pseudo anchor features, calculate the spatio-temporal-feature comprehensive similarity, determine the continuously unmatched targets as leaving the monitoring area and remove them from the tracking list to generate an animal target list.

[0014] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the steps of outputting the attack probability and mapping it to the dynamic threat level include the following, Based on the angular deviation between the movement direction angle in the animal target list and the protected area boundary direction angle, the head height difference and limb opening degree detected by OpenPose, a multi-dimensional time series vector is formed in combination with information entropy; The multi-dimensional time series vector is input into the LSTM-TCN hybrid model to output the attack probability, and the low, medium, and high threat levels are divided according to the proportional points.

[0015] As a preferred solution of the intelligent security video monitoring system for preventing animal attack behaviors described in the present invention, wherein: the triggered hierarchical response mechanism refers to superimposing a dynamic tracking frame through the AR rendering engine and pushing it to the monitoring terminal in the low threat level, triggering an audible and visual alarm and sending an alarm message to the SMS gateway in the medium threat level, and activating drone cruising and ultrasonic interference based on trajectory prediction in the high threat level.

[0016] The beneficial effects of the present invention are as follows: The primary and secondary data streams are dynamically adjusted based on the signal-to-noise ratio detection of the edge computing node, combined with the VSYNC pulse synchronization and clock calibration of the FPGA timer, ensuring the time stamp alignment error of cross-modal data and avoiding the problem of missed detection in low light caused by the fixed threshold switching; The visible light frequency domain phase spectrum retains high-frequency texture details, the infrared three-level wavelet decomposition extracts the diagonal high-frequency sub-band, generates sparse features through the Sobel gradient amplitude and neighborhood non-maximum suppression, and forms a complementary feature matrix after bilinear interpolation alignment, improving the target discrimination in complex backgrounds; The hybrid feature tensor is horizontally divided into equal-width regions, the infrared and visible light feature means are calculated to generate a set of pseudo anchor points, and the YOLOv7 detection head is improved through the same-class aggregation and different-class exclusion loss functions, driving the same-class target features to gather towards the pseudo anchor points and suppressing the cross-block interference, improving the animal target detection accuracy. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0018] Figure 1 It is a system diagram of the intelligent security video monitoring system for preventing animal attack behaviors.

[0019] Figure 2 It is a schematic diagram of the synchronous acquisition module.

[0020] Figure 3It is a schematic diagram of feature fusion and detection tracking.

[0021] Figure 4 It is a schematic diagram of the hierarchical response mechanism. Specific implementation manners

[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0023] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0025] Refer to Figures 1 to 4 , which is an embodiment of the present invention. This embodiment provides an intelligent security video monitoring system for preventing animal attack behaviors, including the following steps: A synchronous acquisition module that synchronously acquires a thermal radiation image sequence and an RGB image sequence to generate a cross-modal video stream data packet.

[0026] Specifically, the power interface of the infrared camera is connected to direct current power supply, and the power interface of the visible light camera is connected to the same specification power supply; the control bus of the infrared camera sends an initialization instruction, and the control bus of the visible light camera receives the initialization instruction to complete device self-check and sensor preheating.

[0027] The CMOS sensor of the visible light camera outputs a brightness histogram in real time and calculates the average illumination intensity value of the scene; when the illumination intensity sensor of the visible light camera detects that the ambient illumination intensity reaches or exceeds the average illumination intensity value, the image signal processor (ISP) of the visible light camera starts to collect RGB images in Bayer format; when the illumination intensity sensor of the visible light camera detects that the ambient illumination intensity is lower than the average illumination intensity value, the uncooled microbolometer of the infrared camera starts the thermal imaging acquisition process, and the working band covers the mid- and far-infrared spectral ranges; the main control chip sends VSYNC vertical synchronization pulses to both the infrared camera and the visible light camera simultaneously; the FPGA timers of the infrared camera and the visible light camera obtain the thermal radiation image sequence and the RGB image sequence at the same time when receiving the rising edge of the VSYNC pulse, record the local timestamps, and align the timestamps through clock calibration.

[0028] Compress the thermal radiation image sequence with aligned timestamps into an infrared video stream, and compress the RGB image sequence into an independent visible light video stream through an encoder; bind the infrared video stream and the visible light video stream to the edge computing node respectively, decode the visible light video stream through a video decoder, and extract the brightness components of each frame of RGB image; based on the histogram distribution of the brightness components, use the signal-to-noise ratio detection algorithm to calculate the signal-to-noise ratio of the current frame, and the calculation formula is: ; where, is the signal-to-noise ratio, is the pixel mean value of the image brightness component, is the noise standard deviation of the image brightness component.

[0029] The logic controller of the edge computing node continuously receives the signal-to-noise ratio calculation results of the visible light video stream; when the consecutive signal-to-noise ratios are all lower than the signal-to-noise ratio critical value, the logic controller sends control instructions to the infrared camera and the visible light camera, marks the thermal radiation image stream of the infrared camera as the main data stream, and marks the RGB image stream of the visible light camera as the auxiliary data stream; add timestamp synchronization marks to the main data stream and the auxiliary data stream, and perform encapsulation to generate a cross-modal video stream data packet.

[0030] Furthermore, the cross-modal data stream includes device ID encoding, the measured value of the current ambient illumination intensity, and the main and auxiliary data stream marking identifiers.

[0031] The signal-to-noise ratio critical value is set according to the noise standard deviation of the brightness component of the human eye. Here, the set standard deviation is 0.03, then the signal-to-noise ratio is approximately equal to 10.4, but in actuality, a safety margin needs to be reserved and it is specifically set in combination with the quantization error of the edge computing node.

[0032] The feature extraction module performs multispectral feature extraction on the cross-modal video stream data packet to generate visible light enhanced features and infrared specific features.

[0033] Specifically, restore the main data stream in the cross-modal video stream data packet to a thermal radiation image sequence, and restore the auxiliary data stream to an RGB image sequence; for a single-frame RGB image generated by the visible light video stream, separate the R-channel matrix, G-channel matrix, and B-channel matrix; calculate the three-channel grayscale image pixel by pixel according to the standard luminance coefficient to generate a three-channel grayscale matrix of the same size as the original RGB image.

[0034] Perform spatial alignment verification on the three-channel grayscale matrix and the original RGB image; perform two-dimensional fast Fourier transform on the R, G, and B channels of the original RGB image after spatial alignment verification respectively, apply the Hanning window function to each color channel matrix to suppress spectral leakage, and obtain the complex frequency domain representation of each channel through the fast Fourier transform respectively. Integrate and output the complex frequency domain matrices of the R, G, and B channels; extract the phase spectrum components from the complex frequency domain matrices, calculate the phase spectrum of each channel respectively. Taking the R channel as an example, the expression is: ; where, is the phase spectrum matrix of the red channel indicating the frequency domain phase information of the red component in the original RGB image, used to retain the high-frequency texture features of the image, is the imaginary part of the Fourier transform result of the red channel indicating the orthogonal component (sine component) of the red component in the frequency domain, which together with the real part constitutes the complex frequency domain representation, is the real part of the Fourier transform result of the red channel

[0035] indicating the in-phase component (cosine component) of the red component in the frequency domain, which together with the imaginary part describes the frequency domain characteristics, and is the complex frequency domain matrix of the red channel after two-dimensional fast Fourier transform (FFT), which converts the red pixel values in the spatial domain into the frequency domain representation, including amplitude and phase information.

[0036] Perform three-level discrete wavelet transform on a single-frame thermal radiation image in the thermal radiation image sequence. The specific operations are as follows: The first-level decomposition uses the Haar wavelet basis to perform row filtering and column filtering on a single-frame thermal radiation image, generating a first-level low-frequency subband, a horizontal high-frequency subband, a vertical high-frequency subband, and a diagonal high-frequency subband; The second-level decomposition repeats the row filtering and column filtering operations on the low-frequency subband obtained from the first-level decomposition, generating a second-level low-frequency subband and its corresponding horizontal, vertical, and diagonal high-frequency subbands; The third-level decomposition further performs row filtering and column filtering on the low-frequency subband obtained from the second-level decomposition, finally outputting a third-level low-frequency subband and its corresponding horizontal, vertical, and diagonal high-frequency subbands.

[0037] Select the diagonal high-frequency subband of the third-level decomposition as the highest-frequency component, enhance the coefficient significance through absolute value operation; perform Sobel operator convolution on the diagonal high-frequency subband after absolute value operation, extract the horizontal direction gradient and the vertical direction gradient, generating the gradient magnitude; perform neighborhood comparison on the gradient magnitude along the gradient direction, perform non-maximum suppression and binarization, generating a binarized sparse feature point matrix; upsample the binarized sparse feature point matrix through the bilinear interpolation algorithm, generating an infrared-specific feature of the same size as the visible light enhanced feature.

[0038] Further explanation, the specific steps of performing neighborhood comparison on the gradient magnitude along the gradient direction are as follows: For the 0° direction, compare the magnitude of the current pixel with the magnitudes of the left and right adjacent pixels; for the 45° direction, compare the magnitude of the current pixel with the magnitudes of the upper-left and lower-right diagonal pixels; for the 90° direction, compare the magnitude of the current pixel with the magnitudes of the upper and lower adjacent pixels; for the 135° direction, compare the magnitude of the current pixel with the magnitudes of the upper-right and lower-left diagonal pixels, retain the local maximum points with magnitudes greater than the adjacent pixels, and set the rest to zero.

[0039] Feature fusion module, the visible light enhanced feature and the infrared-specific feature extract cross-modal shared features through a shared-weight convolutional layer, and are concatenated to form a hybrid feature tensor.

[0040] Specifically, merge the visible light enhanced feature and the infrared-specific feature according to the channel dimension, and splice the three-channel data of the visible light enhanced feature and the single-channel data of the infrared-specific feature into a four-channel input matrix; apply an initial convolutional layer with shared weights to perform dimensionality reduction on the four-channel input matrix, the first layer uses a large convolutional kernel to perform feature downsampling on the four-channel input matrix, reducing the spatial dimension and extracting basic features; the second layer further compresses the size of the basic feature map through max pooling, keeping the number of channels unchanged; the third layer applies a small convolutional kernel to adjust the channel dimension to achieve feature enhancement, and outputs cross-modal shared features.

[0041] Check whether the spatial resolution of the cross-modal shared features is consistent with the infrared-specific features and the visible-light enhanced features. If there are dimensional differences, adjust them to the same size through bilinear interpolation to ensure that all feature maps are strictly aligned in the height and width dimensions.

[0042] Evenly divide the spatial dimension of the cross-modal shared features into grid regions, where each grid covers the pixels in the height direction and the width direction; perform global average pooling on the cross-modal shared features within each grid region to compress them into local feature descriptors; apply a fully connected layer to each local feature descriptor to compress the feature dimension and generate a local feature matrix.

[0043] The infrared-specific features and the visible-light enhanced features respectively adjust the number of channels through convolutional kernels to keep the spatial resolution consistent with the local feature matrix; concatenate the local feature matrix, the adjusted infrared-specific features, and the visible-light enhanced features along the channel dimension to form a mixed feature tensor.

[0044] The detection and tracking module horizontally divides the mixed feature tensor into equally wide regions, calculates the regional pseudo-anchors, and generates multi-scale candidate boxes through an improved YOLOv7 detection head. The pseudo-anchors guide the loss of homogeneous aggregation and heterogeneous repulsion, and output a set of candidate boxes with high confidence.

[0045] Specifically, evenly divide the mixed feature tensor horizontally into equally wide regions, where each region covers the complete vertical range and has the same width, ensuring that the region division has no overlap and completely covers; within each divided region, extract all sets of feature vectors at the corresponding spatial positions from the infrared-specific features and the visible-light enhanced features respectively; based on the target position information in the detection results, screen out the feature vectors belonging to the animal category within each region from all sets of feature vectors, and calculate the average values of the homogeneous target feature vectors in the infrared-specific features and the visible-light enhanced features respectively as the representative features of the corresponding regions; determine the horizontal center coordinates of the pseudo-anchors according to the region positions, and map the average values of the infrared and visible-light features to the corresponding coordinate positions respectively to generate a set of pseudo-anchors with unique region numbers and modality type labels.

[0046] Input the set of pseudo-anchors into the improved YOLOv7 detection head for multi-scale feature map convolution processing. The specific steps are as follows: The first detection layer generates prediction results for large-size targets through convolution, the second detection layer processes medium-size targets through downsampling, and the third detection layer further processes small-size targets through downsampling; each detection layer outputs a prediction vector containing the bounding box coordinates, confidence, and class probabilities, and maps the coordinates to the normalized range, converts the confidence to a probability value, and normalizes the class probabilities through a non-linear transformation.

[0047] The prediction results at different scales are upsampled to a unified resolution and stitched and fused to form a multi-scale comprehensive prediction result. Finally, candidate boxes with a probability higher than the class probability threshold and a qualified confidence level in the animal category are selected.

[0048] Further explanation, the class probability threshold is used to screen the probability that the target belongs to the animal category, indicating the minimum requirement for the confidence level of classifying the target as an animal. It is set according to the classification difficulty and data distribution of the animal category. Generally, the value range is 0.3 - 0.6 (for example, 0.5 is used to balance multi-class detection, and 0.3 is used for small targets or occluded scenarios).

[0049] The evaluation and analysis module converts the bounding box parameters of each candidate box in the high-confidence candidate box set into a state vector, and combines Kalman filtering and cross-frame feature association to generate an animal target list.

[0050] Specifically, the center coordinates of the candidate box are scaled proportionally to the horizontal block index according to the width of the original image, and the validity of the separate numbering is verified; based on the block numbering, infrared pseudo-anchor features and visible-light-enhanced pseudo-anchor features are extracted from the pseudo-anchor set, and the dominant pseudo-anchor is selected according to the modal primary and secondary markers of the candidate box. When the main data stream is infrared, the infrared pseudo-anchor features are selected, and when the auxiliary data stream is visible light, the visible-light-enhanced pseudo-anchor features are selected; the same-class aggregation constraint is performed on the infrared specific features of the candidate box, the cosine similarity between the infrared specific features and the infrared pseudo-anchor features of the same block is calculated and minimized, driving the same-class features to aggregate towards the pseudo-anchor; the different-class exclusion constraint is performed on the left and right adjacent blocks of the block where the candidate box is located, the infrared pseudo-anchors of the adjacent blocks are retrieved, the cosine similarity between the infrared specific features and the adjacent pseudo-anchors is calculated, and the negative contribution of the cosine similarity is maximized to suppress cross-block feature confusion.

[0051] The loss function is constructed by integrating the similarity calculation results of all candidate boxes. The gradient of the loss with respect to the weights of the improved YOLOv7 detection head is calculated through the chain rule, and the convolution kernel parameters are updated using the Adam optimizer to enhance the consistency of the same-class target features, and a high-confidence candidate box set is output.

[0052] Convert the bounding box parameters (center coordinates, width and height) of each candidate box in the high-confidence candidate box set into a state vector. If the candidate box is marked as infrared-dominated, associate it with the infrared pseudo-anchor feature; if it is visible light-dominated, associate it with the visible light-enhanced pseudo-anchor feature. Predict the current target position based on the state vector, and use the candidate boxes in the high-confidence candidate box set as observations to update the state vector through Mahalanobis distance matching, generating a prediction box with motion correction. Perform cross-frame association between the corrected prediction box and the infrared pseudo-anchor feature and the visible light-enhanced pseudo-anchor feature, and fuse the infrared pseudo-anchor feature and the visible light-enhanced pseudo-anchor feature to calculate the spatio-temporal-feature comprehensive similarity. For the targets that have not been matched continuously by the Kalman prediction, decide whether to terminate the tracking according to the similarity result. For the targets that have not been successfully matched for multiple consecutive frames, determine that they have left the monitoring area and remove them from the tracking list, generating an animal target list.

[0053] Further explain that the animal target list includes the animal target ID, the current frame bounding box coordinates, the infrared pseudo-anchor feature, the visible light-enhanced pseudo-anchor feature, the motion trajectory queue, and the target visibility flag (0: normal, 1: prediction state).

[0054] The hierarchical response module extracts the motion trajectory from the animal target list, combines the OpenPose key point detection, and fuses the historical behavior features to input into the LSTM-TCN hybrid model, outputs the attack probability and maps it to the dynamic threat level, triggering the hierarchical response mechanism.

[0055] Specifically, extract the center coordinate data of the latest consecutive multiple frames from the trajectory queue of the animal target list, arrange them in ascending order of timestamp, and calculate the time interval according to the video stream frame rate. Calculate the horizontal displacement and vertical displacement for adjacent two frames. Combine the horizontal displacement, vertical displacement and time interval of adjacent two frames to calculate the instantaneous velocity, forming a velocity sequence, and the expression is: ; where, is the instantaneous velocity of the target at time , is the horizontal displacement, is the vertical displacement, is the time interval; Calculate the instantaneous acceleration between adjacent two frames based on the instantaneous velocity, generating an acceleration sequence corresponding to the velocity sequence; according to the velocity sequence and the acceleration sequence, take the coordinates of the latest two frames, calculate the horizontal direction change amount and the vertical direction change amount, and obtain the motion direction angle.

[0056] Read the protected area boundary line equation from the preset parameter library, and the expression is: ; where, To control the inclination of a straight line in the horizontal direction, To control the inclination of a straight line in the vertical direction, is a constant term that determines the position offset of the straight line, and are the horizontal direction coordinate variable and the vertical direction variable in the rectangular coordinate system of the plane.

[0057] Based on the equation of the protected area boundary line, calculate the boundary direction angle, and the expression is: ; where is the boundary direction angle.

[0058] If the boundary is a curve, then take the tangent direction of the target's current position as the boundary direction angle.

[0059] Calculate the angular deviation between the movement direction angle and the boundary direction angle, and use the angular deviation as the threat assessment parameter for the target approaching the boundary.

[0060] Based on the calculated angular deviation between the movement direction angle and the boundary direction angle, combined with the coordinates (center point, width, and height) of the animal target detection box, perform pose analysis, specifically as follows: Intercept a rectangular area from the four-channel input matrix after fusing the visible light enhanced feature and the infrared specific feature, and output the cropped multi-modal image block; input the multi-modal image block into the pre-trained OpenPose key point detection model, generate multi-scale feature maps through the feature pyramid network, and output the heat map and affinity map of the skeletal key points; perform local maximum screening on the heat map, suppress non-maximum regions and retain high-confidence candidate points; for the candidate points, calculate the weighted response value of the surrounding area through a Gaussian weighted window, further strengthen the spatial consistency and eliminate isolated noise points to complete the heat map optimization; for the affinity map, uniformly sample multiple positions along the limb direction between the candidate points, and extract the candidate limb vectors and affinity vectors of each point; verify the rationality of the connection between the key points by calculating the average matching degree of the candidate limb vectors and the affinity vectors; perform weighted fusion on the heat map weighted response value and the affinity matching degree to obtain the preliminary comprehensive confidence, and after normalization processing, generate the confidence score.

[0061] Determine the confidence threshold according to experimental verification. In the training dataset, when the confidence ≥ 0.5, the key point detection accuracy exceeds 95%, and the false detection rate is lower than 5%. Therefore, the confidence threshold is used as the dividing line between high and low confidence; according to the confidence score, filter out the low-confidence key points; based on the high-confidence nose and hip key points after screening, calculate the vertical height difference of the head. If the confidence of any key point is insufficient, the vertical height difference is set to zero, and record the confidence mean value for the vertical height difference of the head that meets the confidence threshold.

[0062] For each limb side (left front, right front, left rear, right rear), verify whether the confidence levels of the corresponding key points such as shoulders, elbows, hips, and knees meet the confidence threshold; if the confidence levels of all associated key points meet the standard, calculate the straight-line distances from the shoulders to the elbows and from the hips to the knees respectively to measure the extension degree of the unilateral limb; if the confidence level of any key point is insufficient, directly determine that the opening degree of this limb side is invalid and assign a value of zero.

[0063] After completing the calculation of the unilateral limb, based on the average confidence level of each key point, generate a unilateral weight value reflecting the data reliability. The higher the weight, the stronger the credibility of the calculation result; integrate the opening degrees of the four limb sides and the corresponding weights, and fuse them into the overall limb opening degree measurement through weighted average. If the sum of the weights of all limb sides is zero (i.e., there is no reliable data), then determine that the overall opening degree is invalid and the result is reset to zero.

[0064] Based on the angular deviation between the calculated movement direction angle and the boundary direction angle, combined with the animal target ID in the animal target list and the movement trajectory queue, establish a grid coordinate system within the protected area. Each grid covers a fixed geographical range, retrieve the sequence of trajectory points where the animal target ID is located, map the central coordinates of each trajectory point to the corresponding grid number, and count the residence frequency of the target in each grid; based on the residence frequency distribution, calculate the information entropy, and the expression is: ; where is the information entropy value, representing the uncertainty of the target movement trajectory. The lower the entropy value, the stronger the regularity of the movement path, and the higher the entropy value, the greater the randomness of the behavior. is the grid number, that is, each grid represents a fixed area of the protected area. is the residence probability of the target in the th grid, reflecting the distribution density of the target in a specific area.

[0065] Based on the information entropy value, the access frequency of each grid within one minute is statistically counted. The grids with high access frequency are marked as high-risk attention areas and associated with the historical behavior characteristics of the animal target ID. The speed sequence, acceleration sequence, angle deviation, vertical height difference, limb opening degree, and historical behavior characteristics are aligned according to the time window and spliced into a multi-dimensional time series vector, which is input into a pre-trained LSTM-TCN hybrid model. The LSTM unit processes the input multi-dimensional time series vector in sequence, captures the long-term dependence relationship of the target behavior, outputs 128-dimensional time series features, and retains the time step information. The dilated causal convolution layer of the TCN applies the dilation factor, convolution kernel, and extended receptive field in sequence, extracts the local behavior pattern, and outputs time series features with the same dimension as the output of the LSTM. The output features of the LSTM and the TCN are spliced into a 256-dimensional vector, which is input into the fully connected layer. The attack probability is output through the Sigmoid function, representing the possibility of the target launching an attack, and the Softmax function is used to map the attack probability to the threat level.

[0066] Specifically, the dynamic division ratio points are set to 0.3 and 0.7, and the attack probability interval is evenly divided into three segments corresponding to the low threat level, medium threat level, and high threat level respectively. When the attack probability is less than 0.3, it represents the low threat level. When the attack probability is less than 0.7 and greater than or equal to 0.3, it represents the medium threat level. When the attack probability is greater than or equal to 0.7, it represents the high threat level.

[0067] Associate the surrounding real-time environmental data, dynamically correct the threat level, and trigger the fuzzy rules.

[0068] Among them, the real-time environmental data includes the personnel density and the status of the protection facilities.

[0069] If the attack probability is high and the personnel density is also high, the threat level is raised by one level. If the attack probability is high and the status of the protection facilities fails, the threat level is raised by one level.

[0070] The original threat level is raised according to the conditions of the fuzzy rule base. If multiple rules are satisfied simultaneously, the highest level takes effect.

[0071] Integrate the multi-source analysis results to generate a structured threat event list, including the animal target ID associated with the latest detection box parameters and the instantaneous speed. The threat level is updated in real time according to the fuzzy logic correction result. The confidence level is the attack probability, the list of high-risk grid numbers, and the location coordinates and status of the nearest protection facilities.

[0072] Based on the threat levels (low, medium, high) in the threat event list, perform hierarchical responses: Low threat level: Extract the target detection box coordinates, movement speed, and AR identification parameters from the threat event list. Input the video stream after fusing the visible light enhancement features and infrared specific features into the AR rendering engine, overlay a dynamic tracking box (red semi-transparent rectangle) and a risk identification (exclamation mark icon), and push the AR screen to the monitoring terminal in real time through the WebRTC protocol. Synchronously record the timestamp of the AR screen, the animal target ID, and the rendering parameters; Medium threat level: Retrieve the target location and threat level in the event list, trigger the frequency and flashing mode parameters of the audible and visual alarm device, send an alarm message through the SMS gateway, the content includes the target location, threat level, and timestamp, and write the device ID, command time, and response status into the local log; High threat level: Obtain the target coordinates, movement trajectory, and the status of the protection facilities. Based on the nearest target coordinate sequence in the movement trajectory queue, predict the trajectory in the short term in the future through cubic spline interpolation. At the same time, judge the available driving devices according to the status of the protection facilities. When the status of the protection facilities is normal, select the drone to cruise and generate a cruise path centered on the predicted trajectory points. When the status of the protection facilities fails, activate the ultrasonic transmitter to emit pulsed ultrasonic waves regularly, and the lighting interference system activates the stroboscopic white light LED array to flash at a fixed time interval.

[0073] In summary, the present invention dynamically adjusts the primary and secondary data streams based on the signal-to-noise ratio detection of the edge computing node, combines the VSYNC pulse synchronization and clock calibration of the FPGA timer to ensure the time stamp alignment error of cross-modal data, and avoids the problem of missed detection in low light caused by the fixed threshold switching; The visible light frequency domain phase spectrum retains high-frequency texture details, the infrared three-level wavelet decomposition extracts the diagonal high-frequency sub-band, generates sparse features through the Sobel gradient amplitude and neighborhood non-maximum suppression, and forms a complementary feature matrix after bilinear interpolation alignment to improve the target discrimination in complex backgrounds; The mixed feature tensor is horizontally divided into equal-width regions, the mean values of infrared and visible light features are calculated to generate a set of pseudo anchor points, and the improved YOLOv7 detection head drives the features of the same class of targets to gather towards the pseudo anchor points through the same-class aggregation and different-class repulsion loss functions, suppressing cross-block interference and improving the animal target detection accuracy.

[0074] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An intelligent security video surveillance system for preventing animal attack behavior, characterized in that: including a synchronous acquisition module that synchronously acquires a thermal radiation image sequence and an RGB image sequence to generate a cross-modal video stream data packet; a feature extraction module that performs multi-spectral feature extraction on the cross-modal video stream data packet to generate a visible light enhanced feature and an infrared specific feature; a feature fusion module that extracts cross-modal shared features from the visible light enhanced feature and the infrared specific feature through a shared weight convolutional layer, and splices them to form a mixed feature tensor; a detection and tracking module that horizontally divides the mixed feature tensor into equal-width regions, calculates regional pseudo anchor points, and generates multi-scale candidate boxes through an improved YOLOv7 detection head. The pseudo anchor points guide homogeneous aggregation and heterogeneous exclusion losses, and output a set of high-confidence candidate boxes; an evaluation and analysis module that converts the bounding box parameters of each candidate box in the set of high-confidence candidate boxes into state vectors, combines Kalman filtering and cross-frame feature association to generate an animal target list; a hierarchical response module that extracts the motion trajectory from the animal target list, combines OpenPose key point detection, and fuses historical behavior features to input into an LSTM-TCN hybrid model, outputs an attack probability and maps it to a dynamic threat level, and triggers a hierarchical response mechanism.

2. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, characterized in that: The synchronous acquisition of the thermal radiation image sequence and the RGB image sequence includes the following steps detect the ambient light intensity through the light intensity sensor of the visible light camera, start RGB image acquisition when it is higher than the average light intensity value, and start infrared thermal imaging acquisition when it is lower than the average light intensity value; the main control chip synchronously sends a VSYNC vertical synchronization pulse, aligns the timestamps of the thermal radiation image sequence and the RGB image sequence through an FPGA timer, and compresses them into an independent infrared video stream and a visible light video stream; dynamically switch the primary and secondary data streams according to the signal-to-noise ratio detection result of the visible light video stream, and generate a cross-modal video stream data packet.

3. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The multi-spectral feature extraction includes the following steps separate the three-channel matrix of the RGB image, generate a complex frequency domain matrix through two-dimensional fast Fourier transform and Hanning window function, extract the phase spectrum components of each channel and superimpose them with the three-channel grayscale matrix to generate a visible light enhanced feature; perform three-level discrete wavelet transform on the thermal radiation image, decompose it to the third-level diagonal high-frequency subband through Haar wavelet basis, and generate an infrared specific feature through Sobel gradient calculation and non-maximum suppression.

4. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The extraction of cross-modal shared features includes the following steps splice the three-channel data of the visible light enhanced feature and the single-channel data of the infrared specific feature into a four-channel input matrix; through a shared weight convolutional layer, sequentially perform large convolutional kernel feature downsampling, max pooling compression, and small convolutional kernel feature enhancement, and output cross-modal shared features.

5. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The splicing to form a mixed feature tensor includes the following steps uniformly divide the spatial dimension of the cross-modal shared feature into grid regions, perform global average pooling on each grid to generate local feature descriptors, and compress them into local feature matrices through a fully connected layer; adjust the number of channels of the infrared specific feature and the visible light enhanced feature to be consistent with the spatial resolution of the local feature matrix; splice the local feature matrix, the adjusted infrared specific feature, and the visible light enhanced feature along the channel dimension to form a mixed feature tensor.

6. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The generation of multi-scale candidate boxes by improving the YOLOv7 detection head includes the following steps: Divide the mixed feature tensor into non-overlapping equal-width regions, and extract the average values of the infrared and visible light feature vectors within the regions to generate a set of pseudo-anchors; Input the set of pseudo-anchors into the improved YOLOv7 detection head for multi-scale feature map convolution processing to generate multi-scale candidate boxes.

7. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The so-called pseudo-anchor-guided homogeneous aggregation and heterogeneous exclusion loss refers to performing homogeneous aggregation constraints on the infrared specific features of the candidate boxes, driving homogeneous features to gather towards the pseudo-anchors, and performing heterogeneous exclusion constraints on the left and right adjacent blocks of the block where the candidate box is located to suppress cross-block feature confusion.

8. The intelligent security video monitoring system for preventing animal attack behaviors according to claim 1, wherein: The generation of the animal target list includes the following steps: Convert the bounding box parameters of each candidate box in the high-confidence candidate box set into state vectors, and associate the infrared pseudo-anchor features and the visible light enhanced pseudo-anchor features; Predict the current target position based on the state vector, and update the state vector through Mahalanobis distance matching to generate a prediction box after motion correction; Perform cross-frame association on the corrected prediction box with the infrared pseudo-anchor features and the visible light enhanced pseudo-anchor features, calculate the spatio-temporal-feature comprehensive similarity, determine the continuously unmatched targets as leaving the monitoring area and remove them from the tracking list to generate the animal target list.

9. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, characterized in that: The output of the attack probability and mapping to the dynamic threat level includes the following steps: Based on the angular deviation between the motion direction angle in the animal target list and the boundary direction angle of the protected area, the head height difference and limb opening degree detected by OpenPose, and combine the information entropy to form a multi-dimensional time series vector; Input the multi-dimensional time series vector into the LSTM-TCN hybrid model, output the attack probability, and divide the low, medium, and high threat levels according to the proportional points.

10. The intelligent security video surveillance system for preventing animal attack behaviors according to claim 1, wherein: The so-called triggering of the hierarchical response mechanism means that in the low threat level, a dynamic tracking box is superimposed through the AR rendering engine and pushed to the monitoring terminal in real time. In the medium threat level, an audible and visual alarm is triggered and an alarm message is sent to the SMS gateway. In the high threat level, the drone cruise and ultrasonic interference are activated based on the trajectory prediction.

Citation Information

Patent Citations

  • Solar photovoltaic power generation prediction method based on TCN-LSTM

    CN110909926A

  • Analysis method for detecting group pig attack behaviors by adopting convolutional neural network and long-term and short-term memory

    CN111160422A

  • Human body abnormal behavior recognition alarm system and method under panoramic monitoring based on posture estimation

    CN112991656A

  • Crowd counting system and method based on cross-modal feature alignment fusion

    CN117315428A

  • Illumination perception-based dual-band improved YOLOv7 poultry disease detection method and system

    CN117809837A

Cited By

  • Infrared image power transmission equipment target identification method and system based on YOLOv7

    CN121280701A

  • Intelligent monitoring method and system for forest wild animals based on multi-source data fusion

    CN121660505A

  • Intelligent monitoring method and system for forest wild animals based on multi-source data fusion

    CN121660505B