Park monitoring method and system, electronic equipment and storage medium

By using multi-sensor collaborative monitoring technology and data stream information from infrared thermal imaging, visible light, audio, and vibration sensors, a multi-dimensional description vector is constructed, which solves the problems of false alarms and missed alarms in the park monitoring system under environmental changes and achieves highly accurate park security management.

CN120877487AActive Publication Date: 2025-10-31NANJING NANDA SIWEI TECHNOLOGY DEVELOPMENT CO LTD

Patent Information

Application Number
CN202510979826.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing park monitoring systems rely on visible light cameras, which are easily affected by changes in environmental conditions, leading to false alarms and missed alarms, thus reducing the accuracy of monitoring.

Method used

By employing multi-source data stream information from infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors, a multi-dimensional sensing data matrix is ​​formed through a unified timestamp. Cross-modal correlation features are extracted, a multi-dimensional description vector is constructed, and offset distance is calculated to generate early warning information.

Benefits of technology

It enables multi-dimensional collaborative monitoring of target behaviors within the park, improving monitoring accuracy, reducing false alarms and missed alarms, and ensuring the reliability of park security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877487A_ABST
    Figure CN120877487A_ABST
Patent Text Reader

Abstract

The invention discloses a park monitoring method and system, electronic equipment and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring data stream information of a plurality of sensors in a park; adding a uniform timestamp to the data stream information of each sensor, and synthesizing a multi-dimensional sensing data matrix according to the timestamps; based on the multi-dimensional perception data matrix, cross-modal association features are extracted; constructing a multi-dimensional description vector of the target monitoring behavior by using the cross-modal correlation feature, wherein the multi-dimensional description vector comprises a thermal radiation intensity value, a motion velocity component, an acoustic energy value and a vibration amplitude; and calculating an offset distance between the multi-dimensional description vector of the target monitoring behavior and each security description vector in a preset description vector set, and when the offset distance exceeds a preset threshold, generating early warning information. By implementing the technical scheme provided by the invention, the accuracy of park monitoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a park monitoring method, system, electronic device, and storage medium. Background Technology

[0002] With the acceleration of urbanization and the continuous advancement of smart park construction, various industrial parks, science and technology parks, and commercial parks are constantly expanding in scale, and the number of personnel, vehicles, and facilities within these parks is also increasing daily. This places higher demands on park security management. As a crucial link in ensuring the normal operation of the park, park security monitoring not only needs to promptly detect abnormal behavior but also ensure the accuracy and reliability of monitoring to effectively prevent various safety accidents.

[0003] Currently, the park's monitoring system mainly uses visible light cameras to capture images of the park and employs computer vision algorithms to analyze and identify target behaviors in the monitoring videos in real time, thereby detecting potential anomalies.

[0004] However, in practical applications, due to the complexity and variability of environmental conditions, it is difficult to accurately identify various abnormal behaviors by relying solely on visible light images. This can easily lead to false alarms and missed alarms when monitoring the park, thereby reducing the accuracy of park monitoring. Summary of the Invention

[0005] This application provides a campus monitoring method, system, electronic device, and storage medium that can improve the accuracy of campus monitoring.

[0006] Firstly, this application provides a method for monitoring a park, including: Data stream information from multiple sensors within the park is acquired, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors; A unified timestamp is added to the data stream information of each sensor, and the data from different sensors at the same time are combined into a multi-dimensional sensing data matrix according to the timestamp. Based on the multidimensional sensing data matrix, cross-modal correlation features are extracted, including the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration. A multidimensional description vector of the target monitoring behavior is constructed using the cross-modal correlation features. The multidimensional description vector includes thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude. The offset distance between the multidimensional description vector of the target monitoring behavior and each safety description vector in the preset description vector set is calculated. When the offset distance exceeds a preset threshold, an early warning message is generated.

[0007] By adopting the above technical solution, multi-source data stream information from infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors is acquired and a unified timestamp is added to form a multi-dimensional sensing data matrix, thereby achieving multi-dimensional collaborative monitoring of target behavior. Furthermore, by extracting cross-modal correlation features such as the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration, and constructing a multi-dimensional description vector containing thermal radiation intensity values, motion velocity components, acoustic energy values, and vibration amplitudes based on these features, the characteristics of target monitoring behavior can be more comprehensively and accurately characterized. Finally, by calculating the offset distance between the multi-dimensional description vector of target monitoring behavior and the preset safety description vector and performing threshold judgment, abnormal behavior can be effectively identified and early warning information can be generated in a timely manner, avoiding the false alarms and missed alarms that are prone to occur with single visible light monitoring, thereby improving the accuracy of park monitoring.

[0008] Optionally, the system clock information of each sensor is acquired and calibrated to the reference clock of the corresponding monitoring center of the park; the data stream information of each sensor is sampled at a preset sampling period according to the reference clock to generate a sampled data sequence with a unified timestamp; the data elements of each sensor at the same sampling time in the sampled data sequence are arranged and combined according to a preset data structure to obtain a multi-dimensional sensing data matrix.

[0009] Optionally, target detection is performed on the thermal imaging data and visible light image data in the multidimensional sensing data matrix to obtain the first position coordinates of the target in the thermal imaging image and the second position coordinates of the target in the visible light image; the overlap between the first position coordinates and the second position coordinates is calculated to generate a consistency verification value between thermal imaging and visible light; Fourier transform is performed on the audio data and vibration data in the multidimensional sensing data matrix to obtain the corresponding spectral distribution; the cross-correlation coefficient of the spectral distribution is calculated to obtain the frequency domain correlation coefficient between audio and vibration.

[0010] Optionally, a first target bounding box at the first location coordinates and a second target bounding box at the second location coordinates are established; the area of ​​the intersection region and the area of ​​the union region of the first target bounding box and the second target bounding box are calculated; the area of ​​the intersection region is divided by the area of ​​the union region to obtain the region overlap; the region overlap is used as a consistency verification value between thermal imaging and visible light.

[0011] Optionally, based on the consistency verification values ​​of the thermal imaging and visible light, thermal imaging data, visible light images, audio data, and vibration data of the target area are screened. The average temperature of the target is extracted from the screened thermal imaging data as the thermal radiation intensity value. The target is tracked on the screened visible light image sequence, and the displacement change rate of the target on the horizontal and vertical axes is calculated as the motion velocity component. The effective frequency band is determined according to the frequency domain correlation coefficient of the audio and vibration, and the total energy of the screened audio data within the effective frequency band is calculated as the acoustic energy value. The maximum amplitude of the screened vibration data within the effective frequency band is extracted as the vibration amplitude. The thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude are combined into a multi-dimensional description vector of the target monitoring behavior.

[0012] Optionally, based on the multidimensional description vectors, a sequence of multidimensional description vectors for the target monitoring behavior within a continuous preset time window is determined; based on the multidimensional description vector sequence, a steady-state characteristic benchmark value for the target monitoring behavior is calculated using a moving average method; the steady-state characteristic benchmark value is subtracted from each safety description vector in the preset description vector set to obtain a relative offset; variance analysis is performed on the relative offset to determine the weighting coefficients of each characteristic component; the relative offsets are weighted and summed based on the weighting coefficients to obtain the offset distance of the target monitoring behavior relative to each safety description vector.

[0013] Optionally, the number of times the target monitoring behavior deviates beyond a preset threshold within a consecutive preset time window is counted, and the warning level corresponding to the target monitoring behavior is determined based on the number of times; the movement trajectory and activity area information of the target monitoring behavior are extracted; based on the movement trajectory and activity area information, warning information containing a description of abnormal behavior and an area identifier is generated, and the warning information is sent to the staff corresponding to the warning level.

[0014] A second aspect of this application provides a campus monitoring system, the system comprising: The information acquisition module is used to acquire data stream information from multiple sensors within the park, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors. The data matrix determination module is used to add a unified timestamp to the data stream information of each sensor, and combine the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamp. The description vector determination module is used to extract cross-modal correlation features based on the multi-dimensional sensing data matrix. The cross-modal correlation features include the consistency verification values ​​of thermal imaging and visible light at the corresponding target positions, and the frequency domain correlation coefficients of audio and vibration. The module uses the cross-modal correlation features to construct a multi-dimensional description vector of the target monitoring behavior. The multi-dimensional description vector includes thermal radiation intensity values, motion velocity components, acoustic energy values, and vibration amplitudes. The behavior warning module is used to calculate the offset distance between the multi-dimensional description vector of the target's monitored behavior and each safety description vector in the preset description vector set. When the offset distance exceeds a preset threshold, a warning message is generated.

[0015] A third aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, the program being loaded and executed by the processor to implement a campus monitoring method.

[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a campus monitoring method.

[0017] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting the above technical solution, multi-source data stream information from infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors is acquired and a unified timestamp is added to form a multi-dimensional sensing data matrix, thereby achieving multi-dimensional collaborative monitoring of target behavior. Furthermore, by extracting cross-modal correlation features such as the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration, and constructing a multi-dimensional description vector containing thermal radiation intensity values, motion velocity components, acoustic energy values, and vibration amplitudes based on these features, the characteristics of target monitoring behavior can be more comprehensively and accurately characterized. Finally, by calculating the offset distance between the multi-dimensional description vector of target monitoring behavior and the preset safety description vector and performing threshold judgment, abnormal behavior can be effectively identified and early warning information can be generated in a timely manner, avoiding the false alarms and missed alarms that are prone to occur with single visible light monitoring, thereby improving the accuracy of park monitoring. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a park monitoring method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a park monitoring system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0019] Explanation of reference numerals in the attached drawings: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0021] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0022] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0023] This application provides a method for monitoring a park. In one embodiment, please refer to... Figure 1 , Figure 1 This is a flowchart illustrating the campus monitoring method provided in this application embodiment. This method can be implemented using a computer program, which can be integrated into an application or run as a standalone utility application. The method can also be implemented using a microcontroller and can run on a campus monitoring system based on the von Neumann architecture. Specifically, the method may include the following steps: Step 101: Obtain data stream information from multiple sensors within the park, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors.

[0024] Here, data stream information refers to a continuous sequence of data that is continuously collected and transmitted in real time by sensors. Specifically, data stream information includes thermal radiation intensity data streams collected by infrared thermal imaging sensors, image data streams collected by visible light camera sensors, acoustic signal data streams collected by audio sensors, and vibration signal data streams collected by vibration sensors.

[0025] Specifically, the first step is to acquire data stream information from multiple sensors within the park. Considering the complex and variable environmental conditions within the park, a single sensor cannot comprehensively and accurately perceive the target's behavioral characteristics. Therefore, this embodiment employs an infrared thermal imaging sensor, a visible light camera sensor, an audio sensor, and a vibration sensor to construct a multi-dimensional sensing network, enabling comprehensive monitoring of target behavior within the park. Specifically, the infrared thermal imaging sensor detects the target's thermal radiation characteristics and can operate effectively at night or in low-light conditions; the visible light camera sensor acquires visible light image information of the target, obtaining its appearance, movement, and other characteristics; the audio sensor collects acoustic signals from the environment and can identify abnormal sounds; and the vibration sensor detects vibration signals from the ground or equipment, sensing abnormal mechanical vibrations.

[0026] In actual deployment, these sensors are placed in appropriate locations based on the terrain features and key monitoring areas of the park. Infrared thermal imaging sensors and visible light camera sensors are deployed collaboratively to ensure overlapping fields of view for subsequent image matching; audio sensors are placed in areas prone to abnormal noise, such as equipment rooms and near walls; vibration sensors are installed on the ground or important equipment. Each sensor connects to the monitoring center via industrial Ethernet or a wireless network, transmitting data stream information in real time. Specifically, the infrared thermal imaging sensor outputs a thermal imaging video stream with a typical frame rate of 25fps; the visible light camera sensor outputs a high-definition video stream with a resolution of at least 1080P; the audio sensor collects acoustic signals with a sampling rate of 44.1kHz; and the vibration sensor samples at a frequency of 1kHz. By simultaneously acquiring data stream information from these different types of sensors, a multi-dimensional observation of target behavior is achieved. This multi-sensor collaborative sensing scheme overcomes the limitations of a single sensor under different environmental conditions, improving the environmental adaptability of the monitoring system. For example, even at night or in adverse weather conditions, target information can still be obtained through infrared thermal imaging despite the reduced quality of visible light images; when suspicious individuals approach quietly, even if they are not easily detected visually, they can be detected promptly through audio and vibration signals. The acquisition of multidimensional data streams lays the data foundation for subsequent cross-modal feature extraction and abnormal behavior identification, significantly improving the reliability and accuracy of park monitoring.

[0027] Step 102: Add a unified timestamp to the data stream information of each sensor, and combine the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamp.

[0028] The multidimensional sensing data matrix refers to a two-dimensional array that organizes data collected by different sensors at the same time according to a preset structure. Specifically, the row vectors of the multidimensional sensing data matrix represent sampling times, and the column vectors represent data features from different sensors. In this embodiment, the multidimensional sensing data matrix can be understood as an M×N matrix, where M represents the number of sampling times, N represents the dimension of sensor features, and each matrix element contains temperature distribution data from an infrared thermal imaging sensor, image data from a visible light camera sensor, acoustic data from an audio sensor, and vibration data from a vibration sensor. This matrix structure is used to achieve a unified representation of multi-source heterogeneous data. By aligning the data from different sensors in the time dimension, it facilitates subsequent cross-modal feature extraction and correlation analysis, thereby enabling multi-dimensional feature description and anomaly identification of target behavior within the park.

[0029] Specifically, since multiple sensors are distributed in different locations within the park, their system clocks may deviate, and their sampling frequencies and data formats may differ. To achieve effective fusion of multi-sensor data, it is necessary to synchronize and organize the data stream information of each sensor. In this embodiment, a unified timestamp is added to the data stream information collected by the infrared thermal imaging sensor, visible light camera sensor, audio sensor, and vibration sensor to ensure that different types of sensor data can be matched according to time correspondence. In specific implementation, a unified time benchmark is first established, and precise time information is marked for each data sampling point to achieve alignment of multi-source data. Then, the multi-dimensional data at the same moment is organized into a multi-dimensional sensing data matrix according to a preset data structure. Each row of this matrix corresponds to a sampling moment, and each column corresponds to data from a type of sensor. This time synchronization and data organization method not only solves the asynchronous problem of multi-sensor data but also provides a standardized data format for subsequent cross-modal feature extraction, effectively improving the efficiency and accuracy of multi-source data fusion analysis. For example, when abnormal behavior is detected, the corresponding data collected by each sensor at that moment can be quickly located, enabling multi-dimensional behavioral feature analysis and verification.

[0030] Based on the above embodiments, as an optional embodiment, step 102, which involves adding a unified timestamp to the data stream information of each sensor and combining the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamps, may further include the following steps: Step 201: Obtain the system clock information of each sensor and calibrate the system clock information to the reference clock of the corresponding monitoring center in the park.

[0031] Specifically, since the infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors distributed within the park operate at different clock speeds, it is first necessary to obtain the system clock information of each sensor. Specifically, local clock data from each sensor is collected via Network Time Protocol (NTP), compared with the atomic clock reference clock at the park's monitoring center, the clock deviation value is calculated, and a clock synchronization algorithm is used to calibrate the system clock of each sensor, ensuring that all sensor clocks are synchronized with the reference clock. This clock calibration mechanism ensures the time consistency of subsequent data sampling.

[0032] Step 202: Sample the data stream information of each sensor at a preset sampling period according to the reference clock to generate a sampled data sequence with a unified timestamp.

[0033] Specifically, after clock synchronization is completed, the data stream information of each sensor is uniformly sampled according to a preset sampling period, using the reference clock of the monitoring center as the standard. In practice, considering the significant differences in the original sampling rates of different sensors, a suitable uniform sampling period is selected for data resampling. For example, the sampling period can be set to 40ms, in which case infrared thermal imaging and visible light image data are sampled every frame, while audio and vibration data require downsampling. A timestamp corresponding to the reference clock is added to each sampled data, generating a sampled data sequence with a uniform timestamp. This uniform sampling strategy ensures the temporal correspondence of the data while balancing the system's computational overhead.

[0034] Step 203: Arrange and combine the data elements of each sensor at the same sampling time in the sampling data sequence according to the preset data structure to obtain a multi-dimensional sensing data matrix.

[0035] Specifically, after obtaining the sampling data sequence with a unified timestamp, it needs to be organized into a standardized data structure. Data elements at the same sampling time are arranged according to a preset format: infrared thermal imaging data can be represented as a temperature matrix, visible light image data as a pixel matrix, audio data as sound pressure levels, and vibration data as acceleration values. These data are combined according to their timestamp correspondence to form a multidimensional sensing data matrix, where each row corresponds to a multidimensional data set at a sampling time. This structured data organization not only achieves a unified representation of multi-source heterogeneous data but also provides a standardized data interface for subsequent feature extraction and analysis, improving data processing efficiency. This clock-synchronized data sampling and organization method solves the temporal consistency problem in multi-sensor data fusion, providing a reliable data foundation for abnormal behavior identification in park monitoring systems.

[0036] Step 103: Based on the multidimensional sensing data matrix, extract cross-modal correlation features, including the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration.

[0037] Cross-modal correlation features refer to the interrelated feature parameters extracted from heterogeneous data collected from different types of sensors. Specifically, cross-modal correlation features include the consistency verification value of thermal imaging and visible light at the corresponding target location, calculated through image registration, and the frequency domain correlation coefficient of audio and vibration signals, obtained through frequency domain analysis. In this embodiment, cross-modal correlation features can be understood as quantitative indicators reflecting the correlation between multi-dimensional sensing data. The consistency verification value characterizes the degree of matching between thermal imaging and visible light images for the same target detection results, and the frequency domain correlation coefficient characterizes the spectral correlation strength between acoustic and vibration signals. These features are used to achieve collaborative analysis of multi-source heterogeneous data, improving the accuracy and reliability of identifying target behavior characteristics within the park by mining the intrinsic correlation between different sensor data.

[0038] Specifically, to fully utilize the heterogeneous data information acquired by multiple sensors, it is necessary to extract cross-modal correlation features from the multi-dimensional sensing data matrix, focusing on analyzing the spatial correspondence between thermal imaging and visible light images, as well as the frequency domain correlation between audio and vibration signals. First, image registration is performed on the thermal imaging and visible light data. Through projection transformation, the two images are mapped to the same coordinate system, and the consistency verification value between the thermal imaging intensity and visible light brightness value at the corresponding target location is calculated. This value reflects the degree of mutual verification between the two imaging modalities in target detection results. Simultaneously, frequency domain transformation is performed on the audio and vibration data to extract the spectral features of the two signals, and their frequency domain correlation coefficient is calculated. This coefficient characterizes the coupling relationship between acoustic signals and mechanical vibrations. By extracting these cross-modal correlation features, the consistency of observations of the same target behavior by different sensors is verified, and potential correlation patterns between multi-dimensional data are discovered, effectively improving the reliability of abnormal behavior identification. For example, when suspicious intrusion behavior is detected, the consistency verification of thermal imaging and visible light can reduce the false alarm rate, while the correlation analysis of audio and vibration can be used to confirm the authenticity of the abnormal behavior.

[0039] Based on the above embodiments, as an optional embodiment, step 103: extracting cross-modal correlation features based on the multidimensional perceptual data matrix, this step may further include the following steps: Step 301: Perform target detection on the thermal imaging data and visible light image data in the multidimensional sensing data matrix to obtain the first position coordinates of the target in the thermal imaging image and the second position coordinates of the target in the visible light image.

[0040] Specifically, to extract the spatial correlation features between thermal imaging and visible light images, target detection is first required for both thermal imaging and visible light image data in the multidimensional sensing data matrix. In practice, the YOLOv5 target detection algorithm is used to process the thermal imaging and visible light images separately. This algorithm extracts image features through a convolutional neural network and performs target localization and classification on the feature map. For thermal imaging images, the algorithm primarily detects targets based on the boundary features of temperature anomaly regions, outputting the first position coordinates of the target, including the upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2) of the target bounding box. For visible light images, the algorithm detects targets based on their visual appearance features, outputting the second position coordinates of the target, which also includes the bounding box position information. This dual-modal target detection method fully utilizes the advantages of both thermal imaging and visible light, improving the accuracy of target localization.

[0041] Step 302: Calculate the overlap between the first position coordinates and the second position coordinates to generate a consistency verification value between thermal imaging and visible light.

[0042] Specifically, to evaluate the consistency of thermal imaging and visible light images in detecting the same target, it is necessary to calculate the overlap between the first and second position coordinates. In practice, firstly, the two sets of position coordinates are mapped to the same coordinate system through coordinate transformation for comparison. Then, the spatial overlap between these two sets of position coordinates is calculated, yielding an overlap value characterizing the spatial consistency of the detection results from the two images. This overlap value is the consistency verification value between thermal imaging and visible light, used to quantify the matching degree of target detection results under different imaging modalities. This consistency verification method based on position coordinate overlap can effectively evaluate the reliability of multimodal target detection, providing accurate feature support for subsequent abnormal behavior recognition. For example, a higher consistency verification value indicates a good spatial correspondence between the thermal imaging and visible light detection results, thereby improving the system's detection reliability.

[0043] Based on the above embodiments, as an optional embodiment, step 302: calculating the overlap between the first position coordinates and the second position coordinates to generate a consistency verification value between thermal imaging and visible light, may further include the following steps: Step 312: Create a first target bounding box with first location coordinates and a second target bounding box with second location coordinates.

[0044] Specifically, to quantitatively evaluate the spatial consistency of target detection results in thermal imaging and visible light images, it is first necessary to establish standardized target bounding boxes based on the detected position coordinates. In practice, for the first position coordinates (x1_t, y1_t, x2_t, y2_t) in the thermal imaging image, a rectangular first target bounding box is constructed, where (x1_t, y1_t) are the coordinates of the upper left corner, and (x2_t, y2_t) are the coordinates of the lower right corner. Similarly, for the second position coordinates (x1_v, y1_v, x2_v, y2_v) in the visible light image, a rectangular second target bounding box is constructed. Before establishing the bounding boxes, the coordinates in the thermal imaging coordinate system need to be transformed to the visible light coordinate system using a homography matrix to ensure that the two bounding boxes are compared in the same coordinate system. This standardized bounding box representation provides a unified geometric expression for subsequent overlap calculations.

[0045] Step 322: Calculate the area of ​​the intersection region and the area of ​​the union region of the first target bounding box and the second target bounding box.

[0046] Specifically, after obtaining two standardized target bounding boxes, it is necessary to calculate their spatial overlap. First, the area of ​​the intersection region of the two rectangular bounding boxes is calculated using the formula: the upper left corner coordinates are taken as the larger values ​​of the corresponding coordinates of the two boxes, max(x1_t, x1_v) and max(y1_t, y1_v), and the lower right corner coordinates are taken as the smaller values ​​of the corresponding coordinates of the two boxes, min(x2_t, x2_v) and min(y2_t, y2_v). Then, the area of ​​the rectangle determined by these coordinates is calculated. Simultaneously, the area of ​​the union region of the two bounding boxes is calculated using the formula: the sum of the areas of the two bounding boxes minus the area of ​​the intersection region. This area calculation method based on geometric operations can accurately reflect the degree of spatial overlap between the two detection boxes.

[0047] Step 332: Divide the area of ​​the intersection region by the area of ​​the union region to obtain the region overlap; use the region overlap as a verification value for the consistency between thermal imaging and visible light.

[0048] Specifically, to standardize the degree of overlap into a unified metric, the area of ​​the intersection region is divided by the area of ​​the union region to obtain the region overlap. This overlap is the consistency verification value between thermal imaging and visible light, calculated as: Consistency Verification Value = Intersection Area / Union Area. This value ranges from [0, 1], where 0 indicates that the two detection boxes do not overlap at all, and 1 indicates that the two detection boxes completely overlap. This consistency verification method based on the intersection-union ratio (IUU) exhibits good scale invariance and rotation invariance, objectively reflecting the spatial consistency of target detection results under different imaging modalities. For example, when the consistency verification value is greater than a preset threshold (e.g., 0.7), it can be considered that the detection results of thermal imaging and visible light for the same target have high spatial consistency, thereby improving the reliability of the detection results.

[0049] Step 303: Perform Fourier transform on the audio and vibration data in the multidimensional sensing data matrix to obtain the corresponding spectral distribution; calculate the cross-correlation coefficient of the spectral distribution to obtain the frequency domain correlation coefficient between audio and vibration.

[0050] Specifically, to analyze the frequency domain correlation between acoustic signals and mechanical vibrations, frequency domain analysis is performed on audio and vibration data in a multidimensional sensing data matrix. In practice, the audio and vibration signals are first preprocessed, including denoising and framing, with each frame containing 1024 points. Then, the Fast Fourier Transform (FFT) algorithm is used to calculate the spectral distribution of both signals, obtaining complex-form spectral sequences. The amplitude spectrum is obtained by modulo operation on the spectral sequences and normalized. Finally, the cross-correlation coefficient between the normalized audio and vibration spectra is calculated using the Pearson correlation coefficient of the two spectral sequences. This coefficient is the frequency domain correlation coefficient between audio and vibration, ranging from -1 to 1; a larger absolute value indicates a stronger correlation between the two signals in the frequency domain. This correlation calculation method based on frequency domain analysis can effectively identify the coupling relationship between acoustic and vibration events, providing important feature evidence for abnormal behavior detection.

[0051] Step 104: Construct a multidimensional description vector of the target monitoring behavior using cross-modal correlation features. The multidimensional description vector includes thermal radiation intensity value, motion velocity component, acoustic energy value and vibration amplitude.

[0052] A multidimensional descriptive vector (MDV) is a feature vector that organizes the multimodal features of target monitoring behavior. Specifically, an MMV includes thermal radiation intensity values ​​extracted from thermal imaging data, motion velocity components calculated from visible light image sequences, acoustic energy values ​​extracted from audio data, and vibration amplitude values ​​extracted from vibration data. An MMV can be understood as a vector representation integrating features from multimodal sensing data such as heat, light, sound, and vibration, where each component represents a different physical attribute of the target behavior. This vector is used to achieve a unified feature representation of target monitoring behavior, providing a complete feature description for subsequent behavior recognition and anomaly detection by fusing multimodal sensing information.

[0053] Specifically, to comprehensively characterize the monitoring behavior of a target, a multi-dimensional descriptive vector needs to be constructed based on cross-modal correlation features. In practice, firstly, the average thermal radiation intensity value of the target area is extracted from thermal imaging data to characterize the target's thermal features; secondly, the displacement change rate of the target in the horizontal and vertical directions is calculated from visible light image sequences to obtain the motion velocity component, which characterizes the target's motion features; thirdly, the energy value of the acoustic signal is extracted from audio data to characterize the acoustic features generated by the target; and fourthly, the amplitude of the vibration signal is extracted from vibration data to characterize the mechanical vibration features caused by the target. These features are organized into a multi-dimensional descriptive vector according to a preset order, realizing a multi-modal feature expression of the target's monitoring behavior. This feature expression method based on multi-dimensional descriptive vectors can integrate complementary information provided by different sensor data, thereby improving the completeness of the description and the accuracy of the identification of the target's behavioral features.

[0054] Based on the above embodiments, as an optional embodiment, step 104: constructing a multidimensional description vector of the target monitoring behavior using cross-modal correlation features, this step may further include the following steps: Step 401: Based on the consistency verification values ​​of thermal imaging and visible light, filter the thermal imaging data, visible light images, audio data and vibration data of the target area.

[0055] Specifically, to ensure the reliability of the multidimensional description vector, the multimodal data first needs to be filtered based on the consistency verification values ​​between thermal imaging and visible light. In practice, a threshold of 0.7 is set for the consistency verification value. When the consistency verification value of the detected target area in both thermal and visible light images exceeds this threshold, the thermal imaging data, visible light image, audio data, and vibration data within the corresponding time window are retained. This spatial consistency-based data filtering method can effectively remove false detection data caused by factors such as occlusion and changes in illumination.

[0056] Step 402: Extract the average temperature of the target from the screened thermal imaging data as the thermal radiation intensity value; perform target tracking on the screened visible light image sequence, and calculate the displacement change rate of the target on the horizontal and vertical axes as the motion velocity component.

[0057] Specifically, after obtaining the filtered multimodal data, it is necessary to extract the target's thermal and motion features. For thermal imaging data, the average temperature of all pixels within the target area is calculated to obtain the thermal radiation intensity value reflecting the target's thermal radiation characteristics. For visible light image sequences, the KCF tracking algorithm is used to obtain the target's positional changes between consecutive frames, and the displacement changes of the target in the horizontal and vertical directions per unit time are calculated to obtain the velocity components (vx, vy) characterizing the target's motion state. This characterization method, combining thermal and motion features, can simultaneously describe the target's temperature anomalies and behavioral trajectory.

[0058] Step 403: Determine the effective frequency band based on the frequency domain correlation coefficient between audio and vibration, and calculate the total energy of the filtered audio data within the effective frequency band as the acoustic energy value.

[0059] Specifically, to extract the acoustic features of a target, the effective frequency band of the signal needs to be determined based on the frequency domain correlation coefficient between audio and vibration. In practice, from the calculated frequency domain correlation coefficient sequence, the frequency range with an absolute correlation coefficient greater than 0.6 is selected as the effective frequency band. Then, the energy spectrum of the audio data within this frequency band is calculated, and the acoustic energy value is obtained by integrating the energy spectrum. This energy extraction method based on frequency domain correlation can effectively filter out interference signals in irrelevant frequency bands.

[0060] Step 404: Extract the maximum amplitude of the filtered vibration data within the effective frequency band as the vibration amplitude; combine the thermal radiation intensity value, motion velocity component, acoustic energy value and vibration amplitude into a multi-dimensional description vector of the target monitoring behavior.

[0061] Specifically, mechanical features are extracted from vibration data and a multidimensional descriptive vector is constructed. Within a defined effective frequency band, the maximum amplitude of the vibration signal is identified through spectral analysis, serving as the vibration amplitude characterizing the mechanical vibration intensity. Then, the thermal radiation intensity value, velocity components (vx, vy), acoustic energy value, and vibration amplitude are combined in a predetermined order to form a multidimensional descriptive vector describing the target's monitoring behavior. This vector representation method, which integrates multidimensional features, achieves a comprehensive characterization of the target's behavioral features, providing reliable feature support for subsequent behavioral analysis and anomaly identification. For example, when a target exhibits abnormal behavior, its thermal radiation, motion state, acoustic features, and vibration features often change significantly simultaneously; these coordinated changes can be effectively captured through the multidimensional descriptive vector.

[0062] Step 105: Calculate the offset distance between the multidimensional description vector of the target monitoring behavior and each safety description vector in the preset description vector set. When the offset distance exceeds the preset threshold, generate an early warning message.

[0063] The preset description vector set refers to a safety description vector library obtained by sampling and training the target's behavior under normal conditions. Each safety description vector contains thermal radiation intensity, motion velocity components, acoustic energy, and vibration amplitude, used to characterize the target's standard behavioral features under safe conditions.

[0064] Offset distance refers to the vector deviation between the currently monitored multidimensional description vector and each safety description vector in the preset description vector set. It is used to quantify the degree of difference between the target's current behavioral characteristics and normal behavioral patterns.

[0065] Early warning information refers to alarm data generated by the system when abnormal target behavior is detected. This includes the time, location, and type of the anomaly, as well as the corresponding multi-dimensional descriptive vector parameters. This information is used to promptly notify relevant personnel to take countermeasures and achieve real-time monitoring of the park's security status.

[0066] Specifically, to promptly detect abnormal target behavior, a security assessment of the constructed multidimensional description vectors is required. In practice, a set of description vectors containing multiple sets of safe behavioral features is first pre-constructed. Each safe description vector represents the multidimensional feature parameters of the target's behavior under normal conditions. Then, the offset distance between the multidimensional description vector of the current target's monitored behavior and each safe description vector in the set is calculated. This offset distance reflects the degree of deviation between the target's current behavioral characteristics and the normal behavioral pattern. When the calculated offset distance exceeds a preset threshold, it indicates that the target's behavior is abnormal, and the system immediately generates an early warning message. This anomaly detection method based on multidimensional feature offset, by quantifying the difference between the target's behavior and the normal pattern, enables rapid identification and early warning of abnormal behavior, thereby improving the real-time performance and reliability of park security monitoring.

[0067] Based on the above embodiments, as an optional embodiment, step 105, which calculates the offset distance between the multidimensional description vector of the target monitoring behavior and each security description vector in the preset description vector set, may further include the following steps: Step 501: Determine the sequence of multidimensional description vectors of the target monitoring behavior within a continuous preset time window based on the multidimensional description vectors.

[0068] Specifically, to assess the dynamic changes in target behavior, a sequence representation of multidimensional descriptive vectors needs to be constructed along the time dimension. In practice, a preset time window length of 10 minutes is set, and multidimensional descriptive vectors of target behavior are collected within consecutive time windows. These vectors are then organized into a sequence according to time order. Each time window corresponds to a multidimensional descriptive vector, containing the target's thermal radiation intensity, velocity component, acoustic energy value, and vibration amplitude within that time period. This time window-based serialization method preserves the temporal correlation of target behavior characteristics.

[0069] Step 502: Calculate the steady-state characteristic baseline value of the target monitoring behavior using the moving average method based on the multidimensional description vector sequence.

[0070] Specifically, after obtaining the multidimensional description vector sequence, it is necessary to extract the steady-state features of the target behavior. A moving average method is used to smooth each feature component in the sequence, with a window length of 5 sampling points, sliding forward by one sampling point each time. For each sliding window, the mean of all vectors within the window is calculated across all dimensions, yielding the feature mean vector for that window. The feature mean vectors of all windows are then averaged again to obtain the feature benchmark value representing the steady-state characteristics of the target's monitored behavior. This feature extraction method based on moving averages effectively suppresses the influence of short-term fluctuations, obtaining a more stable feature representation of the target behavior. For example, when the target is in a normal state, its behavioral characteristics will fluctuate around a certain steady-state value; by calculating the feature benchmark value, this steady-state behavioral pattern can be accurately characterized.

[0071] Step 503: Subtract the steady-state feature benchmark value from each security description vector in the preset description vector set to obtain the relative offset.

[0072] Specifically, to accurately assess the degree of anomalousness in target behavior, a pre-defined set of description vectors needs to be standardized. In practice, each safe description vector in the set is subtracted from the calculated steady-state feature baseline value to obtain the offset of each feature component relative to the steady-state baseline. This normalization method based on the steady-state baseline eliminates baseline differences between different target behavior patterns, allowing the offset to better reflect the relative degree of change in behavioral characteristics.

[0073] Step 504: Perform variance analysis on the relative offset to determine the weighting coefficients of each feature component; perform weighted summation on the relative offset based on the weighting coefficients to obtain the offset distance of the target monitoring behavior relative to each security description vector.

[0074] Specifically, after obtaining the standardized relative offset, the contribution of different feature components to behavior recognition needs to be considered. A variance analysis is performed on each component of the relative offset to calculate the variance of each feature component across all security description vectors. A larger variance indicates a greater contribution of the feature component to distinguishing different behavior patterns. Based on the variance, each feature component is normalized to obtain its corresponding weighting coefficient. Then, each component of the relative offset is multiplied by its corresponding weighting coefficient, and the summation is applied to all weighted components to obtain the comprehensive offset distance of the target monitoring behavior relative to each security description vector. This weighted distance calculation method based on variance analysis can highlight the influence of key feature components and improve the accuracy of abnormal behavior recognition. For example, when changes in certain feature components have a stronger indicative effect on identifying target behavior, larger weighting coefficients can enhance the role of these features in offset distance calculation.

[0075] Based on the above embodiments, as an optional embodiment, step 105: generating warning information may further include the following steps: Step 505: Count the number of times the target monitoring behavior deviates beyond a preset threshold within a consecutive preset time window, and determine the warning level corresponding to the target monitoring behavior based on the number of times.

[0076] Specifically, to achieve tiered early warning for abnormal behavior, it is first necessary to assess the persistence of the target behavior anomaly. In practice, the number of times the offset distance exceeds a preset threshold is counted within a consecutive preset time window (e.g., 30 minutes). Based on the number of times the threshold is exceeded, the early warning level is divided into three levels: Level 3 warning (minor anomaly) when the number of exceedances is less than 5; Level 2 warning (moderate anomaly) when the number of exceedances is between 5 and 10; and Level 1 warning (serious anomaly) when the number of exceedances is greater than 10. This tiered early warning mechanism based on the persistence of anomalies can effectively filter out sporadic false alarms and improve the reliability of early warnings.

[0077] Step 506: Extract the motion trajectory and activity area information of the target monitoring behavior.

[0078] Specifically, to fully understand the abnormal behavior characteristics of a target, it is necessary to extract its activity information. Specifically, based on the target's position coordinates in a visible light image sequence, a trajectory fitting algorithm is used to reconstruct the target's trajectory; simultaneously, areas where the target appears frequently are statistically analyzed to determine its main activity areas. By analyzing the characteristic parameters of the trajectory (such as velocity, acceleration, and direction changes) and the attributes of the activity area (such as functional area type and security level), the specific type of abnormal behavior can be further determined. This behavior analysis method, combining trajectory and area information, can provide a more detailed description of abnormal behavior for early warning.

[0079] Step 507: Based on the movement trajectory and activity area information, generate an early warning message containing a description of the abnormal behavior and an area identifier, and send the early warning message to the staff corresponding to the warning level.

[0080] Specifically, based on the acquired abnormal behavior information, targeted early warning notifications need to be generated. In practice, the description of the abnormal behavior (including anomaly type, duration, and severity) and area identifiers (including area number and location description) are integrated into the early warning information. Then, the information is sent to personnel at the corresponding level according to the warning level: Level 1 warnings are sent to the security supervisor, on-site patrol personnel, and the monitoring center; Level 2 warnings are sent to on-site patrol personnel and the monitoring center; and Level 3 warnings are sent only to the monitoring center. This hierarchical response mechanism based on warning levels enables rapid handling of abnormal events and rational allocation of resources. For example, when the system detects suspicious gatherings of individuals in a certain area, it can promptly notify relevant personnel to conduct on-site verification, thereby reducing security risks.

[0081] Reference Figure 2 This application provides a park monitoring system, which includes: an information acquisition module, a data matrix determination module, a description vector determination module, and a behavior early warning module, wherein: The information acquisition module is used to acquire data stream information from multiple sensors within the park, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors. The data matrix determination module is used to add a unified timestamp to the data stream information of each sensor, and combine the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamp. The description vector determination module is used to extract cross-modal correlation features based on the multi-dimensional sensing data matrix. The cross-modal correlation features include the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration. The multi-dimensional description vector of the target monitoring behavior is constructed using the cross-modal correlation features. The multi-dimensional description vector includes thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude. The behavior warning module is used to calculate the offset distance between the multi-dimensional description vector of the target's monitored behavior and each safety description vector in the preset description vector set. When the offset distance exceeds the preset threshold, a warning message is generated.

[0082] Based on the above embodiments, the data matrix determination module is also used to acquire the system clock information of each sensor and calibrate the system clock information to the reference clock of the corresponding monitoring center in the park; to perform time sampling on the data stream information of each sensor according to the reference clock at a preset sampling period to generate a sampling data sequence with a unified timestamp; and to arrange and combine the data elements of each sensor at the same sampling time in the sampling data sequence according to a preset data structure to obtain a multi-dimensional sensing data matrix.

[0083] Based on the above embodiments, the description vector determination module is further used to perform target detection on the thermal imaging data and visible light image data in the multidimensional sensing data matrix, to obtain the first position coordinates of the target in the thermal imaging image and the second position coordinates of the target in the visible light image; calculate the overlap of the first position coordinates and the second position coordinates to generate a consistency verification value between thermal imaging and visible light; perform Fourier transform on the audio data and vibration data in the multidimensional sensing data matrix to obtain the corresponding spectral distribution; calculate the cross-correlation coefficient of the spectral distribution to obtain the frequency domain correlation coefficient between audio and vibration.

[0084] Based on the above embodiments, the description vector determination module is further used to establish a first target bounding box with a first position coordinate and a second target bounding box with a second position coordinate; calculate the area of ​​the intersection region and the area of ​​the union region of the first target bounding box and the second target bounding box; divide the area of ​​the intersection region by the area of ​​the union region to obtain the region overlap; and use the region overlap as a consistency verification value between thermal imaging and visible light.

[0085] Based on the above embodiments, the description vector determination module is also used to filter thermal imaging data, visible light images, audio data, and vibration data of the target area based on the consistency verification values ​​of thermal imaging and visible light; extract the average temperature of the target from the filtered thermal imaging data as the thermal radiation intensity value; perform target tracking on the filtered visible light image sequence, calculate the displacement change rate of the target on the horizontal and vertical axes as the motion velocity component; determine the effective frequency band according to the frequency domain correlation coefficient of audio and vibration, calculate the total energy of the filtered audio data within the effective frequency band as the acoustic energy value; extract the maximum amplitude of the filtered vibration data within the effective frequency band as the vibration amplitude; and combine the thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude into a multi-dimensional description vector of the target monitoring behavior.

[0086] Based on the above embodiments, the behavior warning module is further configured to: determine the multidimensional description vector sequence of the target monitoring behavior within a continuous preset time window based on the multidimensional description vector; calculate the steady-state characteristic benchmark value of the target monitoring behavior using the moving average method based on the multidimensional description vector sequence; subtract the steady-state characteristic benchmark value from each safe description vector in the preset description vector set to obtain the relative offset; perform variance analysis on the relative offset to determine the weight coefficient of each characteristic component; and perform weighted summation of the relative offset based on the weight coefficients to obtain the offset distance of the target monitoring behavior relative to each safe description vector.

[0087] Based on the above embodiments, the behavior warning module is also used to count the number of times the target monitoring behavior deviates from a preset threshold within a consecutive preset time window, and determine the warning level corresponding to the target monitoring behavior based on the number of times; extract the movement trajectory and activity area information of the target monitoring behavior; generate warning information containing abnormal behavior description and area identification based on the movement trajectory and activity area information, and send the warning information to the staff corresponding to the warning level.

[0088] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0089] This application also discloses an electronic device. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0090] The communication bus 302 is used to enable communication between these components.

[0091] The user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0092] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0093] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface graphics, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0094] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a campus monitoring method.

[0095] exist Figure 3In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and acquire user input data; while the processor 301 can be used to call an application program storing a campus monitoring method in the memory 305. When executed by one or more processors 301, the electronic device 300 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0096] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0097] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0098] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0101] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practical disclosure.

[0102] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only.

Claims

1. A method for monitoring a park, characterized in that, include: Data stream information from multiple sensors within the park is acquired, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors; A unified timestamp is added to the data stream information of each sensor, and the data from different sensors at the same time are combined into a multi-dimensional sensing data matrix according to the timestamp. Based on the multidimensional sensing data matrix, cross-modal correlation features are extracted, including the consistency verification values ​​of thermal imaging and visible light at the corresponding target locations, and the frequency domain correlation coefficients of audio and vibration. A multidimensional description vector of the target monitoring behavior is constructed using the cross-modal correlation features. The multidimensional description vector includes thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude. The offset distance between the multidimensional description vector of the target monitoring behavior and each safety description vector in the preset description vector set is calculated. When the offset distance exceeds a preset threshold, an early warning message is generated.

2. The park monitoring method according to claim 1, characterized in that, The step of adding a unified timestamp to the data stream information of each sensor and combining the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamp includes: Acquire the system clock information of each sensor and calibrate the system clock information to the reference clock of the corresponding monitoring center in the park; The data stream information of each sensor is sampled at a preset sampling period according to the reference clock to generate a sampled data sequence with a unified timestamp. The data elements of each sensor in the sampling data sequence at the same sampling time are arranged and combined according to a preset data structure to obtain a multidimensional sensing data matrix.

3. The park monitoring method according to claim 1, characterized in that, The extraction of cross-modal association features based on the multidimensional sensing data matrix includes: Target detection is performed on the thermal imaging data and visible light image data in the multidimensional sensing data matrix to obtain the first position coordinates of the target in the thermal imaging image and the second position coordinates of the target in the visible light image; Calculate the overlap between the first position coordinates and the second position coordinates to generate a consistency verification value between thermal imaging and visible light. Fourier transform is performed on the audio data and vibration data in the multidimensional sensing data matrix to obtain the corresponding spectral distribution; The cross-correlation coefficient of the spectral distribution is calculated to obtain the frequency domain correlation coefficient between audio and vibration.

4. The park monitoring method according to claim 3, characterized in that, The calculation of the overlap between the first and second position coordinates to generate a consistency verification value between thermal imaging and visible light includes: Establish a first target bounding box at the first location coordinates and a second target bounding box at the second location coordinates; Calculate the area of ​​the intersection region and the area of ​​the union region of the first target bounding box and the second target bounding box; Divide the area of ​​the intersection region by the area of ​​the union region to obtain the region overlap. The overlap of the regions is used as a verification value for the consistency between thermal imaging and visible light.

5. The park monitoring method according to claim 3, characterized in that, The construction of a multidimensional description vector of target monitoring behavior using the cross-modal correlation features includes: Based on the consistency verification values ​​between thermal imaging and visible light, thermal imaging data, visible light images, audio data, and vibration data of the target area were selected. The average temperature of the target is extracted from the filtered thermal imaging data and used as the thermal radiation intensity value. Target tracking is performed on the filtered visible light image sequence, and the displacement change rate of the target on the horizontal and vertical axes is calculated as the motion velocity component; The effective frequency band is determined based on the frequency domain correlation coefficient between the audio and vibration, and the total energy of the filtered audio data within the effective frequency band is calculated as the acoustic energy value. The maximum amplitude of the filtered vibration data within the effective frequency band is extracted as the vibration amplitude value. The thermal radiation intensity value, motion velocity component, acoustic energy value, and vibration amplitude are combined into a multidimensional descriptive vector of the target monitoring behavior.

6. The park monitoring method according to claim 1, characterized in that, The calculation of the offset distance between the multidimensional description vector of the target monitoring behavior and each security description vector in the preset description vector set includes: Based on the multidimensional description vector, determine the sequence of multidimensional description vectors of the target monitoring behavior within a continuous preset time window; Based on the multidimensional description vector sequence, the steady-state characteristic baseline value of the target monitoring behavior is calculated using the moving average method; Subtract the steady-state feature benchmark value from each security description vector in the preset description vector set to obtain the relative offset; A variance analysis was performed on the relative offset to determine the weighting coefficients of each feature component; The relative offsets are weighted and summed according to the weighting coefficients to obtain the offset distance of the target monitoring behavior relative to each security description vector.

7. The park monitoring method according to claim 1, characterized in that, The generation of early warning information includes: The number of times the target monitoring behavior deviates beyond a preset threshold within a consecutive preset time window is counted, and the warning level corresponding to the target monitoring behavior is determined based on the number of times. Extract the motion trajectory and activity area information of the target monitoring behavior; Based on the movement trajectory and activity area information, an early warning message containing a description of abnormal behavior and an area identifier is generated, and the early warning message is sent to the staff corresponding to the early warning level.

8. A park monitoring system, characterized in that, The system includes: The information acquisition module is used to acquire data stream information from multiple sensors within the park, including infrared thermal imaging sensors, visible light camera sensors, audio sensors, and vibration sensors. The data matrix determination module is used to add a unified timestamp to the data stream information of each sensor, and combine the data from different sensors at the same time into a multi-dimensional sensing data matrix according to the timestamp. The description vector determination module is used to extract cross-modal correlation features based on the multi-dimensional sensing data matrix. The cross-modal correlation features include the consistency verification values ​​of thermal imaging and visible light at the corresponding target positions, and the frequency domain correlation coefficients of audio and vibration. The module uses the cross-modal correlation features to construct a multi-dimensional description vector of the target monitoring behavior. The multi-dimensional description vector includes thermal radiation intensity values, motion velocity components, acoustic energy values, and vibration amplitudes. The behavior warning module is used to calculate the offset distance between the multi-dimensional description vector of the target's monitored behavior and each safety description vector in the preset description vector set. When the offset distance exceeds a preset threshold, a warning message is generated.

9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the campus monitoring method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the park monitoring method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Sea and lake surface target detection method and device, terminal and storage medium

    CN115880292A

  • Thermal power plant security device with artificial intelligence identity authentication

    CN118887615A

  • Intelligent security system for industrial park

    CN119399906A

  • Electronic sentry monitoring system

    CN119420873A

  • Video object tracking device and video object tracking program

    JP2007233798A

Cited By

  • Multi-source data fusion type passenger station intelligent management and control method and system based on AI drive

    CN121544079A