Data fusion-based secret stealing device detection method and system

By integrating multi-device data and AI intelligent analysis technology, the problem of low detection accuracy and poor efficiency of traditional security detection technology in complex environments has been solved, enabling precise positioning and efficient detection of espionage devices.

CN121959409APending Publication Date: 2026-05-01GUANGZHOU MOUNTAIN FRONT MEASUREMENT & CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511931520.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional security detection technologies suffer from low detection accuracy, poor efficiency, and weak adaptability when facing complex environments and diverse espionage devices, making them unable to effectively address the challenges posed by eavesdropping, surreptitious photography, and other espionage devices.

Method used

By using multi-device data fusion and AI intelligent analysis technology, the system can accurately locate hidden eavesdropping devices, cameras, and suspicious devices inside walls in classified locations. It collects infrared thermal imaging, optical images, nonlinear responses, radar echoes, and electromagnetic spectrum data, performs time synchronization, spatial synchronization, and feature fusion, and uses a cross-modal attention network for multimodal fusion model detection.

Benefits of technology

It significantly improves the detection accuracy, environmental adaptability, and work efficiency of espionage devices, and enables precise detection and location of hidden espionage devices in classified locations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121959409A_ABST
    Figure CN121959409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-source detection and auxiliary positioning, in particular to a data fusion-based secret stealing equipment detection method and system. According to the method, multiple kinds of data including infrared thermal imaging original data, optical image original data, nonlinear response original data, radar echo original data and electromagnetic spectrum original data are collected, and the data are fused into a fusion feature vector through time synchronization, space synchronization and feature fusion methods; the fusion feature vector can effectively represent the data features of all the data types, so that effective integration of multi-source heterogeneous data is realized, and the multi-source data can be synthesized to detect secret stealing equipment. And finally, the fusion feature vector is analyzed through the multi-modal fusion model, and multi-source data are integrated to detect the secret-stealing equipment, so that accurate detection and positioning of the hidden secret-stealing equipment in the secret-related place are realized, and the detection precision, the environmental adaptability and the working efficiency of the secret-stealing equipment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for detecting espionage devices based on data fusion Technical Field

[0001] This invention relates to the field of multi-source detection and assisted positioning technology, and in particular to a method and system for detecting espionage devices based on data fusion. Background Technology

[0002] In the field of information security, confidentiality testing is of paramount importance. With the rapid development of science and technology, eavesdropping and covert surveillance devices are becoming increasingly sophisticated and diverse, posing a serious threat to the information security of classified locations. Traditional confidentiality testing technologies are gradually revealing numerous shortcomings when dealing with complex environments and new types of eavesdropping devices, such as low detection accuracy, poor efficiency, and weak adaptability, making it difficult to meet the high standards of information security required by modern classified locations.

[0003] Current security detection technologies primarily rely on single or a small number of devices operating independently. Examples include using infrared thermal imagers to detect object heat, optical anti-eavesdropping detection devices to capture the reflective features of hidden cameras, nonlinear node detectors to detect the nonlinear responses of electronic components, and spectrum analyzers to analyze electromagnetic signals. These devices mostly operate independently, lacking a collaborative system, thus failing to fully utilize the complementary information between multiple devices. Specifically, an infrared thermal imager may detect areas of abnormal heat but cannot determine if they are eavesdropping devices; optical anti-eavesdropping detection devices may detect reflective features but cannot determine if they are related to electronic components; furthermore, manual analysis of electromagnetic spectrum and radar echoes is time-consuming and laborious, further reducing the reliability and efficiency of detection. Therefore, existing technologies are insufficient to effectively address the challenges posed by diverse eavesdropping devices, necessitating a highly efficient and accurate security detection solution that enables intelligent analysis. Summary of the Invention

[0004] This invention aims to provide a data fusion-based method and system for detecting espionage devices, in order to solve the problems of the limitations of traditional security detection equipment in single-device detection and insufficient comprehensive analysis capabilities, thereby achieving accurate detection and location of hidden espionage devices in classified locations and improving the detection accuracy, environmental adaptability and work efficiency of espionage devices.

[0005] To achieve the above objectives, this invention provides a data fusion-based method for detecting espionage devices, comprising the following steps: determining a target area and acquiring raw infrared thermal imaging data, raw optical image data, raw nonlinear response data, raw radar echo data, and raw electromagnetic spectrum data of the target area; performing time and spatial synchronization on the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data; performing feature fusion on the synchronized infrared thermal imaging data, the synchronized optical image data, the synchronized nonlinear response data, the synchronized radar echo data, and the synchronized electromagnetic spectrum data to obtain a fused feature vector; and constructing a multimodal fusion model based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fused feature vector, thereby completing the detection of the espionage device.

[0006] The aforementioned method for detecting espionage devices collects various types of data, including raw infrared thermal imaging data, raw optical image data, raw nonlinear response data, raw radar echo data, and raw electromagnetic spectrum data, in order to overcome the limitations of detecting data theft from a single device.

[0007] For existing technologies, the differences between the aforementioned multi-dimensional data are significant, making them difficult to combine and limiting detection to single data types. Therefore, the aforementioned method for detecting espionage devices utilizes time synchronization, spatial synchronization, and feature fusion to fuse raw infrared thermal imaging data, raw optical image data, raw nonlinear response data, raw radar echo data, and raw electromagnetic spectrum data into a fused feature vector. This fused feature vector effectively characterizes the data features of all the aforementioned data types, thereby achieving effective integration of multi-source heterogeneous data and enabling comprehensive detection of espionage devices.

[0008] Finally, the above-mentioned method for detecting espionage devices analyzes and fuses feature vectors through a multimodal fusion model, and integrates multi-source data to detect espionage devices, thereby achieving accurate detection and location of hidden espionage devices in classified locations, significantly improving the detection accuracy, environmental adaptability and work efficiency of espionage devices. Attached Figure Description

[0009] Figure 1 is a schematic flowchart of a data fusion-based method for detecting espionage devices provided in an embodiment of the present invention; Figure 2 is a schematic flowchart of a method for detecting espionage devices provided in an embodiment of the present invention; Figure 3 is a schematic flowchart of another method for detecting espionage devices provided in an embodiment of the present invention; Figure 4 is a schematic flowchart of yet another method for detecting espionage devices provided in an embodiment of the present invention. Detailed Implementation

[0010] In the field of information security, security detection is of paramount importance. With technological advancements, eavesdropping and covert surveillance devices are becoming increasingly sophisticated and sophisticated, posing a serious threat to the security of classified locations. Traditional security detection technologies are increasingly proving inadequate in the face of complex environments and new types of espionage equipment, failing to meet the high information security requirements of modern classified locations. Therefore, a more efficient and accurate security detection solution is urgently needed.

[0011] Current security detection technologies mainly rely on single or a small number of devices operating independently. For example, infrared thermal imagers are used to detect the heat generated by objects; optical anti-eavesdropping detection devices capture the reflective characteristics of hidden cameras; nonlinear node detectors are used to detect the nonlinear responses of electronic components; and spectrum analyzers are used to analyze electromagnetic signals. These devices mostly operate independently and have not formed a collaborative system.

[0012] Specifically, while infrared thermal imagers can detect the heat generated by objects, the temperature characteristics of normal heat-generating devices overlap with those of eavesdropping and surveillance devices. Furthermore, low-power eavesdropping devices do not generate significant heat, making them easily confused with surrounding objects that are normally generating heat, leading to false positives and false negatives. For example, some low-power eavesdropping devices generate heat similar to that of ordinary indoor electronic devices such as routers, making accurate differentiation based solely on temperature characteristics difficult.

[0013] Optical anti-eavesdropping detection equipment is susceptible to interference and misjudgment under complex lighting conditions, such as direct sunlight or indoor light reflection. For example, in a brightly lit room, reflections from window glass and metal decorations may be misinterpreted as camera reflections, affecting detection accuracy.

[0014] Nonlinear node detectors have limited detection capabilities for electronic components that are specially disguised or shielded. Some eavesdropping devices use special materials to wrap electronic components, reducing the outward radiation of their nonlinear characteristics and making them difficult for detectors to detect. Furthermore, nonlinear node detectors cannot identify the frequency band characteristics of wireless signals, and spectrum analysis equipment is relatively independent and difficult to combine with spatial positioning information.

[0015] The lack of coordinated analysis of the aforementioned thermal imaging, optical, and electromagnetic data resulted in a high rate of missed detections in complex environments such as those with strong electromagnetic interference and complex wall structures.

[0016] It is evident that traditional detection equipment relies on manual recording and analysis of results, a cumbersome and error-prone process. Manually recording the locations of anomalies in infrared thermal imaging images and the reflectivity in optical anti-espionage images is not only time-consuming and labor-intensive but also prone to errors due to human negligence, failing to meet the demands for rapid and efficient security detection. Existing security detection technologies suffer from low accuracy, poor efficiency, and weak adaptability when facing complex environments and diverse eavesdropping devices, failing to meet the high information security requirements of modern classified locations. There is an urgent need for a security detection solution that integrates the advantages of multiple devices and enables intelligent analysis and rapid reporting.

[0017] This invention aims to address the limitations of traditional security detection equipment, such as the limitations of single-device detection, insufficient comprehensive analysis capabilities, and unintelligent report generation. By integrating multiple devices and AI intelligent analysis technology, it achieves precise location of hidden eavesdropping devices, cameras, and suspicious devices embedded in walls in classified locations. Simultaneously, it enables real-time monitoring of spatial electromagnetic signals and source location, providing a comprehensive and efficient security detection solution for classified locations, improving detection accuracy, environmental adaptability, and work efficiency.

[0018] Please refer to Figure 1. The first embodiment of the present invention provides a data fusion-based method for detecting espionage devices, including the following steps: S101, determining a target area and acquiring raw infrared thermal imaging data, raw optical image data, raw nonlinear response data, raw radar echo data, and raw electromagnetic spectrum data of the target area; S102, performing time and spatial synchronization on the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data; S103, performing feature fusion on the synchronized infrared thermal imaging data, the synchronized optical image data, the synchronized nonlinear response data, the synchronized radar echo data, and the synchronized electromagnetic spectrum data to obtain a fused feature vector; S104, constructing a multimodal fusion model based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fused feature vector, thereby completing the detection of the espionage device.

[0019] Referring to Figure 2, in one specific embodiment, the target area is a conference room. Step S101 includes: raw infrared thermal imaging data acquisition: The operator holds the integrated device, starts the infrared thermal imager, sets the temperature resolution to 0.05℃, and the frame rate to 30fps. The conference room is scanned in all directions according to a preset scanning path, covering areas such as the ceiling, walls, tables, and chairs. For example, when scanning the filing cabinet in the northwest corner of the conference room, the thermal imager captures an abnormal area with a temperature of approximately 32℃, which is 2℃ higher than the surrounding environment.

[0020] Raw optical image data acquisition: After detecting an abnormal area using infrared thermal imaging, the system automatically switches to optical anti-eavesdropping detection equipment, adjusting the focus to 3 meters. Optical image acquisition is performed on the conference room filing cabinet area, capturing a faint circular reflective point in the gap of a certain filing box.

[0021] Raw data acquisition for nonlinear response: For this reflective area, a nonlinear node detector was activated, set to high-frequency detection mode, and operating at a frequency of 1.5 GHz. The detector detected a nonlinear response signal in this area, with a signal strength of 45 dBm.

[0022] Raw radar echo data acquisition: Through-wall radar scanning was performed on the east wall of the conference room. The wall is made of concrete and is 25cm thick. The radar transmission frequency is 1.8GHz. An abnormal echo signal with a reflection intensity of 30dB was detected 15cm deep inside the wall.

[0023] Raw electromagnetic spectrum data: In the center of the conference room, a spectrum analyzer performed a real-time scan of the 20MHz-3GHz frequency band with a scan step size of 1MHz. A persistent and stable signal was found in the 2.45GHz band, with a power spectral density of 20dBm / MHz and a modulation method of Gaussian frequency shift keying (GFSK).

[0024] To further explain the implementation details of step S101, the following describes the data acquisition equipment used in step S101.

[0025] The working principle of an infrared thermal imager: Based on Planck's law, the thermal radiation energy of an object is proportional to the fourth power of its temperature. Where W is the radiative exitance. Where is Boltzmann's constant, and T is the absolute temperature. Infrared radiation emitted by an object is detected, focused onto an infrared detector by an optical system, converted into an electrical signal, and then processed to generate a thermal image.

[0026] Key parameters and application examples of infrared thermal imagers: temperature resolution ≤ 0.1℃, response band 814μm, frame rate ≥ 25fps. During detection, a certain eavesdropping device generates heat during operation, with its temperature 1.5℃ higher than the surrounding environment. According to the formula... Its radiative output is approximately 1.06 times that of the surrounding environment, and thermal imagers can capture this temperature difference.

[0027] The working principle of optical anti-eavesdropping detection equipment: Utilizing the light reflection characteristics of a camera lens, it captures the reflective points of a hidden camera by emitting infrared or visible light. Based on the law of reflection of light: ,in Angle of incidence The angle of reflection is denoted by , and the position of the reflective point can be determined by the direction of the reflected light.

[0028] Optical anti-eavesdropping detection equipment uses an adaptive exposure algorithm to automatically adjust the exposure time based on the ambient light intensity. , where t is the exposure time, I is the ambient light intensity, and k is the adjustment coefficient.

[0029] Application example of optical anti-eavesdropping detection equipment: In a strong light environment, the ambient light intensity I=1000 lux, the adjustment coefficient k=500, and the calculated exposure time t=0.5ms, to avoid overexposure; in a weak light environment, I=50 lux, t=10ms, to ensure image clarity.

[0030] The working principle of a nonlinear node detector: Based on the nonlinear characteristics of a semiconductor PN junction, when a high-frequency signal (e.g., 1 GHz) is emitted towards the object being measured, the semiconductor device generates nonlinear response signals such as second harmonics (e.g., 2 GHz) and third harmonics (e.g., 3 GHz). By receiving these harmonic signals, the presence of electronic components can be determined. The harmonic frequencies... Where n is the harmonic order, The frequency of the transmitted signal.

[0031] Application examples of nonlinear node detectors: Transmit signal frequency =1.5GHz, detected A signal of 3 GHz can indicate the presence of nonlinear devices, such as transistors and integrated circuits.

[0032] The working principle of portable through-wall radar: It uses ultra-wideband (UWB) radar technology, which transmits short-pulse electromagnetic waves and receives the reflected echoes from walls and targets. The target depth is calculated based on the echo time delay. Where d is the target depth, c is the propagation speed of electromagnetic waves in the wall, and t is the time delay. The imaging algorithm of the portable through-wall radar is the synthetic aperture radar (SAR) algorithm, which improves azimuth resolution by moving the radar antenna to simulate a large-aperture antenna. ,in For azimuth resolution, λ is the wavelength, and L is the synthesized aperture length.

[0033] Application example of portable through-wall radar: Electromagnetic wave propagation speed in the wall c=0.15m / ns, echo time delay t=200ns, calculated target depth d=15cm; radar travel distance L=1m, operating frequency f=2GHz, wavelength... =0.15m, azimuth resolution =0.075m, can distinguish targets with a diameter ≥8cm.

[0034] The working principle of a spectrum analyzer is as follows: Based on the superheterodyne principle, the input signal is mixed with the signal generated by the local oscillator and converted into a fixed intermediate frequency signal. After filtering, amplification, and detection, the spectrum distribution of the signal is obtained.

[0035] Key specifications for a spectrum analyzer: frequency resolution ≤10kHz, dynamic range ≥80dB.

[0036] Example of spectrum analyzer application: A 2.45GHz signal was detected with a power spectral density of 20dBm / MHz, a bandwidth of 10MHz, and a total power of 20dBm / MHz + 10log(10×10^6) = 10dBm, which can be identified as a relatively strong wireless signal source.

[0037] Please refer to Figure 3. In a specific embodiment, step S102 includes: applying a wavelet transform denoising algorithm to the raw infrared thermal imaging data, selecting the db5 wavelet basis, performing two-level decomposition, and removing thermal noise from the image; processing the raw optical image data through histogram equalization to enhance the reflective point features; and normalizing the raw nonlinear response data, radar echo data, and electromagnetic spectrum data to convert them into values ​​in the [0,1] interval.

[0038] Specifically, for raw infrared thermal imaging data, the wavelet transform denoising algorithm, based on multi-resolution analysis, decomposes the signal into sub-bands of different frequencies, retaining low-frequency approximate components and removing high-frequency noise components. Taking the db5 wavelet as an example, the decomposition formula is as follows: in These are approximate coefficients. are detail coefficients, and h and g are wavelet filter coefficients.

[0039] For raw optical image data, the grayscale values ​​of image pixels are redistributed to make the image grayscale histogram as uniformly distributed as possible. The transformation formula is shown below: in The original grayscale value. The grayscale value after transformation. Let L be the gray level probability density, and L be the number of gray levels.

[0040] Add microsecond-accurate timestamps to the data from each device and calibrate using a hardware clock synchronization module. Establish a three-dimensional coordinate system to map the coordinates of the abnormal area of ​​infrared thermal imaging (X=3.2m, Y=2.5m, Z=1.8m), the coordinates of the optical reflective point (X=3.25m, Y=2.55m, Z=1.78m), the coordinates of the nonlinear node signal source (X=3.22m, Y=2.53m, Z=1.81m), the coordinates of the through-wall radar target (X=5.6m, Y=2.3m, Z=0.15m), and the spectral signal source (X=4.0m, Y=3.0m, Z=2.0m) to the same coordinate system.

[0041] Specifically, the NTP (Network Time Protocol) synchronization mechanism is adopted, and a hardware clock module (accuracy ±1μs) is used to add timestamps to the data of each device. The timestamp format is UTC time (year-month-day-hour:minute:second.microsecond).

[0042] Establish the transformation relationships between the World Coordinate System (WCS), Device Coordinate System (DCS), and Image Coordinate System (ICS). The transformation matrix from Device Coordinate System to World Coordinate System is shown below: in As a world coordinate system, Let R be the device coordinates, R be the rotation matrix, and T be the translation vector.

[0043] In this embodiment, a precise timestamp is added to the raw data of each modality using a network time protocol synchronization mechanism, ensuring accurate alignment of data collected at different times. This effectively solves the data mismatch problem caused by differences in sampling time, providing a consistent foundation in the temporal dimension for subsequent feature fusion. Simultaneously, by constructing a world coordinate system for the target region and the corresponding transformation matrix, and uniformly transforming the data of each modality to this world coordinate system, precise spatial registration of data from different acquisition devices is achieved. This overcomes spatial deviations caused by differences in the viewing angle, position, and resolution of the acquisition devices, enabling pixel coordinates in infrared images, point coordinates in radar echoes, and positioning information coordinates in the electromagnetic spectrum to be associated with the same physical space coordinate system. This provides a unified spatial reference for the spatial association of multi-source information and target positioning. Therefore, this time and space synchronization method ensures a high degree of consistency between multi-source heterogeneous data in the spatiotemporal dimensions, which is beneficial for subsequent effective feature fusion and accurate detection and positioning, significantly improving the accuracy and reliability of multi-modal data fusion.

[0044] Please refer to Figure 3. In a specific embodiment, step S103 includes: using a feature-level fusion method, inputting the preprocessed thermal imaging temperature gradient features, optical edge features, nonlinear node harmonic features, radar echo time-domain features, and spectral power spectrum features into a cross-modal attention fusion network to generate a fused feature vector.

[0045] Specifically, a cross-modal attention fusion network is employed, incorporating both self-attention and cross-attention mechanisms. The self-attention mechanism calculates the similarity between the query and the key, as shown in the following equation: Where Q, K, and V are the query, key, and value vectors, respectively. The dimension of the key vector.

[0046] Taking the fusion process of infrared thermal imaging synchronization data, optical image synchronization data, nonlinear response synchronization data, radar echo synchronization data, and electromagnetic spectrum synchronization data in a classified conference room detection scenario as an example, the following details the specific implementation steps of feature fusion in conjunction with the self-attention mechanism formula.

[0047] Step 1: Feature preprocessing and vector construction.

[0048] First, the infrared thermal imaging synchronization data, optical image synchronization data, nonlinear response synchronization data, radar echo synchronization data, and electromagnetic spectrum synchronization data are standardized and transformed into feature vectors with unified dimensions, laying the foundation for attention mechanism calculation.

[0049] Thermal imaging temperature gradient feature vector (F thEight key temperature gradient parameters of the abnormal region are extracted, such as horizontal gradient, vertical gradient, and diagonal gradient. After standardization, a vector of dimension 1×64 is constructed, with each gradient parameter corresponding to eight sampling points, denoted as... In one embodiment, the abnormal area data is: temperature 32°C, temperature difference +2°C.

[0050] Optical Edge Feature Vector (Fop): For the circular reflective point at the gap in the filing cabinet in the optical image, Canny edge detection is used to extract 12 edge features, such as edge length, curvature, and grayscale contrast. Each feature corresponds to sampled values ​​in 6 directions, which are standardized to construct a 1×72 vector, denoted as... .

[0051] Nonlinear nodal harmonic eigenvectors (F) no For the detected 45dBm nonlinear response signal, six features were extracted, including harmonic frequency, harmonic amplitude, and harmonic phase difference. Each feature corresponds to 10 time sampling points. After standardization, a 1×60 vector was constructed, denoted as... .

[0052] Radar echo time-domain eigenvector (F ra For a 30dB anomalous echo at a depth of 15cm within the wall, 10 temporal features, including echo peak value, rise time, and fall time, were extracted. Each feature corresponds to 8 depth sampling points. After standardization, a 1×80 vector was constructed, denoted as... .

[0053] Spectral power spectrum eigenvector (F sp For 2.45 GHz band signals, eight features are extracted, including center frequency, bandwidth, peak power spectrum, and modulation index. Each feature corresponds to 12 frequency sampling points. After standardization, a 1×96 vector is constructed, denoted as . .

[0054] By employing feature dimension alignment techniques, combining zero-padding and interpolation, the five types of feature vectors mentioned above are uniformly adjusted to a standard dimension of 1×128, resulting in an aligned feature matrix, which serves as the aligned feature matrix. ,in , , , and These are the aligned thermal imaging temperature gradient feature vector, optical edge feature vector, nonlinear node harmonic feature vector, radar echo time-domain feature vector, and spectral power spectrum feature vector, respectively.

[0055] Step 2: Calculation of self-attention mechanism.

[0056] Through 3 learnable linear transformation matrices , and The alignment feature matrix F is mapped to a query matrix, a key matrix, and a value matrix, respectively. The specific mapping process is shown in the following formula: Among them, W Q W K W V The initial parameters are set using the Xavier initialization method and then optimized through backpropagation during model training.

[0057] Calculate the product of the query matrix and the transpose of the key matrix to obtain the original similarity matrix S, then divide by... Scaling is applied to avoid gradient vanishing; where dk is the dimension of the Key vector, here dk=64. =8. The specific calculation process is shown in the following formula: Original similarity matrix: Original similarity matrix; Scaled similarity matrix: Taking this example, the calculated Sscaled matrix is ​​shown below, which serves as the self-attention fusion feature matrix, where the values ​​are for illustrative purposes: In the self-attention fusion feature matrix above, the element Sscaled[i][j] represents the similarity between the i-th type of feature and the j-th type of feature. For example, Sscaled[3][2]=0.7 indicates that the similarity between the radar echo time domain feature and the nonlinear node harmonic feature is 0.7.

[0058] Softmax normalization and attention weight calculation: Softmax normalization is performed on each row of the self-attention fusion feature matrix to obtain the attention weight matrix A, ensuring that the sum of the weights of each feature class to other features is 1. The attention weight matrix A calculated in this example (with values ​​rounded to two decimal places) is as follows: Among them, A[3][3]=0.58 indicates that the radar echo time domain feature has the highest attention weight to itself, and A[2][4]=0.20 indicates that the nonlinear node harmonic feature has an attention weight of 0.20 to the spectral power spectrum feature, which reflects the correlation between the two types of features, such as the correspondence between harmonic signals and electromagnetic spectrum.

[0059] Weighted summation yields self-attention features: Multiplying the attention weight matrix A by the value matrix V yields the self-attention fused feature matrix Fself-attention, completing the self-correlation fusion between unimodal features. Taking the behavior corresponding to the time-domain characteristics of radar echoes as an example, its fused features are calculated as follows: That is, by weighting, the key information of the other four types of features is incorporated into the time domain features of radar echoes. For example, by combining the harmonic information of nonlinear nodes, the echo feature recognition of electronic devices inside the wall can be optimized.

[0060] Step 3: Supplement and integrate the cross-attention mechanism.

[0061] Building upon self-attention fusion, a cross-attention mechanism is introduced to further enhance the complementarity between features from different modalities. The specific process is as follows: First, constructing the cross-modal query matrix and key matrix: The feature matrix F after self-attention fusion is... self-attention Based on modality, they are divided into two groups: "thermal imaging-optical" visual category and "nonlinear node-radar-spectrum" electromagnetic category. The feature matrix of the visual category is denoted as... The electromagnetic characteristic matrix is ​​denoted as .

[0062] Second, cross-similarity calculation: Using visual features as the query matrix and electromagnetic features as the key and value matrices, the cross-attention weights are calculated using the following formula: ,in This is a visual query mapping matrix.

[0063] ,in This is the key mapping matrix for the electromagnetic class.

[0064] ,in This is the Value mapping matrix for the electromagnetic class.

[0065] Cross-similarity matrix: , where dk-cross=48, is the dimension of the cross attention key.

[0066] Cross-attention weights: .

[0067] Third, cross-modal feature fusion: By integrating key information from electromagnetic features into visual features through cross-attention weights, a cross-attention fusion feature matrix F is obtained. cross As shown in the following formula: Among them, optical edge features in the visual category, such as circular reflective points, are further confirmed to determine whether the reflective points are electronic devices by cross-attention and combined with the 2.45GHz signal characteristics of the electromagnetic spectrum, such as the typical frequency band of wireless cameras, in order to exclude non-electronic reflective light sources such as metal ornaments.

[0068] Step 4: Final fusion feature output.

[0069] The self-attention fusion feature matrix F self-attention Fusing feature matrix F with cross-attention cross The features are concatenated and then compressed and integrated using a single convolutional layer (3×3 kernel, stride 1), ultimately outputting a global fusion feature vector F with a dimension of 1×128. final This vector contains key information on five types of original features and the correlation information between modalities. It can be directly input into the subsequent multimodal fusion model to achieve accurate identification and location of espionage devices.

[0070] In this example, the feature vector F is fused. final In the model, the weights of the feature dimensions related to "temperature gradient-optical reflection-harmonic signal-radar echo-spectral band" are significantly higher than those of other dimensions. Based on this feature vector, the model outputs the coordinates of the espionage device (X=3.23m, Y=2.54m, Z=1.8m), the type of espionage device (pinhole camera, confidence level 0.96), and the electromagnetic properties of the espionage device (2.45GHz, GFSK modulation), which perfectly match the actual detection results, verifying the effectiveness of the feature-level fusion method.

[0071] In this embodiment, the five types of data are first standardized and transformed into feature vectors with uniform dimensions, laying the foundation for subsequent self-attention fusion mechanism calculations. Subsequent feature dimension alignment constructs a unified structured representation aligned feature matrix, enabling heterogeneous features to be processed within the same feature matrix framework.

[0072] Subsequently, this invention introduces a self-attention mechanism to delve into the deep correlations between different types of feature vectors, thereby extracting more discriminative feature representations. Building upon this, the invention further employs a cross-attention mechanism to dynamically establish interdependencies between different modal features, achieving effective interaction and complementary enhancement of cross-modal data of different types. This dual attention mechanism enables the model to adaptively focus on key features that corroborate each other across different data, such as the correlation between suspicious contours in optical images and abnormal hotspots in infrared thermal imaging, and the correlation between specific signals in the electromagnetic spectrum, thus greatly improving the representational capability of the fused features. Finally, by integrating and compressing the self-attention and cross-attention features, the resulting fused feature vector highly condenses the key information and complex cross-modal correlations in multi-source heterogeneous data, effectively representing the key information of the five types of original data features and the correlation information between modalities. This provides a more reliable decision-making basis for the subsequent accurate detection and identification of espionage devices, thereby improving the accuracy of espionage device detection.

[0073] Please refer to Figure 4. In a specific embodiment, step S104 includes: constructing a dataset containing 5,000 samples, of which 3,000 samples are obtained from areas with eavesdropping devices and 4,000 samples are obtained from areas without eavesdropping devices. The samples cover 5 different types of eavesdropping devices, 2 types of wall materials, 10 types of ambient light intensity and electromagnetic interference intensity.

[0074] The multimodal fusion model consists of infrared thermal imaging, optical, nonlinear node, radar, and spectral branches. Each branch extracts features through convolutional layers, which are then fused via an attention mechanism and input into a fully connected layer for classification. The loss function is a multi-task loss function combining cross-entropy loss and mean squared error loss, as shown in the following equation: in For classifying losses, The regression loss is represented by α, which is a weighting coefficient of 0.5.

[0075] After obtaining the multimodal fusion model, the fused feature vector is input into the model, which contains 5 convolutional layers, 3 attention mechanism layers, and 2 fully connected layers. The multimodal fusion model includes the coordinates of the eavesdropping device (e.g., X=3.23m, Y=2.54m, Z=1.8m), the type of eavesdropping device (e.g., the probability of it being a pinhole camera), and the electromagnetic properties of the eavesdropping device (e.g., 2.45GHz, GFSK modulation).

[0076] In this embodiment, by constructing feature subnetworks for five modes—infrared thermal imaging, optical images, nonlinear response, radar echo, and electromagnetic spectrum—in a preset cross-modal attention network, the cross-modal attention network can selectively mine the deep features and inherent patterns within each type of data during training, avoiding feature confusion and information loss that may occur when a single network architecture processes heterogeneous data.

[0077] Secondly, the cross-entropy loss function is a standard loss function used for classification tasks such as device type, measuring the difference between the predicted probability distribution and the true distribution; the mean squared error loss function is a standard loss function used for regression tasks such as device coordinates and electromagnetic properties, measuring the squared difference between the predicted and true values. By combining the cross-entropy and mean squared error loss functions to construct a multi-task loss function, the model can simultaneously optimize multiple tasks during training, including classifying espionage device types, locating espionage device coordinates, and regressing espionage electromagnetic properties. This multi-task learning mechanism promotes mutual reinforcement and constraint among different tasks, not only improving the model's detection accuracy for each individual task but also achieving a regularization effect, enhancing the model's generalization ability and robustness, and effectively preventing overfitting. Finally, the multimodal fusion model trained on the dataset can collaboratively analyze the complex information contained in the fused feature vectors, achieving comprehensive and accurate judgment of the coordinates, type, and electromagnetic properties of espionage devices. This enables accurate detection and location of hidden espionage devices in classified locations, improving detection accuracy, environmental adaptability, and work efficiency.

[0078] In another possible embodiment, steps S103 and S104 are implemented using decision-level fusion rules, including: constructing three types of inference models for the detection of espionage devices for infrared thermal imaging synchronization data, optical image synchronization data, and nonlinear response synchronization data, respectively. The weight of each type of data corresponding to the espionage device detection inference model is determined based on its accuracy in different environments. For example, in a strong light environment, the weight of the prediction model corresponding to optical image synchronization data is 0.3, the weight of the model corresponding to infrared thermal imaging synchronization data is 0.5, and the weight of the model corresponding to nonlinear response synchronization data is 0.2. Finally, by combining the espionage device coordinates, espionage device type, and espionage device electromagnetic properties predicted by the three types of inference models, a comprehensive espionage device coordinate, comprehensive espionage device type, and comprehensive espionage device electromagnetic properties are obtained.

[0079] Taking the detection scenario in a high-light environment of a confidential conference room as an example, the specific steps are as follows: Step 1: Clarify the output format of the single data theft device detection model.

[0080] First, ensure that the eavesdropping device detection models for infrared thermal imaging synchronization data, optical image synchronization data, and nonlinear response synchronization data all output structured results. In this example, the output format and actual results for each device are as follows: Infrared thermal imaging synchronization data inference model output results: 1. Target location coordinates: X2=3.20m, Y2=2.50m, Z2=1.80m. 2. Eavesdropping device type: Suspicious electronic device, confidence level c2=0.92. 3. Electromagnetic properties: No output; the infrared thermal imager device has no electromagnetic detection function and is marked as "None".

[0081] Output of the optical image synchronization data inference model: 1. Target location coordinates: X1=3.25m, Y1=2.55m, Z1=1.78m. 2. Type of eavesdropping device: pinhole camera, confidence level c1=0.85. 3. Electromagnetic properties: No output; the optical anti-eavesdropping device has no electromagnetic detection function and is marked as "None".

[0082] Output results of the nonlinear response synchronous data inference model: 1. Target location coordinates: X3=3.22m, Y3=2.53m, Z3=1.81m. 2. Type of eavesdropping device: Electronic device containing semiconductor components, confidence level c3=0.88. 3. Electromagnetic properties: Nonlinear response signal strength 45dBm.

[0083] It should be noted that the output of the device must be consistent with its hardware capabilities. For example, optical / thermal imaging devices can only determine the location and type of the object based on "reflection / temperature" and cannot detect electromagnetic properties; nonlinear node detectors can determine "whether electronic components are present" and signal strength, but cannot identify specific frequency bands, so it is important to avoid forcibly outputting unfounded results.

[0084] Step 2: Weighted calculation of the coordinates of the espionage device.

[0085] For the X / Y / Z axis 3D coordinates output by the three types of inference models, a weighted average method is used for fusion. The core logic is: using the weight of each device as a coefficient, a weighted average is calculated for the X, Y, and Z axis coordinates respectively, which is then used as the final composite coordinate. Specific formulas and example calculations include: for any dimension of the coordinate, taking the X-axis as an example, the Y-axis and Z-axis are calculated similarly, resulting in the composite coordinate X... final The calculation method is as follows: Where w1, w2, and w3 are the weights of the infrared thermal imaging synchronization data, optical image synchronization data, and nonlinear response synchronization data, respectively. In this example, the sum of the weights is 0.3 + 0.5 + 0.2 = 1, and the denominator can be simplified to 1. X1, X2, and X3 are the X-axis coordinates output by each inference model.

[0086] The X-axis composite coordinate calculation is shown in the following formula: The Y-axis composite coordinate is calculated as follows: The Z-axis composite coordinate calculation is shown in the following formula: Based on the weighted calculation results of the three inference models, the final coordinates of the target are: X=3.24m, Y=2.54m, Z=1.80m.

[0087] Step 3: Weighted voting on the types of espionage devices.

[0088] For the eavesdropping device types and confidence levels output by the three types of inference models, a weighted confidence voting method is used for fusion. First, the type descriptions output by the devices are uniformly mapped to standard categories. Then, the weighted score of each category is calculated by multiplying the device weight by the type confidence level. The category with the highest score is the final type. The specific steps and example calculation are as follows: Step 1: Standard Category Mapping. Since the description dimensions of device types differ among devices (e.g., optical devices directly identify pinhole cameras, while thermal imagers can only identify suspicious electronic devices), it is necessary to first establish mapping rules from device output descriptions to standard categories. These rules serve as the output mapping rules for inference models corresponding to different categories of data. The mapping results for this example are shown in Table 1 below: Table 1 Step 2: Weighted Score Calculation and Final Type Determination. The weighted scores for each category are calculated, and the category with the highest score is the final device type. Simultaneously, the overall confidence score is calculated by dividing the sum of the confidence scores by the sum of the weights. In this example, the weight sum is 1, meaning the total score is 1: Category A, pinhole camera total score: 0.255.

[0089] Category B, Unknown Electronic Devices, Total Score: 0.460.

[0090] Category C, Electronic devices containing semiconductors: Total score: 0.176.

[0091] Furthermore, considering the limitations of the equipment capabilities, the result judgment was further optimized: Infrared thermal imagers can only determine whether an item is an electronic device based on "temperature anomalies," but cannot identify the specific type; while optical equipment can directly identify "pinhole cameras," although with lower confidence, the type is more specific. Therefore, a type priority rule was added to further optimize the result judgment. When the total score of type A, which is more specific than type B and type C in the output type description, is greater than 0.2, type A is selected first to avoid type ambiguity due to equipment limitations.

[0092] In this example, the Class A pinhole camera scores 0.255 ≥ 0.2, and its output type description is more specific. Therefore, based on the weighted calculation results of the three inference models, the target eavesdropping device type is determined to be: pinhole camera, with a comprehensive confidence level of 0.255 + 0.460 × 0.3 = 0.393. Here, 0.3 is the support coefficient of "unknown electronic device" for the "specific type," which is predefined in the system.

[0093] Step 4: Weighted Integration of Electromagnetic Attributes of Eavesdropping Devices. For the "electromagnetic attributes" output by the three types of devices (some devices have no output or incomplete output), a "weighted information supplementation method" is used for fusion. The core logic is: only the output of devices with electromagnetic detection capabilities is retained. Using device weights as coefficients, weighted values ​​are calculated for "quantifiable electromagnetic parameters" (such as signal strength), and for "qualitative information" (such as frequency bands and modulation methods), a "high-weight device priority" principle is applied. Specific steps and example calculations are as follows: Step 1: Screening of Effective Electromagnetic Attribute Outputs. Based on device capabilities, only nonlinear node detectors have electromagnetic attribute outputs; optical / thermal imaging devices have no output. The list of effective outputs is shown in Table 2 below: Table 2 The second step is to calculate the weighted quantitative parameters and supplement the qualitative information.

[0094] Signal strength quantization parameter: Overall signal strength = Sum of weighted parameter values ​​ / Sum of effective weights. Wherein, the effective weights are the weights of the output device, i.e., according to Table 2, the effective weight is 0.2. The specific calculation process is as follows: Overall signal strength = 9 / 0.2 = 45 dBm. Since no other device provides electromagnetic parameters, the overall signal strength result of the three inference models is consistent with the signal strength output of the corresponding inference model of the nonlinear node detector.

[0095] Qualitative frequency band and modulation method: Since none of the three types of equipment have frequency band / modulation method output, it is necessary to associate "spectrum analyzer data". That is, if the target area contains electromagnetic spectrum synchronization data of the spectrum analyzer, its output can be supplemented; this example does not have a spectrum analyzer, so it is marked as "to be supplemented".

[0096] Step 3, Final Electromagnetic Attribute Output: Electromagnetic Attribute: Signal Strength 45dBm.

[0097] Step 5: Summarize the final comprehensive results of the decision-level fusion. Combining the fusion results of the above three steps, in this embodiment, the final output of the decision-level fusion for the confidential conference room is as follows: Target location coordinates: X=3.24m, Y=2.54m, Z=1.80m; Eavesdropping device type: pinhole camera, comprehensive confidence level 0.393; Electromagnetic properties: signal strength 45dBm.

[0098] In one specific embodiment, the process involves: acquiring a monitoring signal frequency band based on the electromagnetic properties of the eavesdropping device; collecting signals from the target area based on the monitoring signal frequency band to obtain a monitoring signal data sequence; training a preset hidden Markov model based on the preset sample signal data sequence to obtain a state transition probability matrix, an observation probability matrix, and an initial state probability; then constructing a sample electromagnetic signal model based on the state transition probability matrix, the observation probability matrix, and the initial state probability; analyzing the matching probability between the monitoring signal data sequence and the sample electromagnetic signal model; and evaluating the confidence level of the eavesdropping device coordinates, the eavesdropping device type, and the electromagnetic properties of the eavesdropping device based on the matching probability. This includes: firstly, using a short-time Fourier transform (STFT) to extract the time-frequency characteristics of the target area signal in the monitoring signal frequency band, as shown in the following formula: in For window functions.

[0099] Based on the Hidden Markov Model (HMM), normal electromagnetic signals are modeled as a transition process of multiple states. When the matching probability between the observed signal and the model is lower than a preset probability threshold, it is judged as an abnormal signal.

[0100] Specifically, the construction of a normal electromagnetic signal model based on a Hidden Markov Model (HMM) involves three core steps: state definition, parameter learning, and state transition verification. These steps transform the time-frequency characteristics of the normal signal into quantifiable state transition rules. The specific process and application examples are as follows.

[0101] Step 1: Based on signal feature classification, define the hidden state of normal electromagnetic signals.

[0102] The "hidden states" of a Hidden Markov Model (HMM) need to correspond to the typical operating modes of normal electromagnetic signals and the signal characteristics of normal devices in real-world scenarios, such as routers and projectors. Three types of hidden states are defined, as shown in Table 3 below: Table 3 Example: During the normal signal sample collection phase in a confidential conference room, a spectrum analyzer was used to continuously collect the 2.45GHz band signal of a normal device, such as an office router, for one hour. It was found that the signal was in the "S1 low-power stable state" for more than 90% of the time, and only briefly switched to the "S2 medium-power intermittent state" during data transmission. There was no "S3 high-power short-term state". Therefore, the core states of the normal signal in this scenario are S1 and S2.

[0103] Step 2: Learn the HMM parameters of the normal signal to obtain the state transition probability and observation probability.

[0104] By collecting a large number of normal signal samples, such as 100 sets of normal 2.45GHz frequency band signal data for 10 minutes each, the core parameters of the Hidden Markov Model are trained to form a state transition model for normal signals. The key parameters and example calculations are as follows: Obtain the state transition probability matrix A, which is used to describe the probability of transitioning from one state to another.

[0105] Let A[i][j] represent the probability of transitioning from state Si to state Sj. Based on the statistical results of normal samples, the matrix A in the example is as follows: Where: A[1][1]=0.95 indicates that when the normal signal is in the low-power stable state of S1, the probability of it remaining in S1 at the next moment is 95%; A[1][2]=0.05 indicates that the probability of S1 transitioning to the power intermittent state of S2 is only 5%, corresponding to the occasional data transmission of normal equipment; A[1][3]=0 indicates that S1 will not directly transition to the high-power short-term state of S3, because normal office equipment does not have such high-frequency high-power signals.

[0106] Obtain the observation probability matrix B, which describes the probability of observing a specific signal feature under a certain state.

[0107] The observed values ​​are defined as the power spectral density range acquired by the spectrum analyzer, such as O1: ≤10dBm / MHz, O2: 10-18dBm / MHz, O3: 18-25dBm / MHz. B[i][k] represents the observed power spectral density range of O under state Si. k The probability of observation, and the observation probability matrix B in the example are as follows: Where: B[1][1]=0.98 indicates that the probability of observing O1 (≤10dBm / MHz) under the low power steady state of S1 is 98%, which is highly consistent with the signal characteristics of S1. B[2][2]=0.93 indicates that the probability of observing O2 (10-18dBm / MHz) under the power intermittent state of S2 is 93%, which is consistent with the power characteristics of this state.

[0108] Obtain the initial state probability π, which describes the probability of the signal being in each state at the initial moment.

[0109] Based on the initial signal feature statistics of normal samples, the π matrix in the example is as follows: π=[π1,π2,π3]=[0.99,0.01,0]. Among them, the normal signal is in the S1 low-power stable state with a 99% probability at the initial moment, and is in the S2 state with only a 1% probability due to the brief transmission when the device starts up. There is no S3 initial state.

[0110] Step 3: Verify the effectiveness of the state transition model and ensure it is consistent with normal signals.

[0111] The trained HMM parameters are substituted into the "forward algorithm" to calculate the matching probability of normal signal samples in the model, i.e., the probability that the sample conforms to the normal state transition rule. If the matching probability of more than 95% of normal samples is greater than or equal to the preset probability threshold, then the state transition model is effective, and this state transition model is used as the sample electromagnetic signal model. Example verification: A set of 2.45GHz signal samples from a normal router is taken and observed for 10 minutes, with the sequence [O1,O1,O1,O2,O1,O1,...]. The matching probability calculated by the forward algorithm is 128, which is greater than the preset probability threshold, indicating that the sample conforms to the normal model. This is repeated for 50 sets of normal samples, all of which satisfy the condition that the matching probability d is greater than the preset probability threshold, indicating that the model is effective.

[0112] The power and fluctuation period of normal electromagnetic signals strictly follow the state transition probability A and observation probability B of the sample electromagnetic signal model. For example, router signals will not frequently switch between S1 and S3, nor will the high power feature of O3 be continuously observed. However, the characteristics of abnormal signals such as espionage devices will break these rules, such as continuous high power and irregular fluctuations, which will cause their "matching probability" in the normal model to be far below the threshold. Specifically, the observation sequence of abnormal signals (such as [O3,O3,O3,...]) conflicts with the observation probability B of the normal model (such as B[1][3]=0, it is almost impossible to observe O3 in the S1 state); the state transition of abnormal signals (such as frequent transition from S3 to S2) conflicts with the transition probability A of the normal model (such as A[3][2]=0.02, the probability of transitioning from S3 to S2 is only 2%). The above conflicts will cause the matching probability calculated by the "forward algorithm" to be greatly reduced. When it is below the threshold, it can be judged as abnormal.

[0113] Taking the multimodal fusion model outputting a hidden wireless camera within a wall as the type of eavesdropping device, and the eavesdropping device's electromagnetic property as a 2.45GHz signal as an example, the determination process is explained in detail: Step 1, Acquire the observation sequence of the observed signal. The signal is continuously acquired for 1 minute using a spectrum analyzer, and the observation sequence is obtained by dividing it into 10-second observation units: O 观测 =[O3,O3,O3,O2,O3,O3]. Where O 观测 The signal characteristics are characterized by continuous high power, with only one brief drop to O2, which is completely different from the signal characteristics of normal equipment.

[0114] Step 2: Calculate the matching probability of the observation sequence in the normal HMM model. Using the "forward algorithm", combined with the trained state transition probability matrix A, observation probability matrix B and initial state probability π, calculate the matching probability of observation O: Initial time (t=1): Calculate the probability of observing O3 in each state: α1(1)=π1×B[1][3]=0.99×0=0; α1(2)=π2×B[2][3]=0.01×0.02=0.0002; α1(3)=π3×B[3][3]=0×0.95=0; Total probability at the initial time: α1(total)=0+0.0002+0=0.0002.

[0115] In subsequent time steps (t=2 to t=6), the forward algorithm is used for recursion. Due to the extremely low probability at the initial time step and the continuous occurrence of O3 in the observation sequence (which conflicts with the B matrix of the normal model), the final calculated matching probability of the observation sequence is 28, which is far below the preset probability threshold.

[0116] Step 3: Determine an anomaly and output the result. Since the matching probability of the observed signal is lower than the preset probability threshold, it indicates that the state transition characteristics of the signal do not match the electromagnetic signal model of the normal signal at all. Therefore, it is determined to be an abnormal electromagnetic signal. Further confirmation that the signal comes from a spying device indicates that the confidence level of the output spying device coordinates, the type of spying device, and the electromagnetic properties of the spying device in the multimodal fusion model is very high.

[0117] In this embodiment, by constructing a sample electromagnetic signal model based on a Hidden Markov Model, the present invention can deeply characterize the dynamic behavior and state transition patterns of normal signals in a target area when no eavesdropping device is present. Furthermore, by calculating the matching probability between the real-time acquired monitoring signal data sequence and this sample model, the preliminary detection results can be objectively and quantitatively verified, assessing the degree of matching between the electromagnetic properties of the eavesdropping device and the normal signals in the target area when no eavesdropping device is present. Finally, this matching probability is converted into a confidence assessment of the eavesdropping device's coordinates, type, and electromagnetic properties. A lower matching probability indicates a higher confidence level that the obtained electromagnetic properties of the eavesdropping device represent an abnormal electromagnetic signal, further confirming that the signal originates from the eavesdropping device. Conversely, a higher matching probability indicates a lower confidence level that the obtained electromagnetic properties of the eavesdropping device represent an abnormal electromagnetic signal, potentially indicating a misjudgment. This implementation not only effectively reduces false alarms and missed alarms caused by environmental interference or model misjudgment but also provides operators with reliable decision-making support, significantly improving the accuracy and reliability of eavesdropping device detection.

[0118] In one specific embodiment, the steps include: acquiring several sets of historical coordinates, historical types, and historical electromagnetic attributes of eavesdropping devices in the target area within a preset historical time period; constructing a device coordinate time-series index based on the historical coordinates of the eavesdropping devices; constructing a device type time-series index based on the historical types of the eavesdropping devices; constructing an electromagnetic attribute time-series index based on the historical electromagnetic attributes of the eavesdropping devices; determining the difference order, autoregression order, and moving average order of a preset time series prediction model based on the coordinate time-series index, the device type time-series index, and the electromagnetic attribute time-series index, thereby obtaining an initial probability prediction model; training the initial probability prediction model based on the coordinate time-series index, the device type time-series index, and the electromagnetic attribute time-series index to obtain a target probability prediction model; determining the target prediction time; and obtaining the probability of eavesdropping devices appearing in the target area within the target prediction time based on the target probability prediction model, including: using a data mining algorithm, employing the Apriori algorithm for association rule mining to find association patterns of eavesdropping device appearances in different scenarios. The support calculation formula is shown below: Where A and B are itemsets, and N is the total number of samples. A⇒B means A implies B, representing an association rule, that is, when itemset A appears, itemset B is likely to appear as well. support(A⇒B) represents the prevalence of this association rule across all samples, with the complete formula being support(A⇒B)=count(A∪B) / N, where N is the total number of samples, ranging from [0,1], with higher values ​​indicating a more prevalent rule. count(A∪B) represents the number of occurrences of the union of itemsets A and B.

[0119] A and B are both "itemsets", which are "sets of conditions / results" related to the confidentiality detection scenario, such as scenario type, environmental parameters, equipment type, etc.; A∪B represents a "combined event" that simultaneously contains itemsets A and B; count(A∪B) represents the total number of times the "combined event" actually occurs in all detection samples.

[0120] Based on several sets of historical coordinates, historical types, and historical electromagnetic properties of espionage devices, a time-series ARIMA model (p, d, q) is constructed, where p is the autoregressive order, d is the differencing order, and q is the moving average order. The optimal model parameters are determined using the Akaike Information Criterion (AIC).

[0121] Time-series indicators are extracted from historical monitoring data. Historical monitoring data contains multi-dimensional information such as the historical coordinates, types, and electromagnetic properties of the espionage devices. However, the ARIMA model only requires quantifiable indicators that change continuously over time. Therefore, the raw data needs to be transformed into indicators to select core indicators that meet the requirements of time-series modeling. Taking the historical monitoring data of a classified conference room from January 2023 to June 2024 (18 months) as an example, the extracted time-series indicators are shown in Table 4 below: Table 4 Among them, the historical coordinates, historical types, and historical electromagnetic properties of the espionage devices are discrete static data and cannot be directly used for ARIMA modeling. They need to be transformed into count / duration indicators that change over time, such as the frequency of occurrence of risk areas and the number of times the device type appears, forming a continuous sequence of "time (e.g., monthly) - indicator value" as a time-series indicator.

[0122] The core of the ARIMA model is to determine parameters through four steps: stationarity processing, autocorrelation analysis, partial autocorrelation analysis, and AIC criterion optimization. The following example uses historical data on the number of times a pinhole camera appears in a classified conference room each month (18 monthly data points: Y=[1,0,2,1,3,2,1,2,3,2,4,3,2,3,4,3,5,4]) to illustrate the parameter calculation process: 1. Determine the difference order d: to ensure the time series is stationary. The difference order d eliminates the trend of the time series through "difference operations," ensuring the series meets the stationarity requirements of ARIMA modeling, i.e., the mean and variance do not change over time.

[0123] Perform a first difference (d=1) on the original time series index sequence Y, and calculate the differenced sequence ΔYt=Yt-Yt-1 (t≥2); if the sequence is stationary after the first difference, then d=1; if it is still not stationary, perform a second difference (d=2).

[0124] Example calculation: Original sequence Y: [1,0,2,1,3,2,1,2,3,2,4,3,2,3,4,3,5,4] First difference sequence ΔY: [-1,2,-1,2,-1,-1,1,1,-1,2,-1,-1,1,1,-1,2,-1] Passing the "ADF stationarity test": The ADF statistic of the first difference sequence is -3.8 (less than the 1% significance level critical value of -3.2), the sequence is judged to be stationary, therefore d=1.

[0125] 2. Determine the autoregressive order p: This captures the historical dependencies of the sequence. The autoregressive order p indicates that the model needs to look back at the sequence values ​​of the previous p time points to predict the current value. For example, p=2 means using Yt-1 and Yt-2 to predict Yt. This is determined through autocorrelation function (ACF) analysis. The "lag order exceeding the confidence interval" in the ACF plot is a candidate value for p. Specifically, plot the ACF plot of the sequence ΔY after first differencing and observe the lag order at which the "ACF value first falls into the confidence interval (±1.96 / √n, n=17, approximately ±0.47)".

[0126] Example calculation: Calculate the ACF value (lag order 1-10) for the ΔY sequence: lag 1 = 0.35, exceeding the confidence interval; lag 2 = 0.28, exceeding the confidence interval; lag 3 = 0.15, falling within the confidence interval; lag 4 and above all fall within the confidence interval; combined with the partial autocorrelation function (PACF) for auxiliary verification: PACF decreases significantly after lag 2, therefore the autoregressive order p = 2, and the maximum lag order at which the ACF exceeds the confidence interval is taken.

[0127] 3. Determine the moving average order q: to eliminate random fluctuations in the series. The moving average order q indicates that the model needs to incorporate error terms from the first q time points to correct prediction bias; for example, q=1 means that εt-1 is used to correct the predicted value of Yt. It is determined through partial autocorrelation function (PACF) analysis. The lag order that exceeds the confidence interval in the PACF plot is the candidate value of the moving average order q. Operation method: Plot the PACF plot of the series ΔY after first differencing, and observe the lag order at which the PACF value first falls into the confidence interval.

[0128] Example calculation: Calculate PACF values ​​(lag order 1-10) for the ΔY sequence: lag 1 = 0.42, exceeding the confidence interval; lag 2 = 0.18, falling into the confidence interval; lag 3 and above all fall into the confidence interval; combined with ACF for auxiliary verification: ACF decreases significantly after lag 1, therefore q = 1.

[0129] 4. Determine the optimal parameter combination using the AIC criterion. The initially determined parameter combination (difference order p=2, autoregression order d=1, moving average order q=1) needs to be verified using the Akaike Information Criterion (AIC) – the smaller the AIC value, the better the model's "goodness-complexity balance", avoiding overfitting or underfitting.

[0130] Operation method: Construct multiple sets of candidate parameter combinations, such as (1,1,1), (2,1,1), (2,1,2), substitute them with historical data to calculate the AIC value, and select the combination with the smallest AIC as the optimal parameter.

[0131] Example calculations are shown in Table 5 below: Table 5 The AIC value of (2,1,1) is the smallest (45.6), so the optimal ARIMA model is determined to be ARIMA(2,1,1).

[0132] Based on the determined optimal parameters (2,1,1), the model is trained using the "number of times pinhole cameras appear" over 18 months as historical detection data. This data is then used to predict future indicator values. The specific process is as follows: Model Training: The first 15 months of data from the 18 months are used as the "training set" and substituted into the ARIMA(2,1,1) model to learn the time trend of the "historical occurrence count" (e.g., the average increase in occurrence count in the first week of each month is 0.2 times). Model Validation: The data from the last 3 months are used as the "test set" to verify the model's prediction accuracy—for example, the predicted occurrence count for the 16th month is 3.2 times, while the actual value is 3 times, with an error of only 0.2 times, meeting the "emergency prediction" requirement for confidentiality detection. Prediction Output: Based on the trained model, the "number of times pinhole cameras appear in the confidential conference room in the first week of next month" is predicted (e.g., a predicted value of 4.1 times). Combined with the scene correlation (support 3.2%), the final output is: "The probability of a pinhole camera appearing in the same location (15cm deep on the east wall) in this conference room is relatively high in the first week of next month; it is recommended to strengthen detection."

[0133] In this embodiment, by collecting and analyzing historical detection data of the target area, time-series indicators of device coordinates, type, and electromagnetic properties are constructed. A target probability prediction model is then built to deeply explore the potential patterns and trends of espionage devices appearing in the target area over time. The trained target probability prediction model can scientifically predict the probability of espionage devices appearing within a specific future timeframe, thereby guiding personnel to investigate espionage devices in the target area and enhancing the intelligence level of the security protection system for classified locations.

[0134] In one specific embodiment, the step of establishing a detection report data template based on the Extensible Markup Language (XML) data storage format, wherein the detection report data template includes several data fields; mapping relationships are constructed between the coordinates of the espionage device, the type of the espionage device, the electromagnetic properties of the espionage device, the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data and the several data fields, respectively, to obtain a mapping relationship set; and an espionage device detection report is generated based on the mapping relationship set and the detection report data template, including: 1. A report template management unit. Template structure: The report template is defined using XML format, including information such as chapter structure, data fields, and format styles. For example, the chapter structure definition of a detailed report template is as follows:<sectionid="overview"> Testing Overview<sectionid="targets"> Target device information<sectionid="analysis"> Data analysis process<sectionid="conclusion"> Conclusions and Recommendations 2. Data Filling Algorithm. Establish a mapping relationship between the analysis results of the multimodal fusion model and the report fields. The specific rules are as follows: Coordinates of the espionage device: Mapped to the "Location" field in the "Target Device Information" section, presented in three-dimensional coordinates (X, Y, Z), and the measurement error is marked (e.g., "X=3.23m, Y=2.54m, Z=1.8m, error ±3cm").

[0135] The type and confidence level of the espionage device are mapped to the "Type" field in the "Target Device Information" section, with the format "Type: Pinhole Camera, Confidence Level: 96%".

[0136] Electromagnetic properties of espionage devices: mapped to the "Electromagnetic Signal Analysis" sub-item in the "Data Analysis Process" chapter, including frequency band, power spectral density, and modulation method (e.g., "Frequency band: 2.45GHz, power spectral density: 20dBm / MHz, modulation method: GFSK").

[0137] Infrared thermal imaging synchronous data: Mapped to the "Thermal Imaging Analysis" sub-item in the "Data Analysis Process" section, recording the temperature values ​​of abnormal areas and their temperature difference with the environment (e.g., "Abnormal temperature: 32℃, temperature difference: +2℃").

[0138] Optical image synchronization data: mapped to the "Optical Anti-Eavesdropping Analysis" sub-item in the "Data Analysis Process" section, including the location and shape description of the reflection point (e.g., "Location: gap in the filing cabinet, Shape: circular reflection").

[0139] Nonlinear response synchronization data: mapped to the "Nonlinear Node Analysis" sub-item in the "Data Analysis Process" section, recorded in dBm (e.g., "Signal Strength: 45dBm").

[0140] Radar echo synchronization data: mapped to the "Through-Wall Radar Analysis" sub-item in the "Data Analysis Process" section, including target depth and reflection intensity (e.g., "Depth: 15cm, Reflection Intensity: 30dB").

[0141] Testing time and environmental parameters: Mapped to the "Testing Overview" section, including test start / end time, ambient temperature, humidity, light intensity, etc. (e.g., "Testing time: July 1, 2025, 10:00-10:05, Ambient temperature: 25℃, Humidity: 50%RH, Light intensity: 500lux").

[0142] Multimodal fusion model analysis conclusions: mapped to the "Detection Conclusions" sub-item in the "Conclusions and Recommendations" section, integrating the target's nature, location, and risk level (e.g., "Conclusion: A pinhole camera was found 15cm deep in the east wall of the conference room, and was determined to be a high-risk eavesdropping device").

[0143] The predictive model suggests mapping to the "Improvement Suggestions" sub-item in the "Conclusions and Recommendations" section, and generating preventative measures by combining historical data (e.g., "Recommendation: Strengthen the detection of similar locations in this conference room during the first week of next month, as historical probability shows that the frequency of eavesdropping devices is relatively high during this period").

[0144] Automatic image association: Based on the target location coordinates, automatically match and insert the original image and processed image (such as thermal image, optical image, radar echo image) of the corresponding device, and label "Image X: Thermal image of target area (location: X=3.23m, Y=2.54m)" below the image.

[0145] Risk level identification: Based on the confidence level and target type of the multimodal fusion model analysis, key information in the report is marked with color, such as red for high-risk targets, yellow for medium-risk targets, and green for low-risk targets.

[0146] Data traceability links: Add hyperlinks to key data in the report, such as abnormal electromagnetic signals or nonlinear node responses. Clicking on these links will take you to the original data page of the corresponding device, which includes timestamps, waveforms, etc., to facilitate subsequent traceability and verification.

[0147] In this embodiment, an Extensible Markup Language (XML) format is used to establish a data template and construct a standardized detection report. This ensures the report's structure and machine readability, facilitating data exchange and long-term archiving between different systems. Simultaneously, by establishing a complete mapping relationship between the detection data and the report template data fields, a report on the detection of espionage devices can be generated quickly, significantly improving work efficiency.

Claims

1. A method for detecting espionage devices based on data fusion, characterized in that, include: Determine the target area and acquire the raw infrared thermal imaging data, optical image data, nonlinear response data, radar echo data, and electromagnetic spectrum data of the target area; The raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data are synchronized in time and space to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data. Feature fusion is then performed on the synchronized infrared thermal imaging data, the synchronized optical image data, the synchronized nonlinear response data, the synchronized radar echo data, and the synchronized electromagnetic spectrum data to obtain a fused feature vector. A multimodal fusion model is constructed based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fused feature vector, thereby completing the detection of the espionage device.

2. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The process of synchronizing the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data in time and space to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data includes: adding timestamps to the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data based on a network time protocol synchronization mechanism to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data; determining a world coordinate system based on the target area and then constructing a world coordinate system transformation matrix; and performing coordinate transformations on the synchronized infrared thermal imaging data, the synchronized optical image data, the synchronized nonlinear response data, the synchronized radar echo data, and the synchronized electromagnetic spectrum data based on the world coordinate system transformation matrix to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data.

3. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The process of fusing features from the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data to obtain a fused feature vector includes: performing feature normalization processing on the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data respectively to obtain a thermal imaging temperature gradient feature vector, an optical edge feature vector, a nonlinear node harmonic feature vector, a radar echo time-domain feature vector, and a spectral power spectrum feature vector; aligning the feature dimensions of the thermal imaging temperature gradient feature vector, the optical edge feature vector, the nonlinear node harmonic feature vector, the radar echo time-domain feature vector, and the spectral power spectrum feature vector to construct an aligned feature matrix; mapping the aligned feature matrix to a query matrix, a key matrix, and a value matrix, and then obtaining a self-attention fusion feature matrix based on the query matrix, the key matrix, and the value matrix; obtaining a cross-attention fusion feature matrix based on the self-attention fusion feature matrix; and performing dimensional compression and feature integration on the self-attention fusion feature matrix and the cross-attention fusion feature matrix to obtain a fused feature vector.

4. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The construction of a multimodal fusion model based on a preset cross-modal attention network includes: acquiring training data samples based on a preset training region, and then constructing a dataset based on the training data samples; constructing an infrared thermal imaging feature subnetwork, an optical feature subnetwork, a nonlinear node feature subnetwork, a radar feature subnetwork, and a spectral feature subnetwork in the preset cross-modal attention network, and constructing a multi-task loss function in the cross-modal attention network based on a preset cross-entropy loss function and a preset mean square error loss function, thereby obtaining a multimodal fusion basic model; and training the multimodal fusion basic model based on the dataset to obtain a multimodal fusion model.

5. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The process of constructing a multimodal fusion model based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fusion feature vector, and after completing the detection of the espionage device, further includes: obtaining a monitoring signal frequency band based on the electromagnetic properties of the espionage device; collecting signals from the target area based on the monitoring signal frequency band to obtain a monitoring signal data sequence; training a preset hidden Markov model based on the preset sample signal data sequence to obtain a state transition probability matrix, an observation probability matrix, and an initial state probability, and then constructing a sample electromagnetic signal model based on the state transition probability matrix, the observation probability matrix, and the initial state probability; analyzing the matching probability between the monitoring signal data sequence and the sample electromagnetic signal model; and evaluating the confidence level of the coordinates, type, and electromagnetic properties of the espionage device based on the matching probability.

6. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The process of constructing a multimodal fusion model based on a preset cross-modal attention network, enabling the multimodal fusion model to output the coordinates, type, and electromagnetic properties of the espionage device according to the fusion feature vector, and completing the espionage device detection, further includes: acquiring several sets of historical coordinates, historical types, and historical electromagnetic properties of espionage devices in the target area within a preset historical time period; constructing a time-series index for device coordinates based on several historical coordinates of the espionage devices; constructing a time-series index for device types based on several historical types of the espionage devices; constructing a time-series index for electromagnetic properties based on several historical electromagnetic properties of the espionage devices; determining the difference order, autoregression order, and moving average order of a preset time series prediction model based on the coordinate time-series index, the device type time-series index, and the electromagnetic property time-series index, thereby obtaining an initial probability prediction model; training the initial probability prediction model based on the coordinate time-series index, the device type time-series index, and the electromagnetic property time-series index to obtain a target probability prediction model; determining the target prediction time; and obtaining the probability of the espionage device appearing in the target area within the target prediction time based on the target probability prediction model.

7. The method for detecting espionage devices based on data fusion according to claim 1, characterized in that, The process of constructing a multimodal fusion model based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fusion feature vector, and after completing the detection of the espionage device, further includes: establishing a detection report data template based on an extensible markup language data storage format, the detection report data template including several data fields; constructing mapping relationships between the coordinates, type, electromagnetic properties, infrared thermal imaging synchronization data, optical image synchronization data, nonlinear response synchronization data, radar echo synchronization data, and electromagnetic spectrum synchronization data and several data fields respectively, to obtain a mapping relationship set; and generating an espionage device detection report based on the mapping relationship set and the detection report data template.

8. A data fusion-based system for detecting espionage devices, characterized in that, include: The data acquisition module is used to determine the target area and acquire the raw infrared thermal imaging data, optical image data, nonlinear response data, radar echo data, and electromagnetic spectrum data of the target area. The data synchronization module is used to perform time and space synchronization on the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data, so as to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data. The data fusion module is used to perform feature fusion on the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data to obtain a fused feature vector; The espionage device detection and analysis module is used to construct a multimodal fusion model based on a preset cross-modal attention network, so that the multimodal fusion model outputs the coordinates, type, and electromagnetic properties of the espionage device according to the fusion feature vector, thereby completing the detection of the espionage device.

9. A data fusion-based eavesdropping device detection system according to claim 8, characterized in that, The process of synchronizing the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data in time and space to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data includes: adding timestamps to the raw infrared thermal imaging data, the raw optical image data, the raw nonlinear response data, the raw radar echo data, and the raw electromagnetic spectrum data based on a network time protocol synchronization mechanism to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data; determining a world coordinate system based on the target area and then constructing a world coordinate system transformation matrix; and performing coordinate transformations on the synchronized infrared thermal imaging data, the synchronized optical image data, the synchronized nonlinear response data, the synchronized radar echo data, and the synchronized electromagnetic spectrum data based on the world coordinate system transformation matrix to obtain synchronized infrared thermal imaging data, synchronized optical image data, synchronized nonlinear response data, synchronized radar echo data, and synchronized electromagnetic spectrum data.

10. A data fusion-based eavesdropping device detection system according to claim 8, characterized in that, The process of fusing features from the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data to obtain a fused feature vector includes: performing feature normalization processing on the infrared thermal imaging synchronization data, the optical image synchronization data, the nonlinear response synchronization data, the radar echo synchronization data, and the electromagnetic spectrum synchronization data respectively to obtain a thermal imaging temperature gradient feature vector, an optical edge feature vector, a nonlinear node harmonic feature vector, a radar echo time-domain feature vector, and a spectral power spectrum feature vector; aligning the feature dimensions of the thermal imaging temperature gradient feature vector, the optical edge feature vector, the nonlinear node harmonic feature vector, the radar echo time-domain feature vector, and the spectral power spectrum feature vector to construct an aligned feature matrix; mapping the aligned feature matrix to a query matrix, a key matrix, and a value matrix, and then obtaining a self-attention fusion feature matrix based on the query matrix, the key matrix, and the value matrix; obtaining a cross-attention fusion feature matrix based on the self-attention fusion feature matrix; and performing dimensional compression and feature integration on the self-attention fusion feature matrix and the cross-attention fusion feature matrix to obtain a fused feature vector.