WiFi pinhole camera detection method based on wireless signal feature analysis
By analyzing wireless signal characteristics, extracting key frames and energy change characteristics of video data, and combining deep neural network verification, the misjudgment problem of WiFi pinhole camera detection in existing technologies is solved, and high-precision and reliable detection effects are achieved.
Patent Information
- Application Number
- CN202511102969.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing WiFi pinhole camera detection technology has a high misjudgment rate in complex network environments, making it difficult to distinguish pinhole cameras from ordinary WiFi devices. It lacks a reliable verification mechanism and cannot effectively utilize the characteristics of video data transmission, resulting in insufficient detection accuracy.
By collecting wireless signal data, extracting the data packet features and energy change features of key frames and non-key frames, combining deep neural networks for feature fusion and verification, building a double verification mechanism, generating device identification coefficients and credibility scores, and confirming the existence of pinhole cameras.
It improves the recognition accuracy of the pinhole camera, significantly reduces the false alarm rate, and improves the reliability of the detection system. It is suitable for non-contact detection, easy to operate and has a wide range of applications.
Smart Images

Figure CN120602644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to wireless signal analysis technology, and in particular to a WiFi pinhole camera detection method based on wireless signal feature analysis. Background Art
[0002] Existing WiFi pinhole camera detection technology has the following defects and shortcomings: Existing technologies mainly rely on simple signal strength or traffic analysis, which cannot effectively distinguish the network behavior of pinhole cameras from ordinary WiFi devices. Especially in complex network environments, the misjudgment rate is high and the detection accuracy is insufficient.
[0003] Existing technologies lack in-depth analysis of the characteristics of video data transmission and fail to fully utilize the periodic characteristics of key frames and non-key frames in video encoding, resulting in incomplete feature extraction for identifying pinhole cameras and difficulty in dealing with various types of pinhole camera devices.
[0004] Existing technologies generally lack reliable verification mechanisms and are unable to evaluate the credibility of features. In complex and changeable wireless channel environments, they are easily affected by environmental interference and signals from other devices, resulting in false positives or negative positives, making it difficult to meet the reliability requirements in practical applications. Summary of the Invention
[0005] The embodiments of the present invention provide a WiFi pinhole camera detection method based on wireless signal feature analysis, which can solve the problems in the prior art.
[0006] A first aspect of an embodiment of the present invention provides a WiFi pinhole camera detection method based on wireless signal feature analysis, comprising: Collecting wireless signals within a target area, and obtaining data packet sequences and energy change data of the wireless signals; Based on the data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to the transmission period and size distribution. The occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted to establish the data packet timing characteristics that characterize the video data transmission characteristics; Extracting the mutation point, gradual change interval, and periodic fluctuation interval from the energy change data, and establishing an energy characteristic chain characterizing the working state of the equipment based on the mutation point, the power climbing curve of the gradual change interval, and the time series evolution law of the periodic fluctuation interval; Performing feature fusion on the data packet timing features and the energy feature chain to generate a feature matching coefficient for the device; performing time series analysis on the feature matching coefficient using a deep neural network, and generating a device identification coefficient at the output layer of the deep neural network; Constructing a feature verification network, wherein the feature verification network evaluates the sustained stability of the data packet timing feature and the degree of channel interference in the target area to generate a feature credibility score; When both the device recognition coefficient and the feature credibility score meet preset conditions, it is confirmed that a pinhole camera has been detected.
[0007] In an optional embodiment, Based on the data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to the transmission period and size distribution, the occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted, and the steps of establishing the data packet timing characteristics that characterize the video data transmission characteristics include: Extracting transmission time intervals between adjacent data packets in the data packet sequence, and determining a transmission period level of the data packet according to a clustering result of the transmission time intervals; Adaptively classifying the data packets in the data packet sequence according to the relationship between the data packet size and the mean and standard deviation of the data packet size, and classifying the data packets that are larger than the product of the mean and standard deviation of the data packet size and belong to the same transmission cycle level as candidate key frame data packets; Calculating a time interval sequence between the candidate key frame data packets, verifying the candidate key frame data packets according to the consistency of the time interval sequence, determining the candidate key frame data packets that pass the verification as key frame data packets, and determining the remaining data packets as non-key frame data packets; Extracting periodic features of the key frame data packets, the periodic features including the occurrence period of the key frame data packets, the coefficient of variation of the occurrence period, and the size mean of the key frame data packets; counting the duration and data packet density of non-key frame data packets located between adjacent key frame data packets, the data packet density being the number of non-key frame data packets within a unit time window; An initial timing feature vector is constructed based on the periodic characteristics, the duration of the non-key frame data packets and the data packet density; multiple initial timing feature vectors are calculated within a sliding time window, and a mean vector of the multiple initial timing feature vectors is calculated, and the initial timing feature vectors are weighted according to the degree of deviation between the initial timing feature vector and the mean vector to generate a final timing feature vector, wherein the weight coefficient is inversely proportional to the degree of deviation.
[0008] In an optional embodiment, The steps of establishing an energy characteristic chain representing the working state of the device according to the mutation point, the power climbing curve in the gradual change interval, and the time series evolution law of the periodic fluctuation interval include: Performing signal smoothing processing on the energy change data, segmenting the smoothed signal to obtain a sudden change point, a gradual change interval, and a periodic fluctuation interval; Extracting energy characteristics of the mutation point, the energy characteristics including mutation amplitude and mutation duration, and calculating the ratio of the mutation amplitude to the mutation duration to obtain a mutation response characteristic; Extracting the power ramp-up characteristics of the gradual change interval, calculating a first time constant and a first amplitude coefficient representing initial power supply characteristics, and calculating a second time constant and a second amplitude coefficient representing activation characteristics of the image acquisition module; Obtaining a power change sequence in the periodic fluctuation interval, calculating the amplitude distribution characteristics of the power change sequence, segmenting the power change sequence based on the amplitude distribution characteristics, extracting the time series change characteristics of each segment, performing time-frequency analysis on the power change sequence to obtain a time-frequency characteristic graph, and extracting the main frequency component and phase evolution law of the periodic fluctuation based on the time-frequency characteristic graph; calculating the energy concentration of the main frequency component and the stability index of the phase evolution law; establishing a corresponding relationship between the environmental load and the power response based on the time series change characteristics, the energy concentration and the stability index, and calculating the basic power value and the power response coefficient; An energy characteristic chain is established based on the sudden change response characteristic, the first time constant, the first amplitude coefficient, the second time constant, the second amplitude coefficient, the basic power value and the power response coefficient. The energy characteristic chain characterizes the energy change process of the device from the initial state to the stable working state.
[0009] In an optional embodiment, The step of fusing the data packet timing feature and the energy feature chain to generate a feature matching coefficient for the device includes: Preprocess the data packet timing features to obtain normalized timing features; Calculating the time dimension coupling coefficient, frequency dimension coupling coefficient, and phase dimension coupling coefficient of the normalized time series feature and the energy feature chain, wherein the time dimension coupling coefficient represents the consistency of the feature change rate, the frequency dimension coupling coefficient represents the correlation of the spectrum density, and the phase dimension coupling coefficient represents the stability of the phase difference; constructing a feature correlation matrix based on the time dimension coupling coefficient, the frequency dimension coupling coefficient, and the phase dimension coupling coefficient; Constructing a dual-stream feature extraction network to extract features from the normalized time series features and the energy feature chain, and generating a time series feature mapping matrix and an energy feature mapping matrix; Cross-calculating the temporal feature mapping matrix and the energy feature mapping matrix and weighting them with the feature correlation matrix to obtain a cross-attention weight; enhancing the temporal feature mapping matrix and the energy feature mapping matrix according to the cross-attention weight, and aligning and fusing the enhanced features to obtain a fused feature; A dynamic confidence is calculated according to a gradient change of the fused feature, and a feature matching coefficient is generated based on a similarity between the fused feature and a reference feature and the dynamic confidence.
[0010] In an optional embodiment, The step of using a deep neural network to perform time series analysis on the feature matching coefficients, wherein the output layer of the deep neural network generates a device identification coefficient, includes: Generate time series statistical features based on the time series sequence of feature matching coefficients; calculate a feature stability score based on the time series statistical features, construct a multi-scale memory pool to store feature patterns of different time scales, and determine the update weight of the memory pool according to the feature stability score, where the update weight is inversely proportional to the feature stability score; Extracting time series features from the time series sequence through a deep neural network to obtain a time series feature vector, calculating the similarity between the time series feature vector and the feature pattern stored in the multi-scale memory pool, and selecting the optimal memory feature based on the similarity; Calculating the Mahalanobis distance between the time series feature vector and its mean to obtain an anomaly score, mapping the anomaly score to obtain a suppression weight, multiplying the suppression weight by the time series feature vector to obtain a suppressed feature vector; performing attention calculation on the suppressed feature vector to obtain a device working mode feature, calculating a steady-state weight based on a gradient value of the time series feature vector; and multiplying the steady-state weight by the device working mode feature to obtain an enhanced device feature; Constructing an evaluation index based on historical recognition results, calculating an update amount based on the gradient of the evaluation index, and adaptively adjusting the decision threshold; The optimal memory feature, the enhanced device feature, and the steady-state weight are subjected to feature splicing and nonlinear transformation, a device recognition score is calculated through a multi-layer perceptron, and a device recognition coefficient is generated according to the decision threshold and the device recognition score.
[0011] In an optional embodiment, The steps of constructing a multi-scale memory pool and dynamically storing and updating feature patterns include: The time scale is divided into three levels: short-term, medium-term, and long-term. The sampling period of each level is dynamically adjusted based on the feature change rate, and the time series statistical features of each level are collected; the time series statistical features are extracted to obtain feature patterns, the time domain consistency and frequency domain correlation of the feature patterns are calculated, and the feature stability score is generated based on the time domain consistency and the frequency domain correlation; Calculating the information entropy of the feature stability score, determining the capacity of each level of memory pool according to the information entropy, expanding the memory pool capacity when the information entropy increases, and compressing the memory pool capacity when the information entropy decreases; Calculating an update weight based on the feature stability score, where the update weight is inversely proportional to the feature stability score; calculating the similarity between the feature pattern and the patterns stored in the memory pool to obtain a difference score; calculating the similarity between the feature pattern and the current working state feature to obtain a representative score; and using the difference score and the representative score as adjustment factors for the update weight; The access time and usage frequency of the characteristic pattern are recorded, a timeliness score is calculated based on the access time and usage frequency, and the memory pool storage space is released when the timeliness score is lower than a preset score threshold.
[0012] In an optional embodiment, Constructing a feature verification network, wherein the feature verification network evaluates the sustained stability of the data packet timing feature and the channel interference level of the target area, and the steps of generating a feature credibility score include: Inputting the data packet timing features into a multi-scale feature extraction layer, the multi-scale feature extraction layer including convolution kernels with different receptive fields, extracting features from the data packet timing features to obtain multi-scale features, calculating the coefficient of variation and autocorrelation coefficient of the multi-scale features, constructing a stability matrix, and calculating a weighted stability score based on the stability matrix; Establishing a WiFi channel interference feature library, the WiFi channel interference feature library including a channel overlap feature template, a multipath interference feature template, and a concurrent transmission feature template; extracting channel state information of a target area; calculating a first matching degree between the channel state information and the channel overlap feature template, a second matching degree between the channel state information and the multipath interference feature template, and a third matching degree between the channel state information and the concurrent transmission feature template; and calculating a channel interference degree score based on the first matching degree, the second matching degree, and the third matching degree; Constructing a dual-branch feature verification network, respectively extracting a timing feature vector of the timing feature of the data packet and an energy feature vector of the channel state information, calculating the mutual information between the timing feature vector and the energy feature vector, constructing an attention matrix, and calculating a feature consistency score based on the mutual information and the attention matrix; The weighted stability score, the channel interference degree score and the feature consistency score are weightedly fused to obtain a feature credibility score.
[0013] When both the device recognition coefficient and the feature credibility score meet preset conditions, the step of confirming that the pinhole camera is detected includes: When the device recognition coefficient is greater than the adaptive threshold calculated based on the historical detection accuracy and the feature credibility score meets the minimum credibility requirement, it is confirmed that a pinhole camera has been detected, and a threat level score is calculated comprehensively based on the device recognition coefficient and the feature credibility score.
[0014] According to a second aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0015] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0016] The present invention realizes effective detection of WiFi pinhole cameras by analyzing wireless signal characteristics, which has the following beneficial effects: The present invention extracts timing features such as key frame period, non-key frame duration and the ratio of the two from the data packet sequence, and constructs an energy feature chain based on the mutation points, gradual change intervals and periodic fluctuation intervals in the energy change data, thereby achieving accurate characterization of video transmission characteristics and effectively improving the recognition accuracy of pinhole cameras.
[0017] The present invention uses a deep neural network to perform timing analysis on the fused feature matching coefficients, and evaluates the stability of the data packet timing characteristics and the degree of channel interference through a feature verification network, constructing a double verification mechanism, which significantly reduces the false alarm rate and improves the reliability of the detection system.
[0018] The present invention does not rely on specific protocols or hardware devices, and can achieve non-contact detection of pinhole cameras based solely on wireless signal analysis. It has the characteristics of simple operation and wide applicability, and provides an effective technical means for protecting personal information security. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the process of a WiFi pinhole camera detection method based on wireless signal feature analysis according to an embodiment of the present invention; Figure 2This is the architecture diagram of the WiFi pinhole camera timing analysis system based on deep neural network. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0021] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0022] Figure 1 FIG. 1 is a flow chart of a WiFi pinhole camera detection method based on wireless signal feature analysis according to an embodiment of the present invention. Figure 1 As shown, the method includes: Collecting wireless signals within a target area, and obtaining data packet sequences and energy change data of the wireless signals; Based on the data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to the transmission period and size distribution. The occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted to establish the data packet timing characteristics that characterize the video data transmission characteristics; Extracting the mutation point, gradual change interval, and periodic fluctuation interval from the energy change data, and establishing an energy characteristic chain characterizing the working state of the equipment based on the mutation point, the power climbing curve of the gradual change interval, and the time series evolution law of the periodic fluctuation interval; Performing feature fusion on the data packet timing features and the energy feature chain to generate a feature matching coefficient for the device; performing time series analysis on the feature matching coefficient using a deep neural network, and generating a device identification coefficient at the output layer of the deep neural network; Constructing a feature verification network, wherein the feature verification network evaluates the sustained stability of the data packet timing feature and the degree of channel interference in the target area to generate a feature credibility score; When both the device recognition coefficient and the feature credibility score meet preset conditions, it is confirmed that a pinhole camera has been detected.
[0023] In an optional embodiment, based on data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to transmission period and size distribution, and the occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted to establish the data packet timing characteristics representing the characteristics of video data transmission. The steps include: Extracting transmission time intervals between adjacent data packets in the data packet sequence, and determining a transmission period level of the data packet according to a clustering result of the transmission time intervals; Adaptively classifying the data packets in the data packet sequence according to the relationship between the data packet size and the mean and standard deviation of the data packet size, and classifying the data packets that are larger than the product of the mean and standard deviation of the data packet size and belong to the same transmission cycle level as candidate key frame data packets; Calculating a time interval sequence between the candidate key frame data packets, verifying the candidate key frame data packets according to the consistency of the time interval sequence, determining the candidate key frame data packets that pass the verification as key frame data packets, and determining the remaining data packets as non-key frame data packets; Extracting periodic features of the key frame data packets, the periodic features including the occurrence period of the key frame data packets, the coefficient of variation of the occurrence period, and the size mean of the key frame data packets; counting the duration and data packet density of non-key frame data packets located between adjacent key frame data packets, the data packet density being the number of non-key frame data packets within a unit time window; An initial timing feature vector is constructed based on the periodic characteristics, the duration of the non-key frame data packets and the data packet density; multiple initial timing feature vectors are calculated within a sliding time window, and a mean vector of the multiple initial timing feature vectors is calculated, and the initial timing feature vectors are weighted according to the degree of deviation between the initial timing feature vector and the mean vector to generate a final timing feature vector, wherein the weight coefficient is inversely proportional to the degree of deviation.
[0024] For example, during the data transmission process of video coding, key frame and non-key frame data packets have obvious timing and size characteristics. This embodiment provides a method for adaptively identifying these characteristics and constructing a timing feature vector.
[0025] For a sequence of acquired video data packets, the time difference between each two adjacent packets is calculated to obtain a sequence of transmission time intervals. For example, for a sequence containing 1000 packets, 999 time interval values are obtained. These time interval values are clustered using K-means with the number of clusters set to 3. This yields three transmission period levels: high, medium, and low. Assuming the clustering results show that the time intervals are concentrated around 5ms, 20ms, and 40ms, respectively, the transmission period levels can be classified into three categories: high frequency (5ms), medium frequency (20ms), and low frequency (40ms).
[0026] Adaptive packet size classification is based on statistical characteristics. The mean and standard deviation of the entire packet sequence are calculated. For example, if the mean is 800 bytes and the standard deviation is 300 bytes, a threshold is set as the product of the mean and standard deviation: 800 × 300 = 240,000. Each packet is evaluated. If its size exceeds the threshold and belongs to the same transmission cycle class (e.g., low frequency), it is classified as a candidate key frame packet. This method can initially screen a group of possible key frame packets, for example, 25 candidate key frame packets were selected from 1,000 packets.
[0027] Calculate the time interval sequence between all candidate keyframe packets, for example, the interval sequence [1200ms, 1205ms, 1195ms, 1210ms...]. Calculate the mean and standard deviation of this sequence, for example, if the mean is 1200ms and the standard deviation is 10ms. Define the coefficient of variation as the ratio of the standard deviation to the mean; in this example, it is 10 / 1200 = 0.0083. Set a threshold of 0.1. When the coefficient of variation is less than this threshold, the time intervals are considered consistent. Eliminate candidate keyframe packets that do not meet temporal consistency. For example, if the three anomalous packets from the previously screened 25 candidate packets are eliminated, 22 packets are ultimately identified as keyframe packets, leaving 978 as non-keyframe packets.
[0028] Extracting periodic features from keyframe packets involves calculating the occurrence period, coefficient of variation, and mean size. The occurrence period is the mean time interval between adjacent keyframe packets, such as 1200ms. The coefficient of variation is the ratio of the standard deviation of the time interval to the mean, such as 0.0083. The mean size of a keyframe packet is the average size of all keyframe packets, such as 1500 bytes.
[0029] Non-keyframe packet characteristics include duration and packet density. Duration refers to the time required for non-keyframe packet transmission between two adjacent keyframe packets. It can be calculated as the time difference between the last non-keyframe and the first non-keyframe, such as 950ms. Packet density refers to the number of non-keyframe packets per unit time, such as an average of 4.5 packets per 100ms.
[0030] The keyframe period (1200ms), coefficient of variation (0.0083), mean keyframe size (1500 bytes), non-keyframe duration (950ms), and packet density (4.5) are combined to form the initial feature vector [1200, 0.0083, 1500, 950, 4.5]. Multiple initial time series feature vectors are calculated using a sliding time window. The window size is set to 10 seconds, with a sliding step of 1 second. Each window position is moved across the entire packet sequence, and an initial feature vector is calculated for each window position. Assume that 20 initial feature vectors are calculated. The mean vector of these 20 initial feature vectors is calculated, such as [1210, 0.0085, 1520, 960, 4.6]. For each initial feature vector, the Euclidean distance from the mean vector is calculated as the degree of deviation. For example, if the deviation of the first initial feature vector is 15.2, its weight coefficient is set to 1 / 15.2 = 0.0658. All initial eigenvectors are weighted averaged according to the weight coefficient to obtain the final time series eigenvector [1215, 0.0084, 1510, 955, 4.55].
[0031] This method uses cluster analysis of transmission time intervals to determine transmission cycle levels, combines this with adaptive classification of packet size to accurately identify keyframes, and utilizes a weighted fusion mechanism for multiple feature vectors within a sliding time window to effectively improve the robustness of packet timing features. This method not only accurately captures the periodic characteristics of video data transmission but also adapts to changes in packet transmission timing caused by network environment fluctuations, providing reliable foundational features for subsequent feature analysis.
[0032] In an optional embodiment, the step of establishing an energy characteristic chain characterizing the working state of the device according to the mutation point, the power climbing curve of the gradual change interval, and the time series evolution law of the periodic fluctuation interval includes: Performing signal smoothing processing on the energy change data, segmenting the smoothed signal to obtain a sudden change point, a gradual change interval, and a periodic fluctuation interval; Extracting energy characteristics of the mutation point, the energy characteristics including mutation amplitude and mutation duration, and calculating the ratio of the mutation amplitude to the mutation duration to obtain a mutation response characteristic; Extracting the power ramp-up characteristics of the gradual change interval, calculating a first time constant and a first amplitude coefficient representing initial power supply characteristics, and calculating a second time constant and a second amplitude coefficient representing activation characteristics of the image acquisition module; Obtaining a power change sequence in the periodic fluctuation interval, calculating the amplitude distribution characteristics of the power change sequence, segmenting the power change sequence based on the amplitude distribution characteristics, extracting the time series change characteristics of each segment, performing time-frequency analysis on the power change sequence to obtain a time-frequency characteristic graph, and extracting the main frequency component and phase evolution law of the periodic fluctuation based on the time-frequency characteristic graph; calculating the energy concentration of the main frequency component and the stability index of the phase evolution law; establishing a corresponding relationship between the environmental load and the power response based on the time series change characteristics, the energy concentration and the stability index, and calculating the basic power value and the power response coefficient; An energy characteristic chain is established based on the sudden change response characteristic, the first time constant, the first amplitude coefficient, the second time constant, the second amplitude coefficient, the basic power value and the power response coefficient. The energy characteristic chain characterizes the energy change process of the device from the initial state to the stable working state.
[0033] Exemplarily, the collected energy change data is subjected to signal smoothing. A sliding average filter is used to process the raw energy data, and the filter window size is set to 10 data points. For example, for power data with a sampling rate of 100 Hz, sliding average filtering can effectively eliminate high-frequency noise and retain the main trend of energy change. After smoothing, the slope change detection algorithm is used to segment the signal, identifying points where the power change exceeds a threshold (such as 20 W / s) as mutation points, intervals where the power slowly rises or falls as gradual change intervals, and intervals where the power exhibits periodic fluctuations as periodic fluctuation intervals. For example, during the camera startup process, three mutation points, two gradual change intervals, and one periodic fluctuation interval may be detected.
[0034] For each detected power surge, we extract its energy characteristics, including the surge amplitude and duration. The surge amplitude is calculated as the difference in power before and after the surge. For example, if the power surge changes from 50W to 120W, the surge amplitude is 70W. The surge duration is the time it takes for the power to stabilize, such as 0.5 seconds. By calculating the ratio of the surge amplitude to the surge duration, we obtain the surge response characteristic, which reflects the device's response speed to power changes. For example, a device with a surge response characteristic of 140W / s indicates a fast power response.
[0035] To analyze the power ramp characteristics during a gradual change, the power data for this period must be segmented. During the device startup phase, two distinct gradual changes can be observed: the initial power-on phase and the image acquisition module activation phase. For the initial power-on phase, power data from 0-2 seconds after power-on is selected and smoothed using a sliding window with a window size of 50ms. The smoothed power data is arranged in a time series, and the power ramp curve is analyzed using an exponential fitting method. Specifically, the time required for the power to rise from 10% to 90% is used as the first time constant, typically between 0.5 and 1.2 seconds. The ratio of the final stable power value to the theoretical maximum power value is used as the first amplitude coefficient, typically between 0.6 and 0.8. For example, the first time constant for a certain pinhole camera model during the initial power-on phase is 0.8 seconds, and the first amplitude coefficient is 0.75. For the image acquisition module activation phase, power data from 2-5 seconds is analyzed. First, the power data is denoised using a median filter to eliminate abnormal fluctuations, with a filter window size of 100ms. The rising characteristics of the power curve are then analyzed. The time from power startup to stabilization is used as the second time constant, which is typically between 1.0 and 1.5 seconds. The ratio of the stable power value to the expected power value is used as the second amplitude coefficient, which is typically between 0.6 and 0.7. For example, during the activation phase of the image acquisition module, the second time constant is 1.2 seconds, and the second amplitude coefficient is 0.65.
[0036] In the analysis of periodic fluctuations, a power variation sequence is obtained when the device is operating in a stable state. Statistical analysis is performed on the power data, calculating statistical features such as the mean, standard deviation, skewness, and kurtosis over a 10-second period. For example, the mean of power fluctuations during stable operation for a particular device is 150W, the standard deviation is 5W, the skewness is 0.2, and the kurtosis is 2.8. Based on these statistical features, a density-based clustering algorithm is used to segment the power sequence, categorizing the power levels into three levels: high, medium, and low. Time series features are extracted from the power data for each level, including the duration of the rising edge, the duration of the steady state, and the duration of the falling edge. For example, in the high-load segment, the rising edge duration is 0.3 seconds, the steady state duration is 2.5 seconds, and the falling edge duration is 0.4 seconds. Time-frequency analysis of the power variation sequence is performed using a short-time Fourier transform (SFT) with a window length of 1 second and an overlap ratio of 50%. The resulting time-frequency feature map reflects the frequency characteristics of the power variation, allowing the identification of the main frequency components. Pinhole cameras typically exhibit characteristic frequencies, such as a 2Hz primary frequency component corresponding to the image processing loop and a 5Hz secondary frequency component corresponding to the autofocus function. The energy fraction of the primary frequency component is calculated as an indicator of energy concentration. For example, 0.85 indicates that 85% of the energy is concentrated in the primary frequency component. The variance of the phase sequence is calculated as an indicator of stability. A small variance (such as 0.12) indicates stable phase changes.
[0037] Based on the above analysis results, a mapping relationship between environmental load and power response is established. The base power value is defined as the power consumption of the device when it is running at no load, such as 120W. The power response factor represents the power increase caused by an increase in unit load, such as 0.25W / unit load.
[0038] Finally, the mutation response characteristics (140W / s), time constants (0.8 seconds and 1.2 seconds), amplitude coefficients (0.75 and 0.65), basic power value (120W) and power response coefficient (0.25W / unit load) are combined in a chronological order to form an energy feature chain. That is, these features are constructed into a feature vector according to the chronological order of the mutation characteristics, initial power supply characteristics, image acquisition module activation characteristics and stable working characteristics of the equipment startup phase. The complete working status of the equipment is reflected through the temporal relationship between the features.
[0039] Existing WiFi device detection technologies primarily rely on static information such as MAC addresses, RSSI signal strength, and packet header features for identification, or employ simple power threshold detection methods. These technologies are susceptible to interference in complex electromagnetic environments and have difficulty distinguishing between standard WiFi devices and pinhole cameras. This application, rather than relying on static device characteristics, establishes a dynamic feature model by analyzing the complete energy evolution process from device startup to stable operation. A dual-time constant model is introduced to characterize the initial power supply characteristics and the activation characteristics of the image acquisition module, respectively. Time-frequency analysis is then used to extract the dominant frequency component and phase evolution patterns within the periodic fluctuation interval. This method captures the unique energy variation patterns of pinhole cameras, particularly the energy characteristics of the image sensor and video processing module, enabling accurate identification of pinhole cameras at the energy level. The improvements presented in this invention significantly enhance the accuracy and robustness of pinhole camera detection. Through energy signature chain analysis, the system can effectively distinguish between standard WiFi devices and pinhole cameras, even when the devices employ camouflage techniques. This method significantly enhances its resistance to environmental interference and maintains stable performance in complex electromagnetic environments.
[0040] In an optional embodiment, the step of performing feature fusion on the data packet timing feature and the energy feature chain to generate a feature matching coefficient of the device includes: Preprocess the data packet timing features to obtain normalized timing features; Calculating the time dimension coupling coefficient, frequency dimension coupling coefficient, and phase dimension coupling coefficient of the normalized time series feature and the energy feature chain, wherein the time dimension coupling coefficient represents the consistency of the feature change rate, the frequency dimension coupling coefficient represents the correlation of the spectrum density, and the phase dimension coupling coefficient represents the stability of the phase difference; constructing a feature correlation matrix based on the time dimension coupling coefficient, the frequency dimension coupling coefficient, and the phase dimension coupling coefficient; Constructing a dual-stream feature extraction network to extract features from the normalized time series features and the energy feature chain, and generating a time series feature mapping matrix and an energy feature mapping matrix; Cross-calculating the temporal feature mapping matrix and the energy feature mapping matrix and weighting them with the feature correlation matrix to obtain a cross-attention weight; enhancing the temporal feature mapping matrix and the energy feature mapping matrix according to the cross-attention weight, and aligning and fusing the enhanced features to obtain a fused feature; A dynamic confidence is calculated according to a gradient change of the fused feature, and a feature matching coefficient is generated based on a similarity between the fused feature and a reference feature and the dynamic confidence.
[0041] Exemplarily, when preprocessing the timing features of data packets, the minimum-maximum normalization method is used to map the original timing features to the interval [0, 1]. Specifically, for each element in the original timing features, the minimum value in the feature sequence is subtracted, and then divided by the difference between the maximum and minimum values of the feature sequence. For example, the original timing features are [120, 85, 260, 180, 310], with a minimum value of 85 and a maximum value of 310, then the normalized timing features are [0.156, 0, 0.778, 0.422, 1]. This normalization process can eliminate the influence of different feature dimensions, making subsequent feature fusion more accurate.
[0042] The coupling coefficient between the normalized time series features and the energy feature chain is calculated along three dimensions: time, frequency, and phase. The time-dimensional coupling coefficient characterizes the consistency of feature change rates and is obtained by calculating the similarity of the change trends of the two features within a time window. For example, if the time window size is set to 100ms, the difference series between the normalized time series features and the energy feature chain within the window is calculated. If the dot product of the two difference series is positive and large, it indicates that the change trends are consistent, and the coupling coefficient is close to 1. If the dot product is negative and large in absolute value, it indicates that the change trends are opposite, and the coupling coefficient is close to -1. The frequency-dimensional coupling coefficient characterizes the correlation of spectral density and is obtained by calculating the correlation of the spectral density distributions after performing a frequency domain transformation on the two features. For example, a fast Fourier transform is applied to the normalized time series features and the energy feature chain to obtain the frequency components, and then the mutual correlation coefficient between the two is calculated. This coefficient ranges from -1 to 1, with larger values indicating higher spectral correlation. The phase-dimensional coupling coefficient characterizes the stability of the phase difference and is obtained by calculating the variance of the phase difference between the two features at different times. For example, the phase information of the two types of features is extracted and the standard deviation of the phase difference sequence is calculated. If the standard deviation is small (such as less than 0.2), it means that the phase difference is stable and the coupling coefficient is close to 1; if the standard deviation is large (such as greater than 0.8), it means that the phase difference is unstable and the coupling coefficient is close to 0.
[0043] Based on the coupling coefficients of the three dimensions above, a feature correlation matrix is constructed. This matrix is a three-dimensional tensor with dimensions of time length × number of features × 3, where 3 represents the coupling coefficients of the three dimensions. For example, for a time series and energy feature chain with a length of 1000, each with 10 features, the constructed feature correlation matrix has dimensions of 1000 × 10 × 3.
[0044] The two-stream feature extraction network consists of two parallel feature extraction branches, one for processing temporal features and the other for processing energy features. Each branch consists of multiple convolutional and pooling layers to extract multi-scale features. For example, the temporal feature branch contains three convolutional layers with kernel sizes of 3×3, 5×5, and 7×7, respectively, each followed by a max pooling layer. The energy feature branch also contains three convolutional layers with the same kernel size, but with non-shared weights. These two branches generate the temporal feature map matrix and the energy feature map matrix, respectively. Both matrices have the same dimensions, such as 128×64, where 128 represents the number of feature channels and 64 represents the sequence length.
[0045] The time feature map matrix and the energy feature map matrix are cross-calculated and weighted with the feature correlation matrix to obtain the cross-attention weights. Cross-calculation involves calculating the similarity between each channel of one matrix and each channel of the other matrix to obtain an attention score between the channels. For example, the cosine similarity between the i-th channel of the time feature map matrix and the j-th channel of the energy feature map matrix is calculated to be 0.85, indicating a high correlation between the two channels. These attention scores are weighted averaged with the coupling coefficients of the corresponding positions in the feature correlation matrix to generate the final cross-attention weights.
[0046] The matrix is enhanced based on the cross-attention weights. The attention weights are multiplied by the original feature map matrix to enhance the representation of important feature channels and suppress the influence of unimportant channels. For example, for a channel with a weight of 0.9, 90% of the original eigenvalues are retained; for a channel with a weight of 0.2, only 20% of the original eigenvalues are retained. The enhanced features are aligned and fused, and the two types of features are combined into a fused feature using channel concatenation. For example, the dimensions of the enhanced time series feature map matrix and energy feature map matrix are both 128×64, and the fused feature dimension is 256×64.
[0047] Dynamic confidence is calculated based on the gradient changes of the fused features. Gradient changes refer to the distribution of the rate of change of the fused features over time. When the gradient changes smoothly without mutation points, the feature extraction and fusion process is stable, and the confidence level is high. When gradient mutation points exist, it indicates possible anomalies or interference, and the confidence level decreases. For example, if the standard deviation of the gradient sequence of the fused features over time is less than the preset threshold of 0.1, the dynamic confidence level is 0.95; if the standard deviation is greater than 0.5, the dynamic confidence level decreases to 0.6.
[0048] A reference feature library is constructed, containing feature templates for known pinhole camera devices in different operating states. Specifically, data from multiple known pinhole cameras in typical scenarios is collected and subjected to the same feature extraction and fusion process to generate standard feature templates that are stored in the reference feature library. For example, the feature library contains feature templates for 100 different pinhole camera models in five typical operating states. The cosine similarity between the current fused feature and all reference feature templates in the feature library is calculated, and the value with the highest similarity is selected as the initial matching score, such as 0.88. This initial matching score is multiplied by the dynamic confidence level (e.g., 0.95) to obtain the final feature matching coefficient, such as 0.88 × 0.95 = 0.836. The closer the feature matching coefficient is to 1, the higher the feature matching degree between the current device and the known pinhole cameras in the feature library.
[0049] This paper employs a dual-stream feature extraction network and a cross-attention mechanism to achieve a deep fusion of packet timing features and energy feature chains. By calculating the coupling coefficients across the time, frequency, and phase dimensions and constructing a feature correlation matrix, this approach ensures the preservation of important information and the suppression of noise during the feature fusion process. This method fully exploits the correlation between the two types of features, improving the accuracy and reliability of feature matching.
[0050] In an optional embodiment, a deep neural network is used to perform time series analysis on the feature matching coefficients, and the step of generating a device identification coefficient by an output layer of the deep neural network includes: Generate time series statistical features based on the time series sequence of feature matching coefficients; calculate a feature stability score based on the time series statistical features, construct a multi-scale memory pool to store feature patterns of different time scales, and determine the update weight of the memory pool according to the feature stability score, where the update weight is inversely proportional to the feature stability score; Extracting time series features from the time series sequence through a deep neural network to obtain a time series feature vector, calculating the similarity between the time series feature vector and the feature pattern stored in the multi-scale memory pool, and selecting the optimal memory feature based on the similarity; Calculating the Mahalanobis distance between the time series feature vector and its mean to obtain an anomaly score, mapping the anomaly score to obtain a suppression weight, multiplying the suppression weight by the time series feature vector to obtain a suppressed feature vector; performing attention calculation on the suppressed feature vector to obtain a device working mode feature, calculating a steady-state weight based on a gradient value of the time series feature vector; and multiplying the steady-state weight by the device working mode feature to obtain an enhanced device feature; Constructing an evaluation index based on historical recognition results, calculating an update amount based on the gradient of the evaluation index, and adaptively adjusting the decision threshold; The optimal memory feature, the enhanced device feature, and the steady-state weight are subjected to feature splicing and nonlinear transformation, a device recognition score is calculated through a multi-layer perceptron, and a device recognition coefficient is generated according to the decision threshold and the device recognition score.
[0051] Figure 2 This is the architecture diagram of a deep neural network-based WiFi pinhole camera time series analysis system. For example, a time series of feature matching coefficients is collected, and the mean, standard deviation, and skewness of this series are calculated to generate time series statistical features. Taking a WiFi pinhole camera detection as an example, a 30-second time series of feature matching coefficients is collected, with data points such as [0.72, 0.75, 0.73, 0.71, 0.74, 0.76, 0.75, 0.77, 0.76, 0.78]. The calculated mean is 0.75, the standard deviation is 0.021, and the skewness is 0.15. Based on these time series statistical features, a feature stability score is calculated. The specific method is to calculate the coefficient of variation (i.e., the standard deviation divided by the mean), which is 0.028. A smaller coefficient of variation indicates higher feature stability. Next, a multi-scale memory pool is constructed to store feature patterns at different time scales. This memory pool comprises short-term memory units, medium-term memory units, and long-term memory units. The update weight of the memory pool is determined based on the feature stability score, and the update weight is inversely proportional to the feature stability score. In actual implementation, when the feature stability score is 0.028, the corresponding short-term memory unit update weight is 0.9, the medium-term memory unit update weight is 0.6, and the long-term memory unit update weight is 0.3.
[0052] The time series is fed into a deep neural network for feature extraction. This deep neural network consists of convolutional layers and recurrent layers. The convolutional layers, consisting of 16 3×1 convolution kernels, extract local time series features. The recurrent layers, consisting of 64 LSTM units, capture long-term dependencies. After network processing, a 128-dimensional time series feature vector is obtained.
[0053] The cosine similarity method is used to calculate the similarity between the time series feature vector and the feature patterns stored in the multi-scale memory pool. For each feature pattern in each memory unit, the similarity with the current time series feature vector is calculated. For example, the similarity with the feature pattern in the short-term memory unit is 0.92, the similarity with the feature pattern in the medium-term memory unit is 0.85, and the similarity with the feature pattern in the long-term memory unit is 0.78. The optimal memory feature is selected based on the similarity. In this example, the feature pattern in the short-term memory unit is selected as the optimal memory feature.
[0054] Subsequently, the Mahalanobis distance between the time series feature vector and its mean is calculated to obtain an anomaly score. In the specific implementation, the mean vector of the feature vector is first calculated, and then the Mahalanobis distance between the current feature vector and the mean vector is calculated, resulting in an anomaly score of 0.35. The anomaly score is converted into a suppression weight using a mapping function. The mapping method is to subtract the anomaly score from 1, resulting in a suppression weight of 0.65. The suppression weight is multiplied by the time series feature vector to obtain a suppressed feature vector. This step aims to reduce the influence of abnormal features and enhance the expression of stable features. The specific implementation of attention calculation on the suppressed feature vector includes three steps: first, the autocorrelation score of each element in the feature vector is calculated. By performing a dot product operation on the feature vector with itself, the correlation matrix between the elements is obtained. Second, the correlation matrix is normalized to obtain an attention weight matrix. Finally, the attention weight matrix is used to weight the feature vector. For WiFi pinhole cameras, higher weights are preset for the packet periodicity indicator (0.4), the transmission volume mutation indicator (0.3), and the duration indicator (0.3). The weighted result forms a feature vector that represents the device's operating mode. Calculate the gradient of the time series feature vector along the time dimension. A smaller gradient indicates a more stable device operating state. Use a mapping function to convert the gradient value into a steady-state weight. In this example, the resulting steady-state weight is 0.82. Multiply the steady-state weight by the device operating mode feature to obtain the enhanced device feature. This step aims to enhance the feature representation of the device in a stable operating state and improve recognition accuracy.
[0055] Record the results of the past 100 recognition attempts and calculate precision and recall. The current precision is 0.95 and the recall is 0.92. Updates are calculated based on the gradient of the evaluation metrics. When precision decreases, the decision threshold is increased; when recall decreases, the decision threshold is decreased. This adaptively adjusts the decision threshold to 0.78.
[0056] The optimal memory features, enhanced device features, and steady-state weights are concatenated and nonlinearly transformed. Feature concatenation involves concatenating the three feature vectors along a dimension to form a longer feature vector. The nonlinear transformation uses the ReLU activation function to enhance the expressive power of the features. The device recognition score is calculated using a multilayer perceptron (MLP), which contains three hidden layers with 256, 128, and 64 nodes per layer, respectively. The concatenated feature vector is input and an identification score between 0 and 1 is output. The current identification score is 0.86. The device recognition coefficient is generated based on the decision threshold of 0.78 and the device recognition score of 0.86. When the recognition score is greater than the decision threshold, the device recognition coefficient is 1, indicating that the target device has been detected; otherwise, it is 0.
[0057] Traditional device identification methods mainly rely on static features or simple statistical features, such as mean value, standard deviation, etc., and lack in-depth analysis of the temporal variation law of features. The present invention proposes a deep temporal analysis method based on multi-scale memory pool and adaptive feature enhancement. The core innovation lies in the construction of a three-layer memory structure of short-term, medium-term and long-term, which realizes the multi-scale storage and dynamic update of feature patterns; introduces an abnormality suppression mechanism based on Mahalanobis distance and a steady-state enhancement mechanism based on gradient value, which significantly improves the robustness of feature extraction; designs an adaptive decision threshold adjustment strategy, so that the system can automatically optimize the judgment criteria according to the historical recognition effect. This method not only focuses on the features themselves, but also pays more attention to the temporal evolution law of features. Through the analysis of the stability, abnormality and consistency of features, it realizes the accurate identification of the device working mode.
[0058] In an optional embodiment, the steps of constructing a multi-scale memory pool and dynamically storing and updating characteristic patterns include: The time scale is divided into three levels: short-term, medium-term, and long-term. The sampling period of each level is dynamically adjusted based on the feature change rate, and the time series statistical features of each level are collected; the time series statistical features are extracted to obtain feature patterns, the time domain consistency and frequency domain correlation of the feature patterns are calculated, and the feature stability score is generated based on the time domain consistency and the frequency domain correlation; Calculating the information entropy of the feature stability score, determining the capacity of each level of memory pool according to the information entropy, expanding the memory pool capacity when the information entropy increases, and compressing the memory pool capacity when the information entropy decreases; Calculating an update weight based on the feature stability score, where the update weight is inversely proportional to the feature stability score; calculating the similarity between the feature pattern and the patterns stored in the memory pool to obtain a difference score; calculating the similarity between the feature pattern and the current working state feature to obtain a representative score; and using the difference score and the representative score as adjustment factors for the update weight; The access time and usage frequency of the characteristic pattern are recorded, a timeliness score is calculated based on the access time and usage frequency, and the memory pool storage space is released when the timeliness score is lower than a preset score threshold.
[0059] For example, the time scale is divided into three levels: short-term, medium-term, and long-term. The short-term level corresponds to data within the last 5 minutes, the medium-term level corresponds to data from 5 minutes to 2 hours, and the long-term level corresponds to data from 2 hours to 24 hours. The feature change rate is calculated by calculating the difference between feature vectors in adjacent time windows. When the feature change rate is high, the sampling frequency is increased; when the feature change rate is low, the sampling frequency is reduced. For example, when detecting WiFi pinhole cameras, if the feature change rate in the short term is 0.15 (rapid change), the sampling period of the short-term level is set to 10 seconds; if the feature change rate in the medium term is 0.08 (moderate change), the sampling period of the medium term level is set to 30 seconds; if the feature change rate in the long term is 0.03 (slow change), the sampling period of the long-term level is set to 2 minutes.
[0060] At each sampling point, time series statistical features, including mean, standard deviation, skewness, and kurtosis, are collected. Taking the energy characteristics of a WiFi pinhole camera as an example, the time series statistical features at a certain sampling point are: mean 0.72, standard deviation 0.08, skewness 0.25, and kurtosis 2.86. Feature extraction is performed on these time series statistical features to obtain characteristic patterns. Feature extraction uses dimensionality reduction techniques to convert high-dimensional time series statistical features into low-dimensional characteristic patterns. Temporal consistency is calculated by calculating the average similarity between the current characteristic pattern and the characteristic patterns in the past N time windows. Frequency correlation is calculated by calculating the autocorrelation function of the characteristic pattern in the frequency domain. For the pinhole camera, the temporal consistency at a certain moment is 0.85 (high, indicating stability in the temporal domain), and the frequency correlation is 0.78 (high, indicating stability in the frequency domain). The weighted average of temporal consistency and frequency correlation, with weights of 0.6 and 0.4, respectively, yields a feature stability score of 0.82.
[0061] The feature stability score is discretized, evenly dividing the interval [0, 1] into 10 subintervals. The probability distribution of the feature stability score falling into each subinterval is then calculated. Information entropy is calculated based on this probability distribution. When the feature stability score distribution is uniform, the information entropy is high; when the distribution is concentrated, the information entropy is low. For example, if the distribution of feature stability scores over a certain period is [0.05, 0.08, 0.12, 0.15, 0.18, 0.20, 0.12, 0.06, 0.03, 0.01], the corresponding information entropy is 2.15. When the information entropy increases, it indicates that the feature is highly volatile, so the memory pool capacity should be expanded to store more feature patterns. When the information entropy decreases, it indicates that the feature is relatively stable, so the memory pool capacity should be compressed to conserve storage resources. In the specific implementation, the basic capacity of the short-term memory pool is 100 feature patterns, the basic capacity of the medium-term memory pool is 200 feature patterns, and the basic capacity of the long-term memory pool is 500 feature patterns. When the information entropy is 2.15 (higher), the capacity of the memory pool at each level is expanded to 120, 240 and 600 feature patterns respectively.
[0062] The update weight is calculated based on the feature stability score. The update weight is inversely proportional to the feature stability score and is obtained by subtracting the feature stability score from 1 and then normalizing it. For example, when the feature stability score is 0.82, the initial update weight is 0.18. At the same time, the similarity between the feature pattern and the stored pattern in the memory pool is calculated to obtain a difference score. The higher the difference score, the more new information the current feature pattern contains. In a specific implementation, for a certain feature pattern, the similarity with the most similar pattern in the memory pool is 0.65, then the difference score is 0.35. The similarity between the feature pattern and the current working state feature is calculated to obtain a representative score. The higher the representative score, the more representative the feature pattern is to the current working state. For example, if the similarity between a certain feature pattern and the current working state feature is 0.88, then the representative score is 0.88. The initial update weight is multiplied by the difference score and representativeness score, with weights of 0.3 and 0.7 respectively, and the final update weight is 0.18×(0.3×0.35+0.7×0.88)=0.124.
[0063] Record the access time and usage frequency of the characteristic pattern. The access time refers to the timestamp of the characteristic pattern's most recent access, and the usage frequency refers to the number of times the characteristic pattern has been accessed over the past period of time. A timeliness score is calculated based on the access time and usage frequency. The timeliness score calculation takes into account the time decay factor. Over time, the timeliness score of the characteristic pattern gradually decreases; the higher the usage frequency, the higher the timeliness score. For example, if a characteristic pattern was last accessed 2 hours ago and is used 10 times per day, the calculated timeliness score is 0.45. When the timeliness score falls below the preset score threshold (such as 0.3), the memory pool storage space is released and the characteristic pattern is deleted.
[0064] This invention effectively addresses the time-scale adaptability issue in device identification by constructing a multi-scale memory pool and implementing dynamic storage and updating of feature patterns. Multi-level time partitioning enables simultaneous capture of short-term, medium-term, and long-term feature changes, and dynamically adjusted sampling periods enhance adaptability to varying rates of change. A memory management mechanism based on feature stability, differentiation, representativeness, and timeliness enables efficient utilization of memory resources, improving recognition performance and adaptability in complex and changing environments, and providing reliable technical support for device identification based on temporal features.
[0065] In an optional embodiment, a feature verification network is constructed, wherein the feature verification network evaluates the continuous stability of the data packet timing feature and the channel interference level of the target area respectively, and the step of generating the feature credibility score includes: Inputting the data packet timing features into a multi-scale feature extraction layer, the multi-scale feature extraction layer including convolution kernels with different receptive fields, extracting features from the data packet timing features to obtain multi-scale features, calculating the coefficient of variation and autocorrelation coefficient of the multi-scale features, constructing a stability matrix, and calculating a weighted stability score based on the stability matrix; Establishing a WiFi channel interference feature library, the WiFi channel interference feature library including a channel overlap feature template, a multipath interference feature template, and a concurrent transmission feature template; extracting channel state information of a target area; calculating a first matching degree between the channel state information and the channel overlap feature template, a second matching degree between the channel state information and the multipath interference feature template, and a third matching degree between the channel state information and the concurrent transmission feature template; and calculating a channel interference degree score based on the first matching degree, the second matching degree, and the third matching degree; Constructing a dual-branch feature verification network, respectively extracting a timing feature vector of the timing feature of the data packet and an energy feature vector of the channel state information, calculating the mutual information between the timing feature vector and the energy feature vector, constructing an attention matrix, and calculating a feature consistency score based on the mutual information and the attention matrix; The weighted stability score, the channel interference degree score and the feature consistency score are weightedly fused to obtain a feature credibility score.
[0066] For example, when constructing a feature verification network, a multi-scale feature extraction layer is first designed. This layer contains multiple convolution kernels with different receptive fields, such as 3×3, 5×5, and 7×7 kernels. After the packet's temporal features are input into this layer, each convolution kernel performs a convolution operation on the input features, generating feature maps at different scales. Specifically, 16 3×3 convolution kernels, 8 5×5 convolution kernels, and 4 7×7 convolution kernels can be designed, with the Reinforced Lu (ReLU) function used as the activation function. For the extracted multi-scale features, the coefficient of variation (coefficient of variation) is calculated: the standard deviation divided by the mean. For example, for the temporal features of a particular packet, the extracted feature vector [0.32, 0.35, 0.31, 0.33, 0.34] has a mean of 0.33, a standard deviation of 0.015, and a coefficient of variation of 0.045. The autocorrelation coefficient is also calculated to reflect the correlation between features at different time points. Based on the calculated results, a stability matrix is constructed, with each element representing the stability score of a specific feature at a specific scale. For the above example, a 3×5 stability matrix can be constructed, where rows represent different convolution kernel sizes and columns represent different feature dimensions. After obtaining the stability matrix, weights are assigned based on the importance of features at different scales. For example, features extracted by 3×3, 5×5, and 7×7 convolution kernels are assigned weights of 0.5, 0.3, and 0.2, respectively, to calculate the weighted stability score.
[0067] The WiFi channel interference signature library contains three types of feature templates: channel overlap, multipath interference, and concurrent transmission. The channel overlap template contains a spectrum overlap feature vector, such as [0.8, 0.6, 0.4, 0.2, 0.1], indicating the degree of overlap in different frequency bands. The multipath interference template contains delay and energy attenuation features, such as a delay feature vector of [2.5, 3.1, 4.2, 5.0] nanoseconds and an energy attenuation feature vector of [0.7, 0.5, 0.3, 0.2]. The concurrent transmission feature template contains a collision probability feature vector, such as [0.1, 0.3, 0.5, 0.7, 0.9], indicating the collision probability under different loads. Channel state information of the target area is acquired using a spectrum analyzer, and its spectrum, delay, and collision features are extracted. The degree of match between the extracted channel state information and each template in the signature library is calculated. For example, if the spectral feature vector of the target area is [0.75, 0.55, 0.45, 0.25, 0.15], the cosine similarity between it and the channel overlap feature template is 0.98, resulting in a first matching degree. Similarly, the matching degrees of the delay and collision features with the corresponding templates are calculated, yielding a second matching degree of 0.85 and a third matching degree of 0.76. The channel interference score is calculated based on these three matching degrees using a weighted average method, assigning weights of 0.4, 0.3, and 0.3 to the three matching degrees, respectively, resulting in a comprehensive channel interference score of 0.872.
[0068] When constructing a two-branch feature verification network, one branch processes packet timing features, while the other processes channel state information. In the timing feature branch, a one-dimensional convolutional network is used to extract the timing feature vector. The network consists of three convolutional layers: the first layer has 16 convolution kernels, the second layer has 32 convolution kernels, and the third layer has 64 convolution kernels, ultimately generating a 128-dimensional timing feature vector. In the channel state information branch, an energy detector is used to extract energy features, including spectral energy distribution and time-domain energy variation, resulting in a 128-dimensional energy feature vector. The mutual information between the two feature vectors is calculated to reflect the correlation between the two features. The mutual information value is obtained by calculating the joint probability distribution and marginal probability distribution of the feature vectors. For example, a calculated mutual information value of 0.65 indicates a strong correlation between the two features. An attention matrix is constructed to emphasize important features. The attention matrix is 128×128 in size, with each element representing the attention weight for the corresponding dimension of the timing feature vector and the energy feature vector. A feature consistency score is calculated based on the mutual information and the attention matrix. The score ranges from 0 to 1, with higher values indicating higher consistency.
[0069] Finally, the weighted stability score, channel interference score, and feature consistency score are weighted and fused to generate a feature credibility score. The weights of these three scores are set to 0.35, 0.35, and 0.3, respectively. For the above example, assuming a weighted stability score of 0.92, a channel interference score of 0.872, and a feature consistency score of 0.78, the feature credibility score is 0.92 × 0.35 + 0.872 × 0.35 + 0.78 × 0.3 = 0.863. This score is used to assess feature reliability and provide a confidence reference for subsequent perception tasks.
[0070] Existing technologies typically use a single statistical metric (such as analysis of variance) to assess feature stability, or simply use the signal-to-noise ratio to assess channel quality. This approach is unable to address complex interference factors in WiFi environments, such as channel overlap, multipath effects, and concurrent transmission. This paper proposes a multi-dimensional feature verification network and designs a multi-scale feature extraction layer. This layer uses convolution kernels with different receptive fields to capture feature stability at different time scales, forming a stability matrix for comprehensive evaluation. A WiFi channel interference feature library containing three templates, namely channel overlap, multipath interference, and concurrent transmission, is established to enable refined identification of interference factors in complex wireless environments. Furthermore, a dual-branch feature verification network is constructed to evaluate feature consistency by calculating the mutual information between temporal and energy feature vectors. This multi-dimensional, multi-layered feature verification mechanism comprehensively assesses feature reliability and effectively distinguishes between inherent device feature fluctuations and feature changes caused by environmental interference. This paper significantly improves the reliability of WiFi pinhole camera detection, maintaining high accuracy in various interference environments and effectively reducing false positive and false negative rates. It performs particularly well in areas with dense WiFi devices and complex architectural environments.
[0071] In an optional embodiment, when both the device recognition coefficient and the feature credibility score meet preset conditions, the step of confirming that the pinhole camera is detected includes: When the device recognition coefficient is greater than the adaptive threshold calculated based on the historical detection accuracy and the feature credibility score meets the minimum credibility requirement, it is confirmed that a pinhole camera has been detected, and a threat level score is calculated comprehensively based on the device recognition coefficient and the feature credibility score.
[0072] For example, the device's recognition coefficient is determined to be greater than an adaptive threshold calculated based on historical detection accuracy. The adaptive threshold calculation takes into account the accuracy of historical detection results and dynamically adjusts the judgment criteria based on the accuracy metric. In a specific implementation, the accuracy of the most recent 100 tests, including the number of true positives, false positives, true negatives, and false negatives, is recorded to calculate the detection accuracy.
[0073] For example, in the last 100 detections, 45 true pinhole cameras were successfully identified (true positives), 5 were false positives, 48 non-pinhole camera devices were correctly excluded (true negatives), and 2 were missed (false negatives). Based on this data, the accuracy, precision, and recall were calculated to be 0.9, 0.9, and 0.96. Based on these metrics, an adaptive threshold is calculated. When the accuracy is high, the threshold is appropriately lowered to increase detection sensitivity; when the accuracy is low, the threshold is increased to reduce false positives. In this example, the adaptive threshold is calculated to be 0.75.
[0074] In the current detection, the device identification coefficient is 0.83, which is greater than the adaptive threshold of 0.75, satisfying the first judgment condition. Next, the feature credibility score is checked to see if it meets the minimum credibility requirement. The minimum credibility requirement is preset based on the security level of the application scenario. High-security scenarios have higher minimum credibility requirements, while general scenarios have lower minimum credibility requirements. In this example, the application scenario is a hotel room, and the preset minimum credibility requirement is 0.7.
[0075] The current detected feature confidence score is 0.78, which exceeds the minimum confidence requirement of 0.7 and meets the second judgment condition. Therefore, the pinhole camera is confirmed to have been detected, and the threat level assessment phase begins.
[0076] The threat level score is calculated using a weighted average of the device identification coefficient and the signature credibility score. The device identification coefficient is weighted at 0.6, and the signature credibility score is weighted at 0.4. In this example, the threat level score is calculated as 0.83 × 0.6 + 0.78 × 0.4 = 0.81.
[0077] The threat level is divided into five levels: very low risk (0-0.2), low risk (0.2-0.4), medium risk (0.4-0.6), high risk (0.6-0.8), and extremely high risk (0.8-1.0). The current threat level score is 0.81, which is extremely high risk and will issue the highest level alert and provide an estimated device location.
[0078] In actual applications, the current detection results are also recorded, including the device identification coefficient, feature credibility score, threat level score, and detection time. This information will be used for subsequent adaptive threshold updates and performance evaluations.
[0079] This invention effectively addresses the accuracy and reliability issues in pinhole camera detection by introducing dual criteria and an adaptive threshold mechanism. The combined determination of the device identification coefficient and the feature credibility score ensures both a precise match of the target device's features and the credibility of the detection results.
[0080] According to a second aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0081] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0082] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A WiFi pinhole camera detection method based on wireless signal feature analysis is characterized by: include: Collecting wireless signals within a target area, and obtaining data packet sequences and energy change data of the wireless signals; Based on the data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to the transmission period and size distribution. The occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted to establish the data packet timing characteristics that characterize the video data transmission characteristics; Extracting the mutation point, gradual change interval, and periodic fluctuation interval from the energy change data, and establishing an energy characteristic chain characterizing the working state of the equipment based on the mutation point, the power climbing curve of the gradual change interval, and the time series evolution law of the periodic fluctuation interval; Performing feature fusion on the data packet timing features and the energy feature chain to generate a feature matching coefficient for the device; performing time series analysis on the feature matching coefficient using a deep neural network, and generating a device identification coefficient at the output layer of the deep neural network; Constructing a feature verification network, wherein the feature verification network evaluates the sustained stability of the data packet timing feature and the degree of channel interference in the target area to generate a feature credibility score; When both the device recognition coefficient and the feature credibility score meet preset conditions, it is confirmed that a pinhole camera has been detected.
2. The method according to claim 1, characterized in that Based on the data characteristics of video coding, the data packet sequence is divided into key frame data packets and non-key frame data packets according to the transmission period and size distribution, the occurrence period of the key frame data packets, the duration of the non-key frame data packets, and the size ratio of the key frame data packets to the non-key frame data packets are extracted, and the steps of establishing the data packet timing characteristics that characterize the video data transmission characteristics include: Extracting transmission time intervals between adjacent data packets in the data packet sequence, and determining a transmission period level of the data packet according to a clustering result of the transmission time intervals; Adaptively classifying the data packets in the data packet sequence according to the relationship between the data packet size and the mean and standard deviation of the data packet size, and classifying the data packets that are larger than the product of the mean and standard deviation of the data packet size and belong to the same transmission cycle level as candidate key frame data packets; Calculating a time interval sequence between the candidate key frame data packets, verifying the candidate key frame data packets according to the consistency of the time interval sequence, determining the candidate key frame data packets that pass the verification as key frame data packets, and determining the remaining data packets as non-key frame data packets; Extracting periodic features of the key frame data packets, the periodic features including the occurrence period of the key frame data packets, the coefficient of variation of the occurrence period, and the size mean of the key frame data packets; counting the duration and data packet density of non-key frame data packets located between adjacent key frame data packets, the data packet density being the number of non-key frame data packets within a unit time window; An initial timing feature vector is constructed based on the periodic characteristics, the duration of the non-key frame data packets and the data packet density; multiple initial timing feature vectors are calculated within a sliding time window, and a mean vector of the multiple initial timing feature vectors is calculated, and the initial timing feature vectors are weighted according to the degree of deviation between the initial timing feature vector and the mean vector to generate a final timing feature vector, wherein the weight coefficient is inversely proportional to the degree of deviation.
3. The method according to claim 1, characterized in that The steps of establishing an energy characteristic chain representing the working state of the device according to the mutation point, the power climbing curve in the gradual change interval, and the time series evolution law of the periodic fluctuation interval include: Extracting energy characteristics of the mutation point, the energy characteristics including mutation amplitude and mutation duration, and calculating the ratio of the mutation amplitude to the mutation duration to obtain a mutation response characteristic; Extracting the power ramp-up characteristics of the gradual change interval, calculating a first time constant and a first amplitude coefficient representing initial power supply characteristics, and calculating a second time constant and a second amplitude coefficient representing activation characteristics of the image acquisition module; Obtaining a power change sequence in the periodic fluctuation interval, calculating the amplitude distribution characteristics of the power change sequence, segmenting the power change sequence based on the amplitude distribution characteristics, extracting the time series change characteristics of each segment, performing time-frequency analysis on the power change sequence to obtain a time-frequency characteristic graph, and extracting the main frequency component and phase evolution law of the periodic fluctuation based on the time-frequency characteristic graph; calculating the energy concentration of the main frequency component and the stability index of the phase evolution law; establishing a corresponding relationship between the environmental load and the power response based on the time series change characteristics, the energy concentration and the stability index, and calculating the basic power value and the power response coefficient; An energy characteristic chain is established based on the sudden change response characteristic, the first time constant, the first amplitude coefficient, the second time constant, the second amplitude coefficient, the basic power value and the power response coefficient. The energy characteristic chain characterizes the energy change process of the device from the initial state to the stable working state.
4. The method according to claim 1, wherein The step of fusing the data packet timing feature and the energy feature chain to generate a feature matching coefficient for the device includes: Preprocess the data packet timing features to obtain normalized timing features; Calculating the time dimension coupling coefficient, frequency dimension coupling coefficient, and phase dimension coupling coefficient of the normalized time series feature and the energy feature chain, wherein the time dimension coupling coefficient represents the consistency of the feature change rate, the frequency dimension coupling coefficient represents the correlation of the spectrum density, and the phase dimension coupling coefficient represents the stability of the phase difference; constructing a feature correlation matrix based on the time dimension coupling coefficient, the frequency dimension coupling coefficient, and the phase dimension coupling coefficient; Constructing a dual-stream feature extraction network to extract features from the normalized time series features and the energy feature chain, and generating a time series feature mapping matrix and an energy feature mapping matrix; Cross-calculating the temporal feature mapping matrix and the energy feature mapping matrix and weighting them with the feature correlation matrix to obtain a cross-attention weight; enhancing the temporal feature mapping matrix and the energy feature mapping matrix according to the cross-attention weight, and aligning and fusing the enhanced features to obtain a fused feature; A dynamic confidence is calculated according to a gradient change of the fused feature, and a feature matching coefficient is generated based on a similarity between the fused feature and a reference feature and the dynamic confidence.
5. The method according to claim 1, wherein The step of using a deep neural network to perform time series analysis on the feature matching coefficients, wherein the output layer of the deep neural network generates a device identification coefficient, includes: Generate time series statistical features based on the time series sequence of feature matching coefficients; calculate a feature stability score based on the time series statistical features, construct a multi-scale memory pool to store feature patterns of different time scales, and determine the update weight of the memory pool according to the feature stability score, where the update weight is inversely proportional to the feature stability score; Extracting time series features from the time series sequence through a deep neural network to obtain a time series feature vector, calculating the similarity between the time series feature vector and the feature pattern stored in the multi-scale memory pool, and selecting the optimal memory feature based on the similarity; Calculating the Mahalanobis distance between the time series feature vector and its mean to obtain an anomaly score, mapping the anomaly score to obtain a suppression weight, multiplying the suppression weight by the time series feature vector to obtain a suppressed feature vector; performing attention calculation on the suppressed feature vector to obtain a device working mode feature, calculating a steady-state weight based on a gradient value of the time series feature vector; and multiplying the steady-state weight by the device working mode feature to obtain an enhanced device feature; Constructing an evaluation index based on historical recognition results, calculating an update amount based on the gradient of the evaluation index, and adaptively adjusting the decision threshold; The optimal memory feature, the enhanced device feature, and the steady-state weight are subjected to feature splicing and nonlinear transformation, a device recognition score is calculated through a multi-layer perceptron, and a device recognition coefficient is generated according to the decision threshold and the device recognition score.
6. The method according to claim 5, characterized in that The steps of constructing a multi-scale memory pool and dynamically storing and updating feature patterns include: The time scale is divided into three levels: short-term, medium-term, and long-term. The sampling period of each level is dynamically adjusted based on the feature change rate, and the time series statistical features of each level are collected; the time series statistical features are extracted to obtain feature patterns, the time domain consistency and frequency domain correlation of the feature patterns are calculated, and the feature stability score is generated based on the time domain consistency and the frequency domain correlation; Calculating the information entropy of the feature stability score, determining the capacity of each level of memory pool according to the information entropy, expanding the memory pool capacity when the information entropy increases, and compressing the memory pool capacity when the information entropy decreases; Calculating an update weight based on the feature stability score, where the update weight is inversely proportional to the feature stability score; calculating the similarity between the feature pattern and the patterns stored in the memory pool to obtain a difference score; calculating the similarity between the feature pattern and the current working state feature to obtain a representative score; and using the difference score and the representative score as adjustment factors for the update weight; The access time and usage frequency of the characteristic pattern are recorded, a timeliness score is calculated based on the access time and usage frequency, and the memory pool storage space is released when the timeliness score is lower than a preset score threshold.
7. The method according to claim 1, characterized in that Constructing a feature verification network, wherein the feature verification network evaluates the sustained stability of the data packet timing feature and the channel interference level of the target area, and the steps of generating a feature credibility score include: Inputting the data packet timing features into a multi-scale feature extraction layer, the multi-scale feature extraction layer including convolution kernels with different receptive fields, extracting features from the data packet timing features to obtain multi-scale features, calculating the coefficient of variation and autocorrelation coefficient of the multi-scale features, constructing a stability matrix, and calculating a weighted stability score based on the stability matrix; Establishing a WiFi channel interference feature library, the WiFi channel interference feature library including a channel overlap feature template, a multipath interference feature template, and a concurrent transmission feature template; extracting channel state information of a target area; calculating a first matching degree between the channel state information and the channel overlap feature template, a second matching degree between the channel state information and the multipath interference feature template, and a third matching degree between the channel state information and the concurrent transmission feature template; and calculating a channel interference degree score based on the first matching degree, the second matching degree, and the third matching degree; Constructing a dual-branch feature verification network, respectively extracting a timing feature vector of the timing feature of the data packet and an energy feature vector of the channel state information, calculating the mutual information between the timing feature vector and the energy feature vector, constructing an attention matrix, and calculating a feature consistency score based on the mutual information and the attention matrix; The weighted stability score, the channel interference degree score and the feature consistency score are weightedly fused to obtain a feature credibility score.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Apparatus for monitoring sneak shot device, and control method thereof
CN109410537A
Photographing device detection method
CN112235819A
Hidden camera discovering device and method for hardware device detection based on Internet
CN116017392A
Camera detection method based on wireless signal
CN119676429A
3D imaging device with digital micromirror device and operating method thereof
KR1020220037938A
Cited By
Wifi-based hidden camera judgment method and device, equipment and storage medium
CN122053820A