Screen voyeurism monitoring system based on multi-dimensional feature analysis

Through multi-dimensional feature analysis and deep twin network, the problem of insufficient low disturbance detection of existing screen candid shooting systems is solved, and efficient monitoring and interference resistance of screen candid shooting behavior is achieved, which improves detection accuracy and stability.

CN120318775BActive Publication Date: 2025-08-29BEIJING DATANGSHENGXING TECH DEV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510803771.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-29
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing screen candid shooting system cannot fully capture candid shooting behavior based on single-dimensional feature analysis, especially for low-perturbable candid shooting detection capabilities, and lacks the anti-episode mechanism for lens distortion and flash interference.

Method used

Multi-dimensional feature analysis is adopted, including spatial feature modules, temporal feature modules and space-time fusion modules, combined with deep twin networks, and through spatial phase encoding, temporal phase encoding and space-time phase fusion, phase fusion factors and dynamic anti-perturbation fields are introduced to construct dynamic adversarial modulation to realize monitoring of screen sneak shot behavior.

Benefits of technology

It improves the detection ability of low-disturbance candid shooting behavior, reduces the dependence on data annotation, enhances the ability to combat interference, and improves the monitoring accuracy and stability of screen candid shooting behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318775B_ABST
    Figure CN120318775B_ABST
Patent Text Reader

Abstract

The present invention discloses a screen candid photography behavior monitoring system based on multi-dimensional feature analysis, including: a spatial feature module: collecting screen image data, performing spatial phase encoding on the screen image data, and obtaining spatial image phase features; a temporal feature module: performing temporal phase encoding on n consecutive frames of screen image data, and obtaining temporal image phase features; a spatiotemporal fusion module: performing spatiotemporal phase fusion after optimizing the spatial image phase features and the temporal image phase features, and obtaining the best spatiotemporal encoding image features; a behavior monitoring module: constructing a deep twin network, and obtaining behavior monitoring results from the best spatiotemporal encoding image features based on the deep twin network; making the normal scene features more stable and the candid scene features more significant, and the acquired feature data further amplifies the feature differences between the two types of samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of screen image recognition technology, and in particular to a screen candid photography behavior monitoring system based on multi-dimensional feature analysis. Background Art

[0002] Screens, as the core carrier of information display, are widely used in sensitive scenarios such as office, finance, and healthcare. However, the security protection of screen content faces severe challenges. Secret photography (illegally photographing the screen using mobile phones, cameras, and other devices) may lead to the leakage of confidential information, posing a huge risk to personal privacy and corporate security. Current mainstream screen security technologies mainly include physical protection (such as anti-peep film and camera shielding) and software monitoring (such as anomaly detection based on image recognition).

[0003] Currently, existing screen-based candid photography systems, which are usually based on single-dimensional feature analysis (focusing only on image texture), are unable to fully capture the spatiotemporal feature distortion caused by candid photography. In particular, they have insufficient detection capabilities for low-disturbance candid photography (such as long-distance, low-resolution photography). Furthermore, they lack a countermeasure mechanism against the optical characteristics of candid photography devices, making it difficult to cope with actual interference factors such as lens distortion and flash interference. Therefore, a screen-based candid photography monitoring system based on multi-dimensional feature analysis is proposed here. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned objectives, the present invention proposes the following technical solutions:

[0005] The screen-camera behavior monitoring system based on multi-dimensional feature analysis includes:

[0006] Spatial feature module: collects screen image data, performs spatial phase encoding on the screen image data, and obtains spatial image phase features;

[0007] Time feature module: performs time phase encoding on the screen image data of n consecutive frames to obtain the time image phase feature;

[0008] Spatiotemporal fusion module: Optimizes the spatial image phase features and the temporal image phase features through spatiotemporal phase fusion to obtain the best spatiotemporal coding image features;

[0009] The optimized space-time phase fusion adds a phase fusion factor to the traditional space-time phase fusion, and based on the phase fusion factor, constructs a dynamic anti-disturbance field to realize dynamic countermeasure modulation of space-time phase fusion;

[0010] Behavior monitoring module: Build a deep twin network and obtain behavior monitoring results from the optimal spatiotemporal encoding image features based on the deep twin network;

[0011] Among them, the deep twin network consists of two sub-networks with the same structure and a spatiotemporal correlation block. The spatiotemporal correlation block adjusts the convolution kernel weights of the deep twin network by calculating the spatiotemporal correlation between the output features of the two sub-networks at different layers.

[0012] The process of acquiring the time image phase feature is as follows:

[0013] Collect screen image data, obtain the color information and brightness information of each pixel in the screen image, and use the spatial phase encoding function to obtain the spatial image phase characteristics based on the color information and brightness information of the pixel points .

[0014] The spatial image phase feature acquisition process is as follows:

[0015] Collect n consecutive frames of screen images, and for the same pixel in each frame, obtain the brightness change and color channel change between different frames. Based on the brightness change and color channel change, use the time phase encoding function to obtain the time image phase feature. .

[0016] The phase fusion factor acquisition process is as follows:

[0017] Obtain the fusion weight parameters in traditional spatiotemporal phase fusion;

[0018] Acquiring spatial phase gradient amplitude based on spatial image phase characteristics;

[0019] A Sigmoid function is introduced, which takes the spatial phase gradient amplitude and the temporal image phase features as input and outputs the phase fusion factor.

[0020] The dynamic anti-disturbance field construction process is as follows:

[0021] According to the phase fusion factor, the spatial phase image features and the temporal phase image features are basically fused, and a threat level coefficient is preset. , the threat level coefficient is calculated by the spatial phase gradient amplitude Get, that is:

[0022] When , it indicates low risk;

[0023] When , it means medium danger;

[0024] When , it indicates high risk;

[0025] Introducing PerlinNoise noise function, based on the threat level function Combined with the PerlinNoise noise function to obtain a dynamic anti-disturbance field.

[0026] The overall implementation process of the optimized spatiotemporal phase fusion is as follows:

[0027] Based on the phase features of spatial images and temporal images, the basic fusion of spatiotemporal phase is performed. The basic fusion part is expressed as:

[0028]

[0029] in, is the initial space-time phase characteristic, is the fusion weight parameter, is the spatial image phase feature, the temporal image phase feature, is the time image phase feature;

[0030] The final process of obtaining the optimized spatiotemporal phase characteristics is:

[0031]

[0032] in, To optimize the spatiotemporal phase characteristics, is the phase fusion factor, is the dynamic anti-disturbance field.

[0033] The sub-network includes an input layer, two convolutional layers, and a fully connected layer;

[0034] The input layer receives the optimized spatiotemporal phase features As input data;

[0035] The two convolutional layers use convolution kernels of different sizes, 3×3 and 5×5;

[0036] The pooling layer adopts a combination of maximum pooling and average pooling. The pooling window size is p×q. , directly perform maximum pooling and average pooling operations;

[0037] The fully connected layer integrates the features extracted by the convolution layer and the pooling layer, and maps the scattered features after the maximum pooling operation and the average pooling operation into a unified feature space, which is represented as f.

[0038] The process of adjusting the convolution kernel weights of the deep Siamese network by calculating the spatiotemporal correlation between the output features of the two sub-networks at different layers is as follows:

[0039] Suppose two subnetworks with the same structure are and , where the subnetwork and Both are convolutional neural networks. For sub-networks The output and subnetwork of layer j The output of the jth layer, obtains the spatiotemporal correlation coefficient between the output features :

[0040] The image data collected when there is no hidden camera and the image data collected when there is hidden camera behavior form a sample pair. When the sample pair is normal. , when the sample is abnormal: force ;

[0041] Constructing association loss based on sample-to-state : , where m is the interval parameter;

[0042] According to the associated loss, the convolution kernel weights are adjusted through the back-propagation algorithm;

[0043] Obtaining Contrastive Loss for Deep Siamese Networks Using Contrastive Learning Training Strategies , combined with the association loss , update and adjust the convolution kernel weights through the Adam optimizer: ,in, is the learning rate, is the convolution kernel weight adjusted by the spatiotemporal correlation block.

[0044] The process of obtaining the behavior monitoring results is as follows:

[0045] Generate optimized spatiotemporal phase characteristics for real-time input screen image data , input the deep twin network dual branch, and get the output feature vector , , get the output feature vector , The Euclidean distance d = , preset an abnormal threshold ,when When the user is watching a video, it is judged that there is a voyeuristic act.

[0046] The present invention has the following beneficial effects:

[0047] In this invention, first, spatial phase encoding is used to map the position, color, and brightness of pixels to phase values. Nonlinear transformation is used to highlight texture details and layout features. Temporal phase encoding is used to weight the sum of brightness and color changes between frames, focusing on the dynamic features of recent frames. Spatiotemporal fusion and dynamic adversarial modulation mechanisms are introduced. The optimal spatiotemporal phase features are obtained through the phase fusion factor (dynamically adjusting weights based on spatial gradient and temporal phase) and the PerlinNoise noise function interference field, making the features of normal scenes more stable and the features of candid scenes more prominent. The acquired feature data further amplifies the feature differences between the two subsequent types of samples.

[0048] Secondly, two identical sub-networks extract features from the same input. Through the dual constraints of Euclidean distance and spatiotemporal correlation coefficient, the features of normal samples are forced to be highly consistent, while the features of candid samples are highly different. This design does not require a large amount of labeled data and can learn abnormal patterns solely through the relative differences between sample pairs, reducing data dependence. The convolution kernel weights are adjusted through backpropagation, allowing the network to adaptively enhance its sensitivity to spatiotemporal consistency (normal scenes) or differences (candid scenes). For example, in static document scenes, the network automatically increases its focus on spatial texture stability; in dynamic video scenes, it focuses on capturing abnormal fluctuations in temporal phase.

[0049] Finally, the threat level is divided according to the spatial phase gradient amplitude, and the PerlinNoise noise intensity and frequency are dynamically adjusted. In high-threat scenarios (such as detecting a flash or macro photography), the noise frequency is positively correlated with the temporal phase, actively interfering with the imaging of the hidden camera device, while providing the twin network with more significant detection features (such as the increase in the feature vector distance caused by noise). BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a system block diagram of the screen camera behavior monitoring system based on multi-dimensional feature analysis proposed by the present invention. DETAILED DESCRIPTION

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0052] Example: Figure 1 As shown, the screen candid photography behavior monitoring system proposed by the present invention based on multi-dimensional feature analysis includes:

[0053] Spatial feature module: collects screen image data, performs spatial phase encoding on the screen image data, and obtains spatial image phase features;

[0054] Collect screen image data and obtain color and brightness information of each pixel (x, y) in the screen image. Each pixel (x, y) includes features describing the screen content in the spatial dimension.

[0055] Color information is represented by the values ​​of the three RGB channels (r, g, b). Different color combinations can produce different visual effects. Brightness information is calculated using the formula H = 0.299r + 0.587g + 0.114b, reflecting the brightness of the pixel.

[0056] Based on the color information and brightness information of the pixel points, the spatial phase encoding function is used to convert the screen image data into spatial image phase features. The formula is expressed as:

[0057] ;

[0058] in, is the spatial image phase feature, By using the inverse tangent function, the relationship between the RGB color channel values ​​is nonlinearly transformed. Reflects the difference between the red channel and the green channel value, divided by This is to perform normalization to avoid the situation where the denominator is zero due to a small b value. is a very small constant. The inverse tangent function maps this difference to a specific range, allowing different color combinations to be represented in a relatively uniform and distinguishable way.

[0059] For example, for reddish pixels, rg is a positive value, and after this calculation, a specific phase value will be obtained, while greenish pixels will get a different result, thus distinguishing pixels with different color tendencies in the phase space;

[0060] The item is used to reflect the position information of the pixel in the screen image. x and y are the horizontal and vertical coordinates of the pixel respectively. and H are the width and height of the screen image. By dividing the coordinate value by the size of the image, the pixel position is normalized to the interval [0,1], so that pixels at the same relative position on screens of different sizes can have similar representations in phase encoding;

[0061] For example, the calculated result of the pixel in the upper left corner of the screen is close to 0 on any screen size, which helps to capture the spatial layout characteristics of the image in subsequent analysis;

[0062] The term compresses and transforms the brightness value through a logarithmic function. Low brightness values ​​will get smaller phase values, and high brightness values ​​will get larger phase values. As the brightness increases, the growth rate of the phase value gradually slows down. This can represent the influence of brightness in phase encoding while avoiding excessive dominance of high brightness values ​​on the overall encoding result.

[0063] For example, for very bright pixels, their brightness values ​​are relatively high. After logarithmic transformation, their contribution to phase encoding will not be too prominent, and they can be balanced with other factors to determine the spatial phase value of the pixel.

[0064] Time feature module: performs time phase encoding on the screen image data of n consecutive frames to obtain the time image phase feature;

[0065] Collect n consecutive screen images and obtain the brightness change between different frames for the same pixel (x, y) in each frame. and color channel changes ;

[0066] Based on the changes in brightness and color channels, the temporal phase encoding function is used to convert the screen image data into temporal image phase features. The formula is expressed as:

[0067] ;

[0068] in, is the temporal image phase feature, , , are the average values ​​of the color channel changes, which are used to unify the change values ​​of different features into a suitable range, i is the index, and T is the time constant;

[0069] The time weight coefficient is set according to the time sequence. The closer to the current time, the greater the weight. In the dynamic change process of screen content, recent changes tend to better reflect the current actual situation and potential anomalies.

[0070] For example, during the playback of a video, if a voyeuristic act suddenly occurs, the frame change information near the moment of the voyeurism is more important for detecting the anomaly. By giving greater weight to the pixels corresponding to the recent time, temporal phase encoding can more sensitively capture these key change information.

[0071] Spatiotemporal fusion module: Optimizes the spatial image phase features and the temporal image phase features through spatiotemporal phase fusion to obtain optimized spatiotemporal phase features;

[0072] Based on the phase features of spatial images and temporal images, the basic fusion of spatiotemporal phase is performed. The basic fusion part is expressed as:

[0073] ;

[0074] in, is the initial space-time phase characteristic, To optimize the spatiotemporal phase characteristics, is the fusion weight parameter, The actual situation of the screen image is used to reasonably allocate the weights of the spatial phase and the temporal phase, so that the final spatiotemporal phase encoding result can more comprehensively and accurately reflect the spatiotemporal characteristics of the screen content;

[0075] In traditional spatiotemporal phase fusion, the fusion weight parameter λ is often fixed, or only adjusted through simple experiments, which is difficult to adapt to the actual situation of complex and changing screen content. Replace the weight parameter λ, the phase fusion factor The acquisition process is:

[0076] Introducing a spatial phase gradient amplitude Capturing the sudden change of texture in screen image, the spatial phase gradient amplitude is calculated by the spatial image phase feature. Obtain the partial derivative and calculate the modulus to capture the screen texture mutation and spatial phase gradient amplitude The formula is:

[0077] ;

[0078] in, The partial derivative of the spatial phase image feature along the horizontal direction (x-axis) of the image reflects the rate of change of the spatial phase in the horizontal direction. For example, in a screen image, how the spatial phase of pixels changes from left to right can reflect the severity of the change in the horizontal direction.

[0079] It represents the partial derivative of the spatial phase image feature along the vertical direction (y-axis) of the image, reflecting the rate of change of the spatial phase in the vertical direction, that is, the degree of change of the spatial phase of the pixel from top to bottom;

[0080] Specifically, the spatial phase gradient amplitude Used to perceive changes in the phase characteristics of spatial images. When new elements appear on the screen or when existing elements undergo significant changes, the value will change accordingly.

[0081] A Sigmoid function is introduced to convert the spatial phase gradient amplitude Multiply by a preset parameter (between 0 and 1), the time image phase feature is used as the input and output phase fusion factor of the Sigmoid function, which is expressed as: ;

[0082] The process of obtaining the dynamic anti-disturbance field is as follows:

[0083] Phase fusion factor based on dynamic adjustment , building a dynamic anti-disturbance field , according to the phase fusion factor Spatial phase image features and temporal phase image features Conduct basic integration;

[0084] Specifically, when When the value is large, it means that the spatial characteristics of the current screen content change significantly, and the spatial phase has a larger weight in the fusion result. When the value of is small, the time phase has a greater weight in the fusion result. In this way, the proportion of spatial phase and time phase in the basic fusion can be reasonably determined according to the actual spatiotemporal changes of the screen content, so that the fusion result can more accurately reflect the spatiotemporal characteristics of the screen content.

[0085] Preset a threat level coefficient , the threat level coefficient is calculated by the spatial phase gradient amplitude Get, that is:

[0086] When , the texture is stable and there is no obvious distortion, indicating low risk;

[0087] When the texture changes slightly, it indicates medium danger;

[0088] When the texture changes dramatically, such as local phase distortion caused by flash illumination and loss of details in macro photography, it is a high risk;

[0089] Combined with the PerlinNoise noise function, it is introduced into the construction of the dynamic anti-disturbance field, and its frequency and time phase image characteristics are Positive correlation, as the time phase value changes, the frequency of PerlinNoise will also change accordingly;

[0090] Specifically, this is because the temporal phase value reflects the dynamic changes of the screen content in the time dimension. When the screen content changes rapidly, the temporal phase value changes greatly. At this time, increasing the frequency of PerlinNoise can enhance the interference effect on the imaging of the camera. When the screen content changes slowly, the frequency of PerlinNoise also decreases accordingly, but it can still maintain a certain interference effect. This dynamic change can interfere with the imaging of the camera, making it difficult for the hidden camera to obtain clear screen image content.

[0091] Through the basic fusion part, the threat level function and The dynamic anti-disturbance field is obtained by combining the Perlin noise function, which is expressed as:

[0092]

[0093] Among them, through the dynamic anti-disturbance field The dynamic anti-disturbance field integrates the information of spatial phase, temporal phase and dynamic noise to form a dynamic and adversarial modulation effect, which not only increases the difficulty of sneak photography, but also provides more obvious feature differences for subsequent deep twin network detection;

[0094] The final process of obtaining the optimized spatiotemporal phase characteristics is:

[0095]

[0096] Specifically, when there is voyeurism, the dynamic anti-disturbance field will change the spatiotemporal characteristics of the screen image, making it easier for the subsequent deep twin network to capture these changes and further accurately determine whether there is voyeurism.

[0097] Build a deep twin network and obtain behavior monitoring results from the optimal spatiotemporal encoding image features based on the deep twin network;

[0098] The deep Siamese network consists of two identical sub-networks and composition;

[0099] Specifically, the use of two identical sub-networks is to achieve fair feature extraction and comparison in contrastive learning, ensuring that the two sub-networks can process input data in the same way;

[0100] Each sub-network in the deep twin network includes an input layer, a convolutional layer, and a fully connected layer;

[0101] The input layer receives optimized spatiotemporal phase features ;

[0102] Convolutional layer setup process:

[0103] Subnetwork and It includes two convolutional layers, using convolution kernels of different sizes, 3×3 and 5×5;

[0104] The 3×3 convolution kernel has a strong local feature extraction capability and is used to capture subtle features such as the stroke details of characters in screen images and the texture details of images. The larger convolution kernel, 5×5, is used to obtain more extensive contextual information, such as the overall layout of the screen and the features of large-area patterns.

[0105] Pooling layer setting process:

[0106] Subnetwork and Both methods use a combination of maximum pooling and average pooling. Maximum pooling can highlight the maximum value in a feature and retain the most significant feature information. For example, it can highlight key features such as important logos and bright spots in a screen image. These key features may play a decisive role in judging voyeuristic photography. Average pooling can smooth features, consider the overall information within the area, and effectively process background information, avoiding ignoring changes in the overall background by focusing only on local significant features.

[0107] Assume that the pooling window size is p×q, for the input feature data , directly perform maximum pooling and average pooling operations;

[0108] Fully connected layer setting process:

[0109] The fully connected layer integrates the features extracted by the convolutional layer and the pooling layer. The neurons in the fully connected layer are connected to all the neurons in the pooling layer, mapping the scattered features after the maximum pooling operation and the average pooling operation into a unified feature space, denoted as f;

[0110] In the subnet and A spatiotemporal correlation block is introduced between them;

[0111] The spatiotemporal correlation block adjusts the network parameters of the deep Siamese network by calculating the spatiotemporal correlation between the output features of the two sub-networks at different layers;

[0112] For subnetworks The output and subnetwork of layer j The output of the jth layer, obtains the spatiotemporal correlation coefficient between the output features ;

[0113] Collect image data when there is no covert photography and image data when covert photography occurs to form a sample pair;

[0114] Specifically, normal sample pairs (screen image pairs without hidden camera), such as image frame pairs generated by continuous normal screen operation, or normal screen image pairs collected by different devices, can enable the deep twin network to learn the spatiotemporal characteristics of normal screen content, such as the smooth changes between frames during normal video playback, the stable distribution of pixel phases when static documents are displayed, etc., and the spatiotemporal correlation coefficient is forced during training. ;

[0115] Abnormal sample pairs (including secretly photographed screen image pairs) can be image pairs collected by simulating secretly photographed images, or image pairs obtained by injecting interference into normal images (such as simulating optical distortion of secretly photographed images). Their role is to let the network learn the temporal and spatial feature distortion rules caused by secretly photographed images. , allowing the network to grasp the height difference of features under abnormal conditions;

[0116] When the sample is normal (no hidden camera): mandatory (features are highly consistent);

[0117] When the sample is abnormal (there is a hidden camera): forced (characteristic height difference);

[0118] Constructing association loss based on sample-to-state : ;

[0119] Among them, m is the interval parameter (m=0.5), which is used to ensure that the correlation coefficients of the two types of samples have sufficient discrimination;

[0120] According to the associated loss, the network parameters (convolution kernel weight W) are adjusted through the back-propagation algorithm;

[0121] Based on the sample pairs consisting of image data without candid photography and image data with candid photography, the contrast loss of the deep twin network is directly obtained using the contrast learning strategy. , combined with the association loss , update the parameters through the Adam optimizer: ,in, is the learning rate, is the convolution kernel weight adjusted by the spatiotemporal correlation block;

[0122] Specifically, through the time-space correlation block, when there is no hidden camera on the screen, and Highly consistent spatiotemporal features will be extracted ( ), if the actual If it is too small, the associated blocks will increase the similarity of the two-branch features (adjusting the network parameters through the back-propagation algorithm so that the convolution kernel increases the focus on the common spatiotemporal patterns);

[0123] When there is a hidden camera, the optical interference of the hidden camera device will destroy the spatiotemporal phase encoding. and The difference of extracted features increases ( ), if the actual If it is too large, the associated blocks will strengthen the feature differences (adjust the network parameters through the back propagation algorithm to reduce the convolution kernel and increase the judgment accuracy);

[0124] For the real-time input screen image data, generate optimized spatiotemporal phase features after preprocessing , input the deep twin network dual branch, and get the output feature vector , , , are the output features of the two sub-networks;

[0125] Get the output feature vector , The Euclidean distance d = ;

[0126] Preset an abnormal threshold ,when When the number of times the camera is on, it is judged that there is a candid photo shooting behavior, otherwise it is normal;

[0127] Specifically, through the collaborative judgment of the spatiotemporal correlation block of the deep twin network and the preset threshold, when there is no voyeurism: the Euclidean distance is small (features are close) and the spatiotemporal correlation coefficient is large (linear relationship is strong);

[0128] When there is voyeurism, the Euclidean distance is large (feature dispersion) and the spatiotemporal correlation coefficient is small (linear relationship is weak);

[0129] and The two branches are convolutional neural networks with exactly the same structure and weights (both contain multiple layers of convolution, pooling, and fully connected layers). The design principle is to ensure that for normal, uninterrupted input, the two branches can extract highly similar features (due to the identical network structure and parameters, and consistent feature extraction logic). However, for input interfered by candid photography, the features extracted by the two branches will destroy the spatiotemporal phase encoding due to the interference, resulting in clearly distinguishable differences.

[0130] In the application, several formulas involved are calculated by taking their numerical values ​​after removing the dimensions, and the formulas are established by collecting a large amount of data and performing software simulation to obtain a formula for the most recent real situation. Some coefficients or weights in the formulas are set by technical personnel in this field according to actual conditions, so they will not be elaborated here.

[0131] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0132] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A screen-camera behavior monitoring system based on multi-dimensional feature analysis, characterized by: include: Spatial feature module: collects screen image data, performs spatial phase encoding on the screen image data, and obtains spatial image phase features; Time feature module: performs time phase encoding on the screen image data of n consecutive frames to obtain the time image phase feature; Spatiotemporal fusion module: Optimizes the spatial image phase features and the temporal image phase features through spatiotemporal phase fusion to obtain the best spatiotemporal coding image features; The optimized space-time phase fusion adds a phase fusion factor to the traditional space-time phase fusion, and based on the phase fusion factor, constructs a dynamic anti-disturbance field to realize dynamic countermeasure modulation of space-time phase fusion; Behavior monitoring module: Build a deep twin network and obtain behavior monitoring results from the optimal spatiotemporal encoding image features based on the deep twin network; The deep twin network consists of two sub-networks with the same structure and a spatiotemporal correlation block. The spatiotemporal correlation block adjusts the convolution kernel weights of the deep twin network by calculating the spatiotemporal correlation between the output features of different layers of the two sub-networks. The phase fusion factor acquisition process is as follows: Obtain the fusion weight parameters in traditional spatiotemporal phase fusion; Acquiring spatial phase gradient amplitude based on spatial image phase characteristics; A Sigmoid function is introduced, which takes the spatial phase gradient amplitude and the temporal image phase features as input and outputs the phase fusion factor; The dynamic anti-disturbance field construction process is as follows: According to the phase fusion factor, the spatial phase image features and the temporal phase image features are basically fused, and a threat level coefficient is preset. , the threat level coefficient is calculated by the spatial phase gradient amplitude Get, that is: When , it indicates low risk; When , it means medium danger; When , it indicates high risk; Introducing PerlinNoise noise function, based on the threat level function Combined with PerlinNoise noise function to obtain dynamic anti-disturbance field; The process of obtaining the behavior monitoring results is as follows: Generate optimized spatiotemporal phase characteristics for real-time input screen image data , input the deep twin network dual branch, and get the output feature vector , , get the output feature vector , The Euclidean distance d = , preset an abnormal threshold ,when When the user is watching a video, it is judged that there is a voyeuristic act.

2. The screen candid photography behavior monitoring system based on multi-dimensional feature analysis according to claim 1 is characterized in that: The spatial image phase feature acquisition process is as follows: Collect screen image data, obtain the color information and brightness information of each pixel in the screen image, and use the spatial phase encoding function to obtain the spatial image phase characteristics based on the color information and brightness information of the pixel points .

3. The screen candid photography behavior monitoring system based on multi-dimensional feature analysis according to claim 1 is characterized in that: The process of acquiring the time image phase feature is as follows: Collect n consecutive frames of screen images, and for the same pixel in each frame, obtain the brightness change and color channel change between different frames. Based on the brightness change and color channel change, use the time phase encoding function to obtain the time image phase feature. .

4. The screen-camera behavior monitoring system based on multi-dimensional feature analysis according to claim 1 is characterized in that: The overall implementation process of the optimized spatiotemporal phase fusion is as follows: Based on the phase features of spatial images and temporal images, the basic fusion of spatiotemporal phase is performed. The basic fusion part is expressed as: ; in, is the initial space-time phase characteristic, is the fusion weight parameter, is the spatial image phase feature, is the time image phase feature; The final process of obtaining the optimized spatiotemporal phase characteristics is: ; in, To optimize the spatiotemporal phase characteristics, is the phase fusion factor, is the dynamic anti-disturbance field.

5. The screen candid photography behavior monitoring system based on multi-dimensional feature analysis according to claim 1 is characterized in that: The sub-network includes an input layer, two convolutional layers, a pooling layer and a fully connected layer; The input layer receives the optimized spatiotemporal phase features As input data; The two convolutional layers use convolution kernels of different sizes, 3×3 and 5×5; The pooling layer adopts a combination of maximum pooling and average pooling. The pooling window size is p×q. , directly perform maximum pooling and average pooling operations; The fully connected layer integrates the features extracted by the convolution layer and the pooling layer, and maps the scattered features after the maximum pooling operation and the average pooling operation into a unified feature space, which is represented as f.

6. The screen-camera behavior monitoring system based on multi-dimensional feature analysis according to claim 5 is characterized in that: The process of adjusting the convolution kernel weights of the deep Siamese network by calculating the spatiotemporal correlation between the output features of the two sub-networks at different layers is as follows: Suppose two subnetworks with the same structure are and , where the subnetwork and Both are convolutional neural networks. For sub-networks The output and subnetwork of layer j The output of the jth layer, obtains the spatiotemporal correlation coefficient between the output features : Collect image data when there is no candid shooting and image data when there is candid shooting to form a sample pair. When the sample pair is normal, force , when the sample is abnormal: force ; Constructing association loss based on sample-to-state : , where m is the interval parameter; According to the associated loss, the convolution kernel weights are adjusted through the back-propagation algorithm; Obtaining Contrastive Loss for Deep Siamese Networks Using Contrastive Learning Training Strategies , combined with the association loss , update and adjust the convolution kernel weights through the Adam optimizer: ,in, is the learning rate, is the convolution kernel weight adjusted by the spatiotemporal correlation block.

Citation Information

Patent Citations

  • SAR image change detection method based on twin network

    CN110659591A

  • Monitoring screen area candid shooting detection method based on real-time visual identification

    CN119964252A