Monitoring Screen Area Secret Photography Detection Method Based on Real-Time Visual Recognition
Through multiple acquisition devices combining object contour and lens brightness analysis methods, the limitations of gesture action recognition in the prior art are solved, and high-precision screen candid detection and early warning are realized to ensure personal privacy and security.
Patent Information
- Application Number
- CN202510449887.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the prior art, the characteristics of the object being held cannot be fully measured through simple gesture action recognition, resulting in limitations in screen sneak shot detection.
Multiple acquisition devices are used to collect images, combining object contour analysis and lens brightness calculation, and through edge detection algorithms and convolutional neural network models, gesture movements, object contours and lens brightness are comprehensively measured, secretly shot behaviors are judged and early warnings are issued.
It improves the reliability and accuracy of the detection of candid photography, and can prevent candid photography at the first time, protect personal privacy and avoid information leakage.
Smart Images

Figure CN119964252B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of peeping detection, and more specifically, to a peeping detection method for a monitoring screen area based on real-time visual recognition. Background Art
[0002] In recent years, with the continuous development of intelligent devices, peeping detection for screens has faced new challenges.
[0003] Chinese Patent Application Publication No. CN117953574A discloses a screen anti-peeping method and system based on visual recognition, including: when a user is in front of the screen of a display, obtaining an original image by real-time shooting through a camera; after denoising and enhancing the original image, obtaining an image to be detected; performing human body detection on the image to be detected; when there is a human body area in the image to be inspected, determining whether there is a shooting behavior by gesture recognition on the human body area. It can be seen that only simply through gesture actions, the characteristics of the object being held are not comprehensively measured, which will lead to certain limitations in identifying shooting behaviors.
[0004] Therefore, it is necessary to design a peeping detection method for a monitoring screen area based on real-time visual recognition to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a peeping detection method for a monitoring screen area based on real-time visual recognition, aiming to solve the problem that simply through gesture actions, the characteristics of the object being held are not comprehensively measured, which will lead to certain limitations in identifying shooting behaviors.
[0006] The present invention proposes a peeping detection method for a monitoring screen area based on real-time visual recognition, including:
[0007] Determine the screen area, evenly deploy a plurality of acquisition devices in the screen area, and set the acquisition frequency of each acquisition device;
[0008] Judge whether there is a suspected peeping behavior according to the overlapping result of the current image and the previous image of the acquisition device. When there is the suspected peeping behavior, obtain the current images of all acquisition devices and perform preprocessing, and determine the screen image of the screen area according to the result of the preprocessing;
[0009] Identify the gesture actions of the screen image, determine the object contour of the screen image according to the edge detection algorithm, and search for the object contour value in the preset object contour table. Determine the lens pixel points according to the object contour, analyze all the lens pixel points, determine the effective lens pixel points, and obtain the lens brightness value according to the effective lens pixel points;
[0010] Based on the gesture actions, the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value are output by using the surreptitious shooting behavior model. The surreptitious shooting behavior value is determined according to the contour weight factor, the brightness weight factor, the object contour value and the lens brightness value. The surreptitious shooting behavior value is compared with the surreptitious shooting behavior threshold, and whether to send a surreptitious shooting warning message to the signal jammer is judged according to the comparison result.
[0011] Further, when judging whether there is a suspected surreptitious shooting behavior according to the overlapping result of the current image and the previous image of the acquisition device, it includes:
[0012] When all the current images of all the acquisition devices completely overlap with all the previous images, it is determined that there is no suspected surreptitious shooting behavior in the screen area;
[0013] When there is an incomplete overlap between the current image and the previous image of one acquisition device, it is determined that there is a suspected surreptitious shooting behavior in the screen area.
[0014] Further, when obtaining the current images of all the acquisition devices and performing preprocessing, and determining the screen image of the screen area according to the result of the preprocessing, it includes:
[0015] The preprocessing includes image denoising and geometric correction, feature point extraction is performed on all the current images after preprocessing, all the extracted feature points are matched, and the association data between the current images is determined;
[0016] Based on the association data, all the current images are registered and merged to obtain the to-be-processed screen image of the screen area. The to-be-processed screen image is adjusted to determine the screen image of the screen area, and the image adjustment includes adjusting the image size and removing the stitching gap.
[0017] Further, when determining the object contour of the screen image according to the edge detection algorithm and searching for the object contour value in the preset object contour table, it includes:
[0018] The Sobel operator is used to calculate the gradient of the screen image, non-maximum suppression is performed based on the Canny edge detection algorithm, and the edge detection result is optimized by the global optimal solution method and the dynamic programming algorithm to extract the object contour of the screen image;
[0019] Based on the object contour, the corresponding object contour value is determined in the preset object contour table, and the object contour table includes the mapping relationship between the object contour and the corresponding object contour value.
[0020] Further, when analyzing all the lens pixel points to determine the effective lens pixel points and obtaining the lens brightness value according to the effective lens pixel points, it includes:
[0021] Change all lens pixel points to lens pixel position points, construct a lens pixel coordinate system with all the lens pixel position points, fit all the lens pixel position points, and determine the lens fitting curve;
[0022] Eliminate the lens pixel position points that are not fitted, take the lens pixel position points on the lens fitting curve as valid lens pixel points, obtain the brightness values corresponding to all the valid lens pixel points, and take the average value of all the brightness values as the lens brightness value.
[0023] Further, when changing all the lens pixel points to lens pixel position points and constructing a lens pixel coordinate system with all the lens pixel position points, it includes:
[0024] Take the spatial position of the lens pixel point as the X-axis coordinate value of the lens pixel point, obtain the lens data of the lens pixel point, and establish a lens data sequence based on the lens data;
[0025] Obtain a qualified lens data sequence corresponding to the lens data sequence, and compare the lens data sequence with the qualified lens data sequence;
[0026] When the lens data in the lens data sequence is greater than the qualified lens data in the qualified lens data sequence, divide the lens data into the lens first data set;
[0027] When the lens data in the lens data sequence is equal to the qualified lens data in the qualified lens data sequence, divide the lens data into the lens second data set;
[0028] When the lens data in the lens data sequence is less than the qualified lens data in the qualified lens data sequence, divide the lens data into the lens third data set;
[0029] Calculate the Y-axis coordinate value of the lens pixel point according to the lens first data set, the lens second data set and the lens third data set, and construct a lens pixel coordinate system according to the X-axis coordinate value and the Y-axis coordinate value.
[0030] Further, when calculating the Y-axis coordinate value of the lens pixel point according to the lens first data set, the lens second data set and the lens third data set, it includes:
[0031] Count the lens first quantity of the lens first data set, the lens second quantity of the lens second data set and the lens third quantity of the lens third data set;
[0032] The Y-axis coordinate value of the lens pixel point is obtained from the following formula:
[0033] ;
[0034] Among them, R represents the Y-axis coordinate value of the lens pixel point, YA represents the mean value of the lens data in the first lens data set, n represents the first number of lenses, Yi represents the i-th lens data in the first lens data set, Yg represents the qualified lens data corresponding to the i-th lens data in the first lens data set, YB represents the mean value of the lens data in the third lens data set, m represents the third number of lenses, Yj represents the j-th lens data in the third lens data set, and Yh represents the qualified lens data corresponding to the j-th lens data in the third lens data set.
[0035] Furthermore, when adopting the secret photography behavior model to output the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value based on the gesture action, it includes:
[0036] Obtain the historical suspected secret photography data set, and divide the historical suspected secret photography data set into a training set and a test set;
[0037] Pre-select a convolutional neural network model, and substitute the training set into the convolutional neural network model for training;
[0038] Verify the trained convolutional neural network model according to the test set to determine the secret photography behavior model. When the verification value is greater than or equal to the verification value after the previous training, then use the currently trained convolutional neural network model as the secret photography behavior model;
[0039] When the verification value is less than the verification value after the previous training, use grid search to find the hyperparameters of the convolutional neural network model, and replace the hyperparameters of the currently trained convolutional neural network model. Cluster the training set to determine the training sample set, re-determine the training set according to the training sample set, substitute the re-determined training set into the convolutional neural network model with replaced hyperparameters, and continue training until the secret photography behavior model is determined;
[0040] Substitute the gesture action, the object contour value, and the lens brightness value into the secret photography behavior model to obtain the contour weight factor and the brightness weight factor.
[0041] Furthermore, when clustering the training set to determine the training sample set and re-determining the training set according to the training sample set, it includes:
[0042] S30: Determine the initial number of clusters K according to the training set, and randomly select K points as the initial cluster centers;
[0043] S31: Assign each data point in the training set to the initial cluster center closest to it;
[0044] S32: Calculate the mean of all data points in the cluster and use it as the new cluster center of this cluster;
[0045] S33: Repeat S31 and S32 until the cluster center no longer changes or reaches a predetermined number of iterations, and use the finally determined clusters of the clustering as the training sample set;
[0046] S34: Interpolate between every two data points in the training sample set, determine new data points and add them to the training set.
[0047] Further, when determining the secretly taken picture behavior value according to the contour weight factor, the brightness weight factor, the object contour value and the lens brightness value, comparing the secretly taken picture behavior value with the secretly taken picture behavior threshold, and judging whether to send a secretly taken picture warning message to the signal jammer according to the comparison result, it includes:
[0048] The secretly taken picture behavior value is obtained according to the following formula:
[0049] ;
[0050] where Q represents the secretly taken picture behavior value, w1 represents the contour weight factor, S1 represents the object contour value, w2 represents the brightness weight factor, and S2 represents the lens brightness value;
[0051] Preset a secretly taken picture behavior threshold Qa;
[0052] When Q ≥ Qa, it is determined to send a secretly taken picture warning message to the signal jammer;
[0053] When Q < Qa, it is determined not to send a secretly taken picture warning message to the signal jammer.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the image acquisition of multiple acquisition devices by real-time visual recognition, combined with the analysis of the object contour, the calculation of the lens brightness value and the capture of gesture actions, it can accurately identify secretly taken picture behaviors, avoid relying on single data or simple rule matching, improve the reliability of secretly taken picture behavior detection, adopt the method of evenly distributing multiple acquisition devices to cover different angles of the screen area, improve the detection ability of secretly taken picture behaviors in complex environments, avoid the problems of occlusion or viewing angle limitation of a single acquisition device, determine the contour weight factor and the brightness weight factor based on the secretly taken picture behavior model, comprehensively measure the object contour and the lens brightness, improve the accuracy of calculating the secretly taken picture behavior value, compare the secretly taken picture behavior value with the secretly taken picture behavior threshold in real time, and can send a warning message to the signal jammer at the first time when a secretly taken picture behavior occurs, thereby preventing the normal operation of the secretly taken picture device, effectively protecting personal privacy, and at the same time avoiding information leakage. Description of the Drawings
[0055] Upon reading the following detailed description of the preferred embodiments, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0056] Figure 1 It is a flowchart of a method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition provided by an embodiment of the present invention. Specific embodiments
[0057] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0058] In some embodiments of the present application, referring to Figure 1 as shown, a method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition includes:
[0059] S100: Determine the screen area, evenly deploy a plurality of acquisition devices in the screen area, and set the acquisition frequency of each acquisition device.
[0060] S200: Determine whether there is a suspected surreptitious photographing behavior according to the overlap result of the current image and the previous image of the acquisition device. When there is a suspected surreptitious photographing behavior, obtain the current images of all acquisition devices and perform preprocessing, and determine the screen image of the screen area according to the result of the preprocessing.
[0061] S300: Identify the gesture actions in the screen image, determine the object contour of the screen image according to the edge detection algorithm, and look up the object contour value in the preset object contour table. Determine the lens pixel points according to the object contour, analyze all the lens pixel points to determine the effective lens pixel points, and obtain the lens brightness value according to the effective lens pixel points.
[0062] S400: Based on the gesture actions, use the surreptitious photographing behavior model to output the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value. Determine the surreptitious photographing behavior value according to the contour weight factor, the brightness weight factor, the object contour value, and the lens brightness value. Compare the surreptitious photographing behavior value with the surreptitious photographing behavior threshold, and determine whether to send a surreptitious photographing warning message to the signal jammer according to the comparison result.
[0063] Specifically, with the left edge of the screen as the center and the screen-facing direction as the reference, 180 degrees centered is set as the screen area, ensuring that all possible secretly photographing devices are within the detection range of the screen area while avoiding detection blind spots, thus improving the detection accuracy. Multiple acquisition devices are evenly deployed in the screen area. Preferably, there are 5 acquisition devices, which capture images at the azimuths of 0°, 45°, 90°, 135°, and 180° centered respectively. The acquisition devices are high-definition cameras, and the cameras are equipped with infrared sensors to enhance the image capture ability in low-light environments. The acquisition devices capture images of the screen area at a set acquisition frequency of 30 frames per second, and the acquisition frequency can be adjusted according to the actual screen area environment. According to the overlapping result of the current image of the acquisition device and the previous frame image, the change situation of the current image is analyzed. If there is an abnormal overlapping result, it is judged as a suspected secretly photographing behavior. Once a suspected secretly photographing behavior is found, the current images of all acquisition devices are obtained and preprocessed to accurately obtain the screen image and lay a foundation for subsequent analysis. As mentioned in the background art, the gesture action is determined according to the trained gesture recognition model, and the edge detection algorithm is used to determine the object contour of the screen image, so as to accurately judge the possible secretly photographing devices (such as mobile phones, cameras, etc.). The lens pixel points of the object contour are extracted, and the effective lens pixel points are determined by means of fitting curves. Then, the lens brightness value is obtained according to the effective lens pixel points, and further, the lens brightness characteristics of the secretly photographing device are utilized to improve the accuracy of secretly photographing detection. The object contour value is searched in the preset object contour table, and the object contour table is obtained based on various object contours, such as (the dimensions of a cylinder and the dimensions of a cube, etc.). By combining the lens brightness value and the object contour value, the detection error caused by the specular reflection of other interfering objects can be avoided. For example: If only the lens brightness value is determined, detection deviation may be caused by the specular reflection of mirrors, watches, etc. If only the object contour value is determined, detection deviation may be caused by similar objects (mobile phones and portable notebooks). Therefore, when detecting secretly photographing in the screen area, it is necessary to analyze in combination with the actual situations of both, and comprehensive detection cannot be achieved only for simple gesture actions.
[0064] It can be understood that according to the determined gesture actions, the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value are output by using the candid photography behavior model. The candid photography behavior value is determined based on the contour weight factor, the brightness weight factor, the object contour value, and the lens brightness value. The brightness weight factor and the contour weight factor can balance the corresponding lens brightness value and the object contour value, avoiding data deviation caused by single data, thereby improving the accuracy of the candid photography behavior value. The candid photography behavior value is compared with the candid photography behavior threshold to dynamically determine whether to send a candid photography warning message to the signal jammer. When the candid photography warning message is sent, the signal jammer can cut off the network interconnection between the communication device and the outside world, further preventing information leakage caused by candid photography and improving the reliability of the security protection for the screen area.
[0065] In some embodiments of the present application, when determining whether there is a suspected candid photography behavior based on the overlapping result of the current image and the previous image of the acquisition device, it includes: when all the current images of all the acquisition devices completely overlap with all the previous images, it is determined that there is no suspected candid photography behavior in the screen area; when there is an incomplete overlap between the current image and the previous image of one acquisition device, it is determined that there is a suspected candid photography behavior in the screen area.
[0066] Specifically, when all the current images of all the acquisition devices completely overlap with all the previous images, at this time, there is no movement of the human body or object in the screen area, and it is determined that there is no suspected candid photography behavior in the screen area. When there is an incomplete overlap between the current image and the previous image of one acquisition device, it indicates that there is movement of the human body or object in the screen area, and it is determined that there is a suspected candid photography behavior in the screen area. Only when the current image changes, the suspected candid photography behavior that appears will be dynamically determined, avoiding misjudgment caused by factors such as environmental light changes and screen content changes, and improving the accuracy and reliability of candid photography detection.
[0067] In some embodiments of the present application, when acquiring the current images of all the acquisition devices and performing preprocessing, and determining the screen image of the screen area according to the preprocessing result, it includes: the preprocessing includes image denoising and geometric correction, and feature point extraction is performed on all the current images after preprocessing. All the extracted feature points are matched to determine the associated data between the current images. Based on the associated data, all the current images are registered and merged to obtain the screen image to be processed of the screen area. The screen image to be processed is adjusted to determine the screen image of the screen area, and the image adjustment includes adjusting the image size and removing the splicing gap.
[0068] Specifically, for image denoising, bilateral filtering is adopted. Bilateral filtering preserves the image details of the current image. Geometric correction corrects the distortions existing in all current images to ensure the alignment between current images, thereby eliminating unnecessary information that affects subsequent processing and improving the image quality. The Scale-Invariant Feature Transform (SIFT) is used for feature point extraction. SIFT can detect the same feature points at different scales. Regardless of how the scaling of the current image changes, SIFT can effectively identify feature points. Based on all the feature points extracted by SIFT, the Random Sample Consensus (RANSAC) algorithm is used for matching to determine the associated data between current images. The associated data includes the relative position and rotation relationship between images. According to the affine transformation, all current images are registered to eliminate interference factors such as rotation and deformation caused by different shooting angles of the acquisition device, improving the accuracy of the screen image to be processed. Multi-band fusion is used to merge all the registered current images to obtain the screen image to be processed. At this time, the screen image to be processed may have problems such as certain stitching gaps or inconsistent sizes due to being processed by multiple algorithms. Through image adjustment, the stitching quality of the screen image can be improved, laying a foundation for subsequent analysis.
[0069] In some embodiments of the present application, when determining the object contour of the screen image according to the edge detection algorithm and looking up the object contour value in the preset object contour table, it includes: calculating the gradient of the screen image using the Sobel operator, performing non-maximum suppression based on the Canny edge detection algorithm, and optimizing the edge detection result through the global optimal solution and the dynamic programming algorithm to extract the object contour of the screen image, and determining the corresponding object contour value in the preset object contour table based on the object contour. The object contour table includes the mapping relationship between the object contour and the corresponding object contour value.
[0070] Specifically, the Sobel operator is an edge detection operator that detects the edges of an image by calculating the gradient of the image. The Sobel operator convolves the screen image through two convolution kernels to detect the changes in the horizontal and vertical directions respectively, obtaining the gradient information of the screen image and then identifying the edges of the screen image. Based on the Canny edge detection algorithm, by suppressing the non-local maxima in the gradient direction and only retaining the local maxima of the edges, the edges are refined to avoid generating noise or false edges, and then an accurate edge contour is obtained. Through a global optimization algorithm (the global optimization method in image segmentation), the edge detection result is further adjusted to make the edge contour smoother and more consistent. The global optimal solution considers the global information of the entire screen image, can suppress local errors, and improve the accuracy of edge detection. The dynamic programming algorithm considers the relationship between adjacent pixels, eliminates the errors caused by local noise or complex object shapes during the edge detection process, and improves the integrity and consistency of object contour extraction. The preset object contour table records the object contours of different objects and their corresponding object contour values. The object contour values include length, angle, curvature, etc. By matching the extracted object contour with the object contour table, the object contour values in the screen image can be identified and determined, improving the accuracy and efficiency of obtaining the object contour values.
[0071] In some embodiments of the present application, when analyzing all lens pixel points to determine effective lens pixel points and obtaining the lens brightness value based on the effective lens pixel points, it includes: changing all lens pixel points to lens pixel position points, constructing a lens pixel coordinate system for all lens pixel position points, fitting all lens pixel position points to determine a lens fitting curve, removing the lens pixel position points that are not fitted, taking the lens pixel position points on the lens fitting curve as effective lens pixel points, obtaining the brightness values corresponding to all effective lens pixel points, and taking the mean value of all brightness values as the lens brightness value.
[0072] Specifically, in an actual screen image, although various image algorithms and preprocessing are used to avoid interference factors in the current image, it may still be affected by the reflection of other objects, resulting in deviations in the lens brightness value. For example, the optical reflection between a mobile phone camera and a mirror. Determine the lens pixel points based on the object contour, fit all the lens pixel position points, and the fitting process is implemented through OriginLab or Matlab. And remove the lens pixel position points that are not fitted, eliminating some invalid lens pixel position points existing in the screen image due to optical reflection. The lens pixel coordinate system is represented as a rectangular coordinate system. Obtain the brightness values corresponding to all effective lens pixel points, and take the mean value of all brightness values as the lens brightness value, comprehensively measuring the influence of all effective lens pixel points and avoiding the detection errors caused by the brightness value corresponding to a single effective lens pixel point.
[0073] In some embodiments of the present application, when changing all lens pixel points to lens pixel position points and constructing a lens pixel coordinate system for all lens pixel position points, it includes: taking the spatial position of the lens pixel point as the X-axis coordinate value of the lens pixel point, obtaining the lens data of the lens pixel point, establishing a lens data sequence based on the lens data, obtaining a qualified lens data sequence corresponding to the lens data sequence, comparing the lens data sequence with the qualified lens data sequence. When the lens data in the lens data sequence is greater than the qualified lens data in the qualified lens data sequence, the lens data is classified into the first lens data set. When the lens data in the lens data sequence is equal to the qualified lens data in the qualified lens data sequence, the lens data is classified into the second lens data set. When the lens data in the lens data sequence is less than the qualified lens data in the qualified lens data sequence, the lens data is classified into the third lens data set. Calculating the Y-axis coordinate value of the lens pixel point according to the first lens data set, the second lens data set, and the third lens data set, and constructing a lens pixel coordinate system according to the X-axis coordinate value and the Y-axis coordinate value.
[0074] In some embodiments of the present application, when calculating the Y-axis coordinate value of the lens pixel point according to the first lens data set, the second lens data set, and the third lens data set, it includes: counting the first lens quantity of the first lens data set, the second lens quantity of the second lens data set, and the third lens quantity of the third lens data set. The Y-axis coordinate value of the lens pixel point is obtained from the following formula:
[0075] ;
[0076] wherein, R represents the Y-axis coordinate value of the lens pixel point, YA represents the mean value of the lens data in the first lens data set, n represents the first lens quantity, Yi represents the i-th lens data in the first lens data set, Yg represents the qualified lens data corresponding to the i-th lens data in the first lens data set, YB represents the mean value of the lens data in the third lens data set, m represents the third lens quantity, Yj represents the j-th lens data in the third lens data set, and Yh represents the qualified lens data corresponding to the j-th lens data in the third lens data set.
[0077] Specifically, the spatial position of the lens pixel points serves as the X-axis coordinate value of the lens pixel points. For example, only considering the spatial position of the lens pixel points along the x-axis (aligned with the screen direction), or the spatial position of the lens pixel points along the y-axis (perpendicular to the screen direction), or the spatial position of the lens pixel points along the z-axis (perpendicular to the ground direction), the lens data includes hue, saturation, etc. The qualified lens data in the qualified lens data sequence corresponds one-to-one with the lens data. By dynamically partitioning all the lens data, the intelligent level of sneak shot detection is improved, thereby improving the accuracy of the Y-axis coordinate value and ensuring the reliability of the lens pixel coordinate system.
[0078] In some embodiments of the present application, when outputting the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value based on the gesture action using the sneak shot behavior model, it includes: obtaining the historical suspected sneak shot data set, dividing the historical suspected sneak shot data set into a training set and a test set, pre-selecting a convolutional neural network model, substituting the training set into the convolutional neural network model for training, verifying the trained convolutional neural network model according to the test set to determine the sneak shot behavior model. When the verification value is greater than or equal to the verification value after the previous training, the currently trained convolutional neural network model is used as the sneak shot behavior model. When the verification value is less than the verification value after the previous training, grid search is used to find the hyperparameters of the convolutional neural network model and replace the hyperparameters of the currently trained convolutional neural network model. The training set is clustered to determine the training sample set, the training set is re-determined according to the training sample set, the re-determined training set is substituted into the convolutional neural network model with the replaced hyperparameters, and training is continued until the sneak shot behavior model is determined. Then, the gesture action, the object contour value, and the lens brightness value are substituted into the sneak shot behavior model to obtain the contour weight factor and the brightness weight factor.
[0079] Specifically, the historical suspected sneak shot data set includes various gesture actions, various object contour values, various lens brightness values, various contour weight factors, various historical brightness weight factors, and the overall brightness change data (ambient light intensity) under different illuminations in the screen area. The historical suspected sneak shot data set is divided into a training set and a test set. Usually, 50%-70% of the data is used as the training set, and the rest is used as the test set. The training set is used to train the convolutional neural network model, and the test set is used to verify the performance of the trained model. A convolutional neural network model is pre-selected. The convolutional neural network model contains multiple layers and different types of activation functions, aiming to capture the complex relationships in the data. The data in the training set is used to train the convolutional neural network model. During the training process, the convolutional neural network model will learn the patterns and relationships in the data and iterate to improve its prediction or classification ability. After each training, the data in the test set is used to verify the model. The verification includes accuracy rate, loss function value, etc., which are used to comprehensively measure the performance of the model.
[0080] It is understandable that if the current verified value after training is greater than or equal to the verified value after the previous training, it indicates that the performance of the model has improved or remained relatively stable. At this time, the training is stopped and the currently trained convolutional neural network model is used as the peeping behavior model. If the current verified value after training is less than the verified value after the previous training, it indicates that the performance of the model has declined. Grid search is used to find the hyperparameters of the convolutional neural network model and replace the hyperparameters of the currently trained convolutional neural network model to change the learning rate of the model, which helps to find the most suitable network architecture and training strategy. The training process is repeated until the verified value is greater than or equal to the verified value after the previous training, and then the currently trained convolutional neural network model is used as the peeping behavior model, ensuring the data extraction and output capabilities of the model, enabling the peeping behavior model to make judgments from different angles, thereby improving the accuracy of the output contour weight factor and brightness weight factor of the model.
[0081] In some embodiments of the present application, when clustering the training set to determine the training sample set and re-determining the training set according to the training sample set, it includes:
[0082] S30: Determine the initial number of clusters K according to the training set, and randomly select K points as the initial cluster centers.
[0083] S31: Assign each data point in the training set to the nearest initial cluster center.
[0084] S32: Calculate the mean of all data points in the cluster and use it as the new cluster center of the cluster.
[0085] S33: Repeat S31 and S32 until the cluster centers no longer change or reach a predetermined number of iterations, and use the finally determined clusters of the clustering as the training sample set.
[0086] S34: Interpolate between every two data points in the training sample set to determine new data points and add them to the training set.
[0087] Specifically, the k-means clustering algorithm can identify potential patterns or structures in the training set. The initial number of clusters K can be adjusted according to the specific application scenario of the screen area. By interpolating to generate new data points, the number of data samples in the training set is effectively increased, enabling the model to be trained in a wide range of scenarios, thereby improving its adaptability to complex scenarios. The SMOTE algorithm is used for interpolation, aiming to generate new data points between the original data points. It also allows the model to be trained on the re-determined training set, avoiding overfitting of the model, enabling the model to make predictions on unseen data, and thus improving the generalization ability. The clustering algorithm and the interpolation algorithm help to address the problem of class imbalance in the training set. When there are fewer samples of certain classes in the training set, more samples can be generated through interpolation, thereby enhancing the class balance. The re-determined training set effectively avoids the deviation of the model's accuracy caused by bias towards a certain class during the training process, improves the accuracy of the contour weight factor and the brightness weight factor, and ensures the reliability of subsequent calculations.
[0088] In some embodiments of the present application, when determining the peeping behavior value according to the contour weight factor, the brightness weight factor, the object contour value, and the lens brightness value, comparing the peeping behavior value with the peeping behavior threshold, and judging whether to send a peeping warning message to the signal jammer according to the comparison result, it includes:
[0089] The peeping behavior value is obtained according to the following formula:
[0090] ;
[0091] Wherein, Q represents the peeping behavior value, w1 represents the contour weight factor, S1 represents the object contour value, w2 represents the brightness weight factor, S2 represents the lens brightness value. A peeping behavior threshold Qa is preset. When Q ≥ Qa, it is determined to send a peeping warning message to the signal jammer. When Q < Qa, it is determined not to send a peeping warning message to the signal jammer.
[0092] Specifically, the contour weight factor and the brightness weight factor can comprehensively consider the characteristics of the object contour and the characteristics of the lens brightness, avoiding errors in the peeping behavior value caused by a single data. For example, for optical reflections and other similar peeping objects, the corresponding object contour value and the lens brightness value are balanced through the contour weight factor and the brightness weight factor, improving the accuracy of peeping behavior detection. When the peeping behavior value exceeds the peeping behavior threshold, it is determined that there is a peeping behavior of the screen behavior. At this time, a peeping warning message is sent to the signal jammer, and the signal jammer receives the peeping warning message and cuts off the mobile network and the wireless network, avoiding the network interconnection between the communication device and the outside world, preventing information leakage and network dissemination. In addition, a screen controller is added to the host or LED large screen of the screen for peeping detection, and the refresh frequency of the screen, inserting black frames, and PWM dimming are changed through the screen controller. Using the frequency range between the highest recognizable update frequency of the human eye and the shutter frequency of the photographing device, the update frequency is set for the controllable display window of the screen, so that the update frequency of the controllable display window of the screen is between the highest recognizable update frequency of the human eye and the shutter frequency of the photographing device, and using the frequency difference between the human eye, the peeping device and the controllable display window of the screen, the peeping device and the controllable display window of the screen are out of sync without affecting the viewing effect of the human eye, so that a complete picture cannot be obtained, further preventing information leakage and effectively reducing the peeping behavior.
[0093] It can be understood that when the refresh frequency is 275Hz, there is no abnormal visual effect when the human eye is watching, but it is out of sync with the shutter speeds of common mobile phones and camera devices (such as 50Hz / 60Hz or 1 / 120 seconds). And, a short black frame is inserted in the refresh cycle of the screen LED (such as inserting 1 black frame every 10 frames). The human eye ignores the black frame due to visual persistence, but mobile phones and camera devices will flicker due to the synchronization of the black frame and the shutter during shooting, avoiding information leakage. When dimming with PWM, when the brightness of the screen LED is high-frequency PWM (such as 2000Hz), the human eye cannot perceive the stroboscopic effect, but when mobile phones and camera devices shoot, there will be brightness fluctuations in the shooting results due to the out-of-sync sampling interval and PWM cycle. When using low-frequency PWM (such as 200Hz), the human eye may perceive a certain stroboscopic effect at low brightness, and mobile phones and camera devices will enhance the stroboscopic effect during shooting, resulting in poor image quality of the shot. Through collection by the collection device and intelligent judgment, the peeping behavior is detected in real time using image recognition and gesture actions, improving the accuracy and real-time performance of peeping detection. Changing the refresh frequency of the screen, inserting black frames, and PWM dimming through the screen controller effectively prevent information leakage and enhance the security protection ability of the screen.
[0094] In summary, the beneficial effects of the present invention are as follows: Through the image acquisition of multiple acquisition devices by real-time visual recognition, combined with the analysis of the object contour, the calculation of the lens brightness value, and the capture of gesture actions, it is possible to accurately identify the secret shooting behavior, avoiding relying on a single piece of data or simple rule matching, improving the reliability of secret shooting behavior detection. By adopting the method of evenly distributing multiple acquisition devices to cover different angles of the screen area, the detection ability of secret shooting behavior in complex environments is improved, avoiding the problems of occlusion or perspective limitation of a single acquisition device. Based on the secret shooting behavior model, the contour weight factor and the brightness weight factor are determined, comprehensively measuring the object contour and the lens brightness, improving the accuracy of calculating the secret shooting behavior value. By comparing the secret shooting behavior value with the secret shooting behavior threshold in real time, a warning message can be sent to the signal jammer at the first time when the secret shooting behavior occurs, thereby preventing the normal operation of the secret shooting device, effectively protecting personal privacy, and avoiding information leakage.
[0095] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0096] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention. Any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition, characterized in that, Including: Determine the screen area, evenly deploy multiple acquisition devices in the screen area, and set the acquisition frequency of each acquisition device; Judge whether there is a suspected secret shooting behavior according to the overlapping result of the current image and the previous image of the acquisition device. When there is the suspected secret shooting behavior, obtain the current images of all acquisition devices and perform preprocessing, and determine the screen image of the screen area according to the result of the preprocessing; Identify the gesture actions of the screen image, determine the object contour of the screen image according to the edge detection algorithm, and look up the object contour value in the preset object contour table. Determine the lens pixel points according to the object contour, analyze all the lens pixel points, determine the effective lens pixel points, and obtain the lens brightness value according to the effective lens pixel points; Based on the gesture actions, use the secret shooting behavior model to output the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value. Determine the secret shooting behavior value according to the contour weight factor, the brightness weight factor, the object contour value and the lens brightness value, compare the secret shooting behavior value with the secret shooting behavior threshold, and judge whether to send a secret shooting warning message to the signal jammer according to the comparison result; When analyzing all the lens pixel points to determine the effective lens pixel points and obtaining the lens brightness value according to the effective lens pixel points, it includes: Change all the lens pixel points to lens pixel position points, construct a lens pixel coordinate system with all the lens pixel position points, fit all the lens pixel position points, and determine the lens fitting curve; Eliminate the lens pixel position points that are not fitted, use the lens pixel position points on the lens fitting curve as the effective lens pixel points, obtain the brightness values corresponding to all the effective lens pixel points, and use the average value of all the brightness values as the lens brightness value. The lens pixel coordinate system is represented as a rectangular coordinate system; When changing all the lens pixel points to lens pixel position points and constructing a lens pixel coordinate system with all the lens pixel position points, it includes: Use the spatial position of the lens pixel point as the X-axis coordinate value of the lens pixel point, obtain the lens data of the lens pixel point, and establish a lens data sequence based on the lens data; Obtain a qualified lens data sequence corresponding to the lens data sequence, and compare the lens data sequence with the qualified lens data sequence; When the lens data in the lens data sequence is greater than the qualified lens data in the qualified lens data sequence, divide the lens data into the lens first data set; When the lens data in the lens data sequence is equal to the qualified lens data in the qualified lens data sequence, divide the lens data into the lens second data set; When the lens data in the lens data sequence is less than the qualified lens data in the qualified lens data sequence, divide the lens data into the lens third data set; Calculate the Y-axis coordinate value of the lens pixel point according to the lens first data set, the lens second data set and the lens third data set, and construct a lens pixel coordinate system according to the X-axis coordinate value and the Y-axis coordinate value; When calculating the Y-axis coordinate value of the lens pixel points based on the first lens data set, the second lens data set and the third lens data set, it includes: Count the first lens quantity in the first lens data set, the second lens quantity in the second lens data set and the third lens quantity in the third lens data set; The Y-axis coordinate value of the lens pixel points is obtained by the following formula: ; Wherein, R represents the Y-axis coordinate value of the lens pixel points, YA represents the mean value of the lens data in the first lens data set, n represents the first lens quantity, Yi represents the i-th lens data in the first lens data set, Yg represents the qualified lens data corresponding to the i-th lens data in the first lens data set, YB represents the mean value of the lens data in the third lens data set, m represents the third lens quantity, Yj represents the j-th lens data in the third lens data set, and Yh represents the qualified lens data corresponding to the j-th lens data in the third lens data set.
2. The method for detecting surreptitious photographing of a monitored screen area based on real-time visual recognition according to claim 1, wherein, When judging whether there is a suspected secret photographing behavior according to the overlapping result of the current image and the previous image of the acquisition device, it includes: When all the current images of all the acquisition devices completely overlap with all the previous images, it is determined that there is no suspected secret photographing behavior in the screen area; When there is an incomplete overlap between the current image and the previous image of one acquisition device, it is determined that there is a suspected secret photographing behavior in the screen area.
3. The method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition according to claim 2, characterized in that, When obtaining the current images of all the acquisition devices and performing preprocessing, and determining the screen image of the screen area according to the preprocessing result, it includes: The preprocessing includes image denoising and geometric correction, feature point extraction is performed on all the preprocessed current images, all the extracted feature points are matched, and the associated data between the current images is determined; Based on the associated data, all the current images are registered and merged to obtain the to-be-processed screen image of the screen area, and the to-be-processed screen image is adjusted to determine the screen image of the screen area, and the image adjustment includes adjusting the image size and removing the splicing gap.
4. The method for detecting secret photographing of a monitoring screen area based on real-time visual recognition according to claim 3, wherein When determining the object contour of the screen image according to the edge detection algorithm and searching for the object contour value in the preset object contour table, it includes: Use the Sobel operator to calculate the gradient of the screen image, perform non-maximum suppression based on the Canny edge detection algorithm, and optimize the edge detection result through the global optimal solution method and the dynamic programming algorithm to extract the object contour of the screen image; Based on the object contour, determine the corresponding object contour value in the preset object contour table, and the object contour table includes the mapping relationship between the object contour and the corresponding object contour value.
5. The method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition according to claim 4, wherein, When outputting the contour weight factor of the object contour value and the brightness weight factor of the lens brightness value by using the secret photographing behavior model based on the gesture action, it includes: Obtain the historical suspected secret photographing data set, and divide the historical suspected secret photographing data set into a training set and a test set; Pre-select a convolutional neural network model, and substitute the training set into the convolutional neural network model for training; Verify the trained convolutional neural network model according to the test set to determine the peeping behavior model. When the verification value is greater than or equal to the verification value after the previous training, the currently trained convolutional neural network model is used as the peeping behavior model; When the verification value is less than the verification value after the previous training, use grid search to find the hyperparameters of the convolutional neural network model, replace the hyperparameters of the currently trained convolutional neural network model, cluster the training set to determine the training sample set, re-determine the training set according to the training sample set, substitute the re-determined training set into the convolutional neural network model with replaced hyperparameters, and continue training until the peeping behavior model is determined; Substitute the gesture action, the object contour value, and the lens brightness value into the peeping behavior model to obtain the contour weight factor and the brightness weight factor.
6. The method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition according to claim 5, characterized in that, When clustering the training set to determine the training sample set and re-determining the training set according to the training sample set, it includes: S30: Determine the initial number of clusters K according to the training set, and randomly select K points as the initial cluster centers; S31: Assign each data point in the training set to the nearest initial cluster center; S32: Calculate the mean of all data points in the cluster and use it as the new cluster center of the cluster; S33: Repeat S31 and S32 until the cluster centers no longer change or reach the predetermined number of iterations, and use the finally determined clusters of the clustering as the training sample set; S34: Interpolate between every two data points in the training sample set to determine new data points and add them to the training set.
7. The method for detecting surreptitious photographing of a monitoring screen area based on real-time visual recognition according to claim 6, wherein, When determining the peeping behavior value according to the contour weight factor, the brightness weight factor, the object contour value, and the lens brightness value, comparing the peeping behavior value with the peeping behavior threshold, and judging whether to send a peeping warning message to the signal jammer according to the comparison result, it includes: The peeping behavior value is obtained according to the following formula: ; where Q represents the peeping behavior value, w1 represents the contour weight factor, S1 represents the object contour value, w2 represents the brightness weight factor, and S2 represents the lens brightness value; Preset the peeping behavior threshold Qa; When Q≥Qa, it is determined to send a peeping warning message to the signal jammer; When Q<Qa, it is determined not to send a peeping warning message to the signal jammer.
Citation Information
Patent Citations
Screen secret photography prevention method and system based on visual identification
CN117953574A
Aero-engine anomaly detection method based on deep attention data enhancement
CN116776265A
Behavior discrimination method and device, computer equipment and storage medium
CN119229535A