Low-energy-consumption intelligent anti-screen-shooting method and system based on multi-thread concurrency and key frame extraction

Through the low-energy intelligent anti-shot screen method based on multi-threaded concurrency and keyframe extraction, the problems of low information acquisition efficiency and inaccurate keyframe extraction in the existing technology are solved, and the effect of reducing energy consumption and improving the intelligence level of equipment management while ensuring information security by the Public Security Bureau is achieved.

CN120183159APending Publication Date: 2025-06-20WUJIANG DISTRICT PUBLIC SECURITY BUREAU SUZHOU CITY
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510654976.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology has many shortcomings in data security and intelligent management, including low information acquisition efficiency, inaccurate keyframe extraction, rough analysis of object motion states, incomplete feature extraction, in-depth risk assessment and incoordinated energy consumption management.

Method used

The low-energy intelligent anti-shoot screen method based on multi-threaded concurrency and keyframe extraction is adopted. By obtaining target screen information, preset multi-threaded concurrent image detection model, training sample set and verification set, the target multi-threaded concurrent image detection model is trained, and the target screen information is processed in combination with the keyframe extraction algorithm, real-time motion state and screen-sight risk characteristics of the target screen object are generated, stable frame images are selected, target detection is performed, and early warning information is determined by combining the screen-sight risk characteristics. At the same time, real-time status of the device and preset energy consumption management information are obtained, and energy consumption management instructions are comprehensively generated.

Benefits of technology

It realizes the ability to respond to low-energy consumption requirements while ensuring information security, improves the intelligence level of equipment management, integrates multi-threaded concurrency, keyframe extraction and other technologies to realize anti-shot screen, anti-peeping and energy consumption reduction functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183159A_ABST
    Figure CN120183159A_ABST
Patent Text Reader

Abstract

The invention provides a low-energy-consumption intelligent anti-screen-shooting method and system based on multi-thread concurrency and key frame extraction, and is applied to the technical field of data processing. According to the method, key frame extraction is performed on target screen information based on a target multi-thread concurrent image detection model and a key frame extraction algorithm, and continuous three-frame images of a target screen object are generated; filtering the continuous three-frame image of the target screen object based on a non-maximum suppression method to generate a target key frame image; processing the target key frame image to generate a real-time motion state and a screen peeping risk feature of the target screen object; processing the real-time motion state of the target screen object based on a preset dynamic threshold value and the screen peeping risk feature to generate a stable frame image; the stable frame image is processed, abnormal operation and screen peeping early warning information is generated, and energy consumption management is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a low-power intelligent anti-screen capture method and system based on multi-threaded concurrency and key frame extraction. Background Art

[0002] With the rapid development of information technology, data security has received increasing attention, especially in scenarios involving sensitive information. There are many deficiencies in the existing technologies: the information acquisition technology is inefficient and vulnerable to interference, making it difficult to ensure real-time performance and accuracy; the key frame extraction algorithm cannot accurately reflect the object motion state, affecting the judgment of abnormal behaviors; the object motion state analysis method is rough and difficult to identify complex motion patterns; the features extracted by the image screening and analysis technology are incomplete and inaccurate, resulting in large deviations in the object motion trajectory information and affecting the motion state assessment; the risk assessment and early warning do not deeply analyze the screen peeping risk characteristics and cannot comprehensively evaluate the risk accurately considering multiple factors and give early warnings in a timely manner; the energy consumption management cannot combine the real-time state of the device and the safety warning information, making it difficult to achieve the coordination of security monitoring and energy consumption optimization.

[0003] At the level of risk assessment and early warning, the analysis of the screen peeping risk characteristics is not deep enough, and it is impossible to comprehensively consider multiple factors such as the sensitivity of the screen display content and the distribution of surrounding personnel, resulting in inaccurate screen peeping risk assessment and inability to give early warnings effectively and in a timely manner. At the same time, in terms of energy consumption management, the existing technologies fail to fully combine the real-time state of the device and the safety warning information, cannot reasonably generate energy consumption management instructions, and it is difficult to achieve the coordination of security monitoring and energy consumption optimization. These deficiencies of the existing technologies urgently require innovative technical solutions to solve in order to meet the growing needs of data security and intelligent management.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of this application is to provide a low-power intelligent anti-screen capture method and system based on multi-threaded concurrency and key frame extraction, which at least overcome the problems existing in the prior art to a certain extent. By obtaining the target screen information, preset models, training sample sets, and validation sets. Through the processing of the sample sets, a target multi-threaded concurrent image detection model is trained to lay the foundation for subsequent detections. Then, the key frame extraction algorithm is used to process the target screen information. First, the model is used for preprocessing to generate image data blocks, and then through processing such as the three-frame difference method, continuous multi-frame data is generated, the object motion state information is analyzed, potential key frames are screened out, and then continuous three-frame images of the target screen object are generated. After that, the non-maximum suppression method is used to filter these images to obtain the target key frame images.

[0006] Extract features from the target key-frame image, compare the feature vectors of adjacent frames to determine the object's motion trajectory, evaluate the motion state in combination with the preset behavior pattern, and at the same time generate the screen-peeping risk features based on the environmental information to form a real-time motion state and screen-peeping risk feature report. Based on this, use the preset dynamic threshold and the screen-peeping risk features to process the real-time motion state and filter out the stable frame images. Finally, perform object detection on the stable frame images, judge whether there are abnormal operations according to the preset rules, and generate warning information in combination with the screen-peeping risk features. At the same time, obtain the real-time state of the device and the preset energy consumption management information, comprehensively generate energy consumption management instructions, control the device operation, and achieve early warning of abnormal operations and energy consumption optimization.

[0007] Other features and advantages of the present application will become apparent from the following detailed description, or be learned in part through the practice of the present invention.

[0008] According to one aspect of the present application, there is provided a low-energy intelligent anti-screen-peeping method based on multi-threaded concurrency and key-frame extraction, including: obtaining target screen information, a preset multi-threaded concurrent image detection model, a training sample set, and a validation set; training and processing the preset multi-threaded concurrent image detection model based on the training sample set and the validation set to generate a target multi-threaded concurrent image detection model; performing key-frame extraction on the target screen information based on the target multi-threaded concurrent image detection model and the key-frame extraction algorithm to generate three consecutive frames of images of the target screen object; performing filtering processing on the three consecutive frames of images of the target screen object based on the non-maximum suppression method to generate target key-frame images; processing the target key-frame images to generate the real-time motion state and screen-peeping risk features of the target screen object; processing the real-time motion state of the target screen object based on the preset dynamic threshold and the screen-peeping risk features to generate stable frame images; processing the stable frame images to generate abnormal operation and screen-peeping warning information and optimize energy consumption management.

[0009] Another aspect of the present application is a low-power intelligent anti-screen capture device based on multi-threaded concurrency and key frame extraction, which is characterized by including: an acquisition module for acquiring target screen information, a preset multi-threaded concurrent image detection model, a training sample set, and a validation set; a processing module for training and processing the preset multi-threaded concurrent image detection model based on the training sample set and the validation set to generate a target multi-threaded concurrent image detection model; extracting key frames from the target screen information based on the target multi-threaded concurrent image detection model and a key frame extraction algorithm to generate three consecutive frames of images of the target screen object; filtering the three consecutive frames of images of the target screen object based on a non-maximum suppression method to generate a target key frame image; processing the target key frame image to generate the real-time motion state and screen peeping risk characteristics of the target screen object; processing the real-time motion state of the target screen object based on a preset dynamic threshold and the screen peeping risk characteristics to generate a stable frame image; processing the stable frame image to generate abnormal operation and screen peeping warning information and optimize energy consumption management.

[0010] According to still another aspect of the present application, an electronic device is characterized by including: a first processor; and a memory for storing executable instructions of the first processor; wherein, the first processor is configured to execute the above-mentioned low-power intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction by executing the executable instructions.

[0011] According to yet another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a second processor, it implements the above-mentioned low-power intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction.

[0012] A low-power intelligent anti-screen capture method and system based on multi-threaded concurrency and key frame extraction provided by the present application acquire target screen information, a preset model, a training sample set, and a validation set. Through the processing of the sample set, a target multi-threaded concurrent image detection model is trained, laying a foundation for subsequent detection. Then, the key frame extraction algorithm is used to process the target screen information. First, the model is used for preprocessing to generate image data blocks, and then through processing such as three-frame difference method, multiple consecutive frames of data are generated. By analyzing, the object motion state information is obtained, potential key frames are screened out, and then three consecutive frames of images of the target screen object are generated. After that, a non-maximum suppression method is used to filter these images to obtain the target key frame image. Feature extraction is performed on the target key frame image, the object motion trajectory is determined by comparing the adjacent frame feature vectors, the motion state is evaluated in combination with a preset behavior pattern, and at the same time, the screen peeping risk characteristics are generated based on the environmental information to form a real-time motion state and screen peeping risk characteristics report. Based on this, the real-time motion state is processed using a preset dynamic threshold and the screen peeping risk characteristics, and stable frame images are screened out.

[0013] Finally, perform object detection on the stable frame image, determine whether there is an abnormal operation according to the preset rules, and generate a warning message in combination with the screen peeping risk characteristics. At the same time, obtain the real-time status of the device and the preset energy consumption management information, comprehensively generate an energy consumption management instruction, control the device operation, and achieve early warning of abnormal operations and energy consumption optimization. This technology ensures the information security of the public security bureau while meeting the requirements of low energy consumption and improving the intelligent level of device management. Centering on the needs of the public security bureau, integrating technologies such as multi-threaded concurrency and key frame extraction, it realizes functions of anti-screen shooting, anti-peeping, and energy consumption reduction, ensuring information security and reducing energy consumption.

[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The flowchart showing a low-energy consumption intelligent anti-screen shooting method based on multi-threaded concurrency and key frame extraction provided by an embodiment of the present application;

[0016] Figure 2 The structural schematic diagram showing a low-energy consumption intelligent anti-screen shooting device based on multi-threaded concurrency and key frame extraction provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0018] The following combines Figure 1 to describe a low-energy consumption intelligent anti-screen shooting method based on multi-threaded concurrency and key frame extraction according to an exemplary embodiment of the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application are applicable to any applicable scenario.

[0019] In one embodiment, the present application also proposes a low-energy consumption intelligent anti-screen shooting method and system based on multi-threaded concurrency and key frame extraction. Figure 1 Schematically shows a flowchart of a low-energy consumption intelligent anti-screen shooting method based on multi-threaded concurrency and key frame extraction according to an embodiment of the present application. As Figure 1 shown, this method is applied to a server and includes:

[0020] S101, obtain target screen information, a preset multi-threaded concurrency image detection model, a training sample set, and a validation set.

[0021] In one implementation, the target screen information is sourced from various computer screens within the public security bureau, such as office computers, conference room display screens, etc. The information covers the content of the real-time display on the screen, including various types of data such as text, images, videos, etc., as well as the basic parameters of the screen, such as resolution (commonly 1920×1080), frame rate (generally 60Hz), color mode (such as RGB mode), etc. These information are the basic data for subsequent analysis. For example, in the office scenario of the public security bureau, the screen may display case materials, surveillance videos, etc. Obtaining these real-time picture contents and screen parameters is crucial for detecting whether there are abnormal operations. Through screen capture technology, the screen is sampled at a specific time interval (such as every 100 milliseconds) to obtain continuous screen image data, ensuring the real-time and integrity of the information.

[0022] The preset multi-threaded concurrent image detection model adopts a convolutional neural network (CNN) architecture based on deep learning, such as the improved YOLO (You Only Look Once) model. This model is divided into an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The input layer receives the preprocessed image data; the convolutional layer extracts image features, such as edges, textures, etc., through convolutional kernels; the pooling layer is used to reduce the amount of data and retain key features; the fully connected layer integrates the extracted features; the output layer outputs the detection results, including information such as object category, location, confidence, etc. The multi-threaded concurrent processing ability is achieved by dividing the image data into multiple sub-tasks and allocating them to different threads for parallel processing during the operation of the model, improving the detection efficiency. In terms of related parameters, the size of the convolutional kernel is set according to experiments and experience, such as the commonly used 3×3 convolutional kernel for capturing local features; the stride is set to 1 or 2 to control the sliding stride of the convolutional operation; the learning rate is set to 0.001 to control the step size of parameter update during model training and avoid falling into local optimal solutions during the training process; the batch size is set to 32, that is, 32 images are processed simultaneously each time during training to balance memory usage and training efficiency. During the training process, the model parameters are continuously adjusted through the backpropagation algorithm to enable the model to accurately identify various objects.

[0023] The training sample set and the validation set are made by collecting a large number of image data in different scenarios. The image sources include the daily office scenarios, meeting scenarios, surveillance videos, etc. of the public security bureau, covering a variety of lighting conditions, different angles, and different personnel behaviors. The sample set contains images in normal operation scenarios, such as people using computers for office work normally and the pictures during meeting discussions, and also contains images in abnormal operation scenarios, such as taking pictures of the screen with a mobile phone and people approaching the screen abnormally, etc. Each image is accurately annotated, and the annotation information includes the category of the object in the image (such as mobile phone, person), the location (represented by the coordinates of a rectangular box), the behavior type (normal operation, abnormal operation), etc. The training sample set and the validation set are divided according to a certain ratio, such as 8:2, that is, 80% of the data is used to train the model, and 20% of the data is used to verify the performance of the model. During the production process, the images are preprocessed, including operations such as cropping, scaling, and normalization, to make them meet the model input requirements. For example, all images are uniformly scaled to the size required for model input (such as 416×416 pixels), and the pixel values are normalized to the interval [0,1] to improve the stability and efficiency of model training. These sample sets and validation sets provide rich data for model training, enabling it to learn the features in different scenarios, thereby improving the accuracy and generalization ability of detection.

[0024] S102, perform training processing on a preset multi-threaded concurrent image detection model based on the training sample set and the validation set to generate a target multi-threaded concurrent image detection model.

[0025] In one implementation, count the number of image features in the training sample set and generate a sampling ratio based on the feature distribution balance. When making the training sample set, image data from multiple channels such as the daily office scenarios, meeting scenarios, and surveillance videos of the public security bureau are collected. These images contain rich features, such as colors, textures, shapes, and the spatial relationships between objects, etc. For example, in the office scenario images, there are document texts on the computer screen (manifested as specific texture and shape features), the colors and styles of the clothing of the personnel (color and texture features); in the meeting scenario images, there are the projected picture contents of the projector (a mixture of various features), the postures of the participants (shape and spatial relationship features), etc.

[0026] Perform a detailed statistics on various types of features in these images and record the frequency of occurrence of each feature. Suppose when counting the color features, it is found that blue appears 100 times in the image, green appears 80 times, etc. By analyzing the distribution of these features, if the frequency of some features is too high or too low, it may lead to bias in model training. In order to enable the model to learn various features more evenly, calculate the sampling ratio based on the balance of feature distribution. For features with a lower frequency of occurrence, give a higher sampling ratio to increase the probability of being selected in subsequent training; for common features, appropriately reduce the sampling ratio. For example, if the image features under a certain rare lighting condition have a low frequency of occurrence, increase its sampling ratio to ensure that the model can fully learn the features in this special case. The generated sampling ratio can ensure that all types of features are reasonably represented during the training process, avoiding over-learning or under-learning of certain features by the model.

[0027] Perform multi-modal stratified sampling on the training sample set based on the sampling ratio to generate a preset number of sampled feature combinations. Multi-modal means sampling from different feature dimensions, such as color, texture, shape, etc. Stratified sampling is to sample separately according to different feature levels. For example, first stratify according to the overall scene of the image (such as office scene, meeting scene), and then sample according to specific feature categories within each scene layer. Taking color and shape features as an example, in the office scene layer, select a certain number of images with specific color features according to the color sampling ratio, such as selecting a certain number of images containing blue computer casings; then according to the sampling ratio of the shape feature, select images containing specific-shaped objects (such as mobile phone shape, document shape). Combine the features sampled from different modalities and levels to form a preset number of sampled feature combinations. These combinations contain diverse image feature information. For example, a sampled feature combination may contain information such as the shape and position of a mobile phone of a specific color in a specific office scene. Through this multi-modal stratified sampling process, rich and diverse and representative sampled feature combinations can be generated, providing more comprehensive data for subsequent model training, helping the model learn the rules under different scenarios and feature combinations, and improving the generalization ability of the model.

[0028] Based on any key image feature and each combination of sampling features, statistical tests are conducted to divide the data into a group with good detection effect and a group with poor detection effect. Each group contains a preset number of data samples, and at least one sample has error identification information. Select a key image feature from numerous image features. For example, select "the relative distance between the position of the mobile phone in the image and the screen" as the key feature. Combine this key feature with each sampling feature for statistical tests. Statistical tests can use methods such as hypothesis testing to analyze the distribution and differences of the key feature in different combinations of sampling features. According to the test results, the data is divided into a group with good detection effect and a group with poor detection effect. When dividing, it is set that each group contains a preset number of data samples. For example, each group is set to 100 samples. And to enable the model to learn about error situations, at least one sample in each group is guaranteed to have error identification information. The error identification information can be annotation errors (such as mislabeling a person using a computer normally as using a mobile phone to take a screen shot), data loss (such as incomplete annotation information of key objects in some images), etc. For example, when conducting statistical tests on the key feature "the relative distance between the position of the mobile phone in the image and the screen", it is found that in some combinations of sampling features, the distribution of the distance between the mobile phone and the screen shows a large difference from the normal situation, and these combinations are divided into the group with poor detection effect; while those combinations that conform to the normal distribution range are divided into the group with good detection effect. After such division, the two groups of data respectively represent the data situations of the model under different detection effects, providing different sample sets for subsequent targeted training of the model.

[0029] Iteratively train the preset multi-threaded concurrent image detection model based on the group with good detection effect and the group with poor detection effect, optimize the model parameters through cross-validation, and generate the trained model and performance metrics. Use the data of the well-divided group with good detection effect and the group with poor detection effect to iteratively train the preset multi-threaded concurrent image detection model. During the training process, adopt the method of cross-validation to optimize the model parameters. Cross-validation divides the data into multiple subsets. For example, divide the data into 5 subsets, select 4 of them as the training set and 1 subset as the validation set each time. In each training, the model learns according to the data in the training set, adjusts its own parameters (such as weights and biases in the convolutional neural network), and then tests on the validation set to evaluate the performance of the model on this group of data. Through multiple cross-validations, continuously try different parameter settings, and select the parameter combination that makes the model perform optimally on the validation set. For example, when adjusting the learning rate, try different learning rate values (such as 0.001, 0.01, 0.1, etc.), observe the changes in performance metrics such as accuracy and recall rate of the model on the validation set, and finally determine the optimal learning rate. After multiple iterative trainings and cross-validations, generate the trained model and the corresponding performance metrics, such as accuracy, recall rate, F1 value, etc. These performance metrics are used to evaluate the model's ability to detect abnormal operations and screen peeping risks in the target screen information captured by the mobile phone, etc.

[0030] If the error identification information in the training result is recognized by the model as a key factor affecting the detection effect evaluation, then regard the trained model as the target multi-threaded concurrent image detection model. After the training is completed, conduct an in-depth analysis of the training result of the model. Check whether the model can identify the error identification information and determine whether it regards the error identification information as a key factor affecting the detection effect evaluation. If the model can accurately identify the error identification information, and during the detection process, these error identification information have a significant impact on evaluation metrics such as the accuracy and reliability of the detection result, it indicates that the model has good robustness and the ability to identify abnormal situations. For example, when testing the model, if the model can discover those mislabeled or data-lacking samples and make reasonable judgment adjustments for similar situations in subsequent detections, it shows that the model has effectively learned the abnormal patterns represented by the error identification information. At this time, determine this trained model as the target multi-threaded concurrent image detection model, which can more accurately detect abnormal operations and screen peeping risks in the target screen information, meet the requirements of the public security bureau for functions such as anti-screen shooting and anti-peeping, and provide reliable support for subsequent key frame extraction, abnormal operation judgment, etc.

[0031] S103, extract key frames from the target screen information based on the target multi-threaded concurrent image detection model and the key frame extraction algorithm, and generate three consecutive frames of images of the target screen object.

[0032] In one implementation, the target screen information is preprocessed based on the target multi-threaded concurrent image detection model to generate a number of image data blocks. In the actual scenarios of public security bureaus, the target screen information comes from the screen contents of devices such as office computers and conference room large screens, and this information is transmitted in real time in the form of video streams. The target multi-threaded concurrent image detection model adopts an improved YOLO model architecture, and uses its multi-threaded concurrent processing ability to split the input video stream into consecutive image frames in chronological order. For example, assuming the video stream frame rate is 60fps, the model obtains one frame of image every 1 / 60 seconds.

[0033] For these image frames, the model first performs noise reduction processing, using the Gaussian filtering algorithm to remove the noise points in the images and improve the image quality. Then, size normalization is carried out to uniformly adjust the images with different resolutions to the input size required by the model, such as 416×416 pixels. Taking an office computer screen image with an original resolution of 1920×1080 as an example, after scaling, the length and width of the image are both adjusted to 416 pixels, while maintaining the integrity of the image content. Finally, the normalized images are divided into multiple image data blocks according to certain rules. For example, each image is evenly divided into 16 8×8 small blocks. These image data blocks provide the basic data for subsequent key frame extraction, facilitating more efficient processing and analysis.

[0034] Based on the key frame extraction algorithm, a number of image data blocks are processed to generate consecutive multi-frame data. The three-frame difference method is used as the core part of the key frame extraction algorithm. For the image data blocks generated in the previous step, three consecutive data blocks are sequentially selected for processing in chronological order. Calculate the difference between the corresponding pixel points of adjacent two-frame data blocks, and set the connected region area threshold to 50 pixel². If the change in the pixel difference of a certain region causes the connected region area to exceed this threshold, it is determined that the object in this region is in a moving state and is marked as "1"; if it is lower than the threshold, it is determined to be in a stationary state and is marked as "0".

[0035] For example, in a video of a meeting scenario, a participant operates a mobile phone and the mobile phone moves in front of the screen. When processing the corresponding image data blocks, it is found through calculation that in a certain three-frame data blocks, the change in the pixel difference in the area where the mobile phone is located causes the connected region area to reach 80 pixel², exceeding the threshold, and this region is marked as the moving state "1". Arrange these state marks in chronological order to form a continuous state sequence, such as "1010", etc., and correspondingly generate consecutive multi-frame data. These data record the change in the motion state of the object in front of the screen, providing a basis for subsequent analysis of the object's motion.

[0036] Process consecutive multiple frames of data to generate object motion state information. Analyze information such as the motion trajectory and speed of the object based on the state marker sequence in the consecutive multiple frames of data. If in the consecutive multiple frames of data, the state marker of a certain area (such as the area where the mobile phone is located) continuously appears as "1", and the position changes significantly in different frames, by calculating the coordinate changes of this area in different frames, its motion direction and speed can be determined. For example, in 5 consecutive frames of data, the coordinates of the area where the mobile phone is located change from (100, 100) to (120, 105), (140, 110), (160, 115), (180, 120) in sequence. By calculating the coordinate differences and time intervals between adjacent frames (assuming the interval between each frame is 1 / 60 seconds), it can be obtained that the mobile phone moves in the lower right direction at a speed of about 1200 pixels per second during this period. At the same time, combined with the time information, record the changes in the object motion state over time, such as whether it accelerates, decelerates, or moves at a constant speed, and comprehensively obtain the object motion state information in front of the screen, providing key data support for judging whether there are abnormal operations.

[0037] Process the object motion state to generate potential key frames, where the potential key frame is the frame with the highest clarity and in a stationary state. Based on the analysis of the object motion state, determine the potential key frames according to the preset discrimination patterns (such as "10", "110", etc.). The discrimination pattern represents the change of the object motion state. For example, "10" means moving first and then stationary. When a frame sequence that conforms to the discrimination pattern is detected, select the potential key frame from the stationary frames (frames marked as "0"). Use a clarity detection algorithm (such as the Laplace algorithm) to calculate the clarity of the stationary frames. The Laplace algorithm measures the clarity of an image by calculating the second derivative of the image, and select the frame with the highest clarity index as the potential key frame.

[0038] Suppose in a video segment, the discrimination pattern of "110" is detected, that is, the object moves continuously twice and then stops. Apply the Laplace algorithm to calculate the gradient magnitude of the stationary frames among them. The larger the gradient magnitude, the clearer the image. After calculation, it is found that the gradient magnitude of a certain stationary frame is the largest among these stationary frames, and this frame is the potential key frame. In addition, if it is determined to be stationary for more than 15 consecutive frames, also extract one frame for processing to avoid missing the screen capture behavior with smaller movements. In this way, the selected potential key frames can more clearly present the possible abnormal operation moments, providing strong support for subsequent accurate judgment.

[0039] Process the potential key frames to generate three consecutive frames of the target screen object. Taking the determined potential key frame as the center, one frame is taken forward and one frame is taken backward to form three consecutive frames of the target screen object. For example, if the potential key frame is the Nth frame, then the (N - 1)th frame, the Nth frame, and the (N + 1)th frame are selected as the three consecutive frames of the target screen object. These three frames cover the pre, middle, and post states of the occurrence of the potential key action, providing more complete information for subsequent filtering processing based on the non-maximum suppression method and object detection. In the scenario of a mobile phone photographing a screen, the (N - 1)th frame may show the mobile phone approaching the screen, the Nth frame captures the moment of the mobile phone taking a photo (i.e., the potential key frame, at this time the mobile phone is relatively stationary and has high clarity), and the (N + 1)th frame presents the state after the mobile phone takes the photo. Through these three consecutive frames, the movement process and behavioral intention of the object can be analyzed more comprehensively, improving the detection accuracy of abnormal operations.

[0040] S104, filter the three consecutive frames of the target screen object based on the non-maximum suppression method to generate the target key frame image.

[0041] In one implementation, a preliminary category screening is performed on the three consecutive frames of the target screen object based on the non-maximum suppression method to generate a preliminary screening image set. In the actual application scenario of the public security bureau, the three consecutive frames of the target screen object are obtained through the previous key frame extraction step and may contain the detection box information of various objects. The non-maximum suppression method first performs a preliminary category screening on the detection boxes in these images. According to the preset category criteria, only the detection boxes of the categories of mobile phones and people are retained, and the detection boxes of other irrelevant objects such as documents, desks, and chairs are filtered out. For example, in the three consecutive frames obtained at a certain moment, after the object detection algorithm, there are initially detection boxes containing objects such as mobile phones, people, documents, and water cups. After the preliminary category screening, only the detection boxes of mobile phones and people are retained, and the detection boxes of documents and water cups are removed, thus obtaining the preliminary screening image set. This can reduce the amount of data for subsequent processing, improve the processing efficiency, and make the subsequent analysis more focused on the objects related to possible abnormal operations (such as a mobile phone photographing a screen, a person approaching the screen abnormally, etc.).

[0042] Process the same-class detection boxes in the initially screened image set to generate a first-optimized image set. In the initially screened image set, there may be multiple detection boxes for the same class of objects (such as mobile phones or people), which may be caused by issues such as the accuracy of the detection algorithm or different angles of the objects. In response to this situation, the non-maximum suppression (NMS) algorithm is used to process the same-class detection boxes. Calculate the intersection over union (IoU) between the same-class detection boxes. For detection boxes with an IoU greater than the set threshold (assumed to be 0.5), retain the detection box with a higher confidence level and remove the detection box with a lower confidence level. For example, in the initially screened image set, for the detection boxes of mobile phones, there are three detection boxes A, B, and C. If the IoU between A and B is greater than 0.5, and the confidence level of A is 0.9 and the confidence level of B is 0.7, then retain detection box A and remove detection box B. Through such processing, only the most representative detection box is retained for each same-class object, optimizing the image set, avoiding repeated detections and misjudgments, generating a first-optimized image set, and making the detection results more accurate and concise.

[0043] Process the different-class detection boxes in the first-optimized image set to generate a second-optimized image set. In the first-optimized image set, there are different-class detection boxes for mobile phones and people. At this time, compound conditions are used to further process these detection boxes. Condition 1 is that detection box A is completely contained in detection box B and 0 < IoU < 0.2, or the IoU between detection boxes A and B > 0.77 and the center point distance is less than dis (assumed dis is set to 50 pixels). When these compound conditions are met, delete the detection box with a lower confidence level. For example, in the first-optimized image set, there is a mobile phone detection box M and a person detection box N. If M is completely contained in N and their IoU is 0.1, and at the same time the confidence level of M is lower than that of N, then delete detection box M; or if the IoU between M and N > 0.77 and the distance between their center points is less than 50 pixels, also delete the detection box with a lower confidence level. In this way, further screen out the effective detection boxes, generate a second-optimized image set, improve the accuracy of the detection, and more accurately identify the objects related to abnormal operations.

[0044] Perform grouping processing on the detection boxes of personnel in the secondarily optimized image set to generate a target sequence and an intermediate image set. In the secondarily optimized image set, check whether there are detection boxes with the category of person. If so, divide all the detection boxes with the category of person into different groups according to the condition that the IoU in the same category is not equal to 0. Then, find the detection box with the highest confidence in each group to form the sequence R; if there are no detection boxes with the category of person in the remaining detection boxes, directly return all the remaining detection boxes as the intermediate image set. For example, in the secondarily optimized image set, there are five detection boxes of personnel P1, P2, P3, P4, and P5. Among them, the IoU of P1 and P2 is not equal to 0, the IoU of P3 and P4 is not equal to 0, and the IoU of P5 and other detection boxes is 0. Then divide P1 and P2 into one group, P3 and P4 into one group, and P5 into a separate group. Select the detection box with the highest confidence from each of these three groups, assume they are P1, P3, and P5, to form the sequence R. The remaining image set (including the detection boxes of mobile phones and other unselected detection boxes of personnel) is used as the intermediate image set. Such grouping processing helps to more accurately judge the relationship between the mobile phone and the personnel in the subsequent steps and determine whether there are abnormal operations.

[0045] Adjust the confidence of the detection boxes of mobile phones based on the target sequence to generate the target key-frame image. If the sequence R is not empty and there are detection boxes with the category of mobile phone in the intermediate image set, adjust the confidence of the detection boxes of mobile phones according to the adjustment rule. The adjustment rule is: the new confidence conf1 of each detection box with the category of mobile phone = λ * conf, where conf is the original confidence, and 0 < λ < 1 is a parameter related to (X, Y) and (X0, Y0). (X0, Y0) is the center point of the detection box with the category of mobile phone, and (X, Y) is the center point of the detection box in the sequence R that is closest to the center point (X0, Y0). When (X0, Y0) deviates further from (X, Y), λ is smaller; when (X0, Y0) deviates closer to (X, Y), λ is larger, and λ is more sensitive to the deviation in the X direction. If the sequence R is empty, return all the remaining detection boxes.

[0046] For example, in sequence R, there are detection boxes P1, P3, and P5, and in the intermediate image set, there is a mobile phone detection box Q. Calculate the distances between the center point (X0, Y0) of Q and the center points of each detection box in R. Assume that the distance is the closest to the center point (X, Y) of P1. If the deviation between (X0, Y0) and (X, Y) in the X direction is small and the deviation in the Y direction is large, according to the rule, λ is calculated to be 0.8, and the original confidence level of Q is 0.9. Then the adjusted confidence level conf1 = 0.8 * 0.9 = 0.72. Finally, based on the adjusted confidence level threshold conf1, the box of the mobile phone category is retained or deleted to generate the target key frame image. In this process, if the distance between the center point of the mobile phone detection box and the center point of the person detection box is too large, the adjusted confidence level will decrease, which may lead to the deletion of the mobile phone detection box because this situation may be a false detection; while when the distance is relatively close, the confidence level adjustment is small, and the detection box is retained to more accurately determine the behavior of the mobile phone taking pictures of the screen, generating a more accurate target key frame image for subsequent analysis.

[0047] S105. Process the target key frame image to generate the real-time motion state and screen-peeping risk characteristics of the target screen object.

[0048] In one implementation, feature extraction is performed based on the target key frame image to generate a set of feature vectors. The target key frame image is obtained after a series of processes and contains images related to possible abnormal operations (such as a mobile phone taking pictures of the screen). A dedicated feature extraction algorithm, such as the HOG (Histogram of Oriented Gradients) algorithm, is used to process the target key frame image. This algorithm calculates the histogram of gradient directions in local regions of the image to describe the texture and shape features of the image. For example, in a target key frame image containing a person and a mobile phone, the HOG algorithm divides the image into multiple small cells, calculates the gradient direction and amplitude of the pixels in each cell, and then statistically analyzes the histogram of these gradient directions. For the mobile phone part, the HOG algorithm can extract unique edge shapes, textures, etc. of the mobile phone, and combine these features into a feature vector; for the person part, it can also extract feature vectors representing the human body contour, posture, etc. Integrating all these feature vectors together forms the set of feature vectors of the target key frame image. These feature vectors provide a data basis for subsequent analysis of the motion and behavior of objects.

[0049] Compare the feature vector sets of adjacent key-frame images to determine the position changes of the same object in different frames and generate object motion trajectory information. After obtaining the feature vector sets of adjacent key-frame images respectively, use a feature matching algorithm (such as a matching algorithm based on Euclidean distance) to compare these sets to determine the position changes of the same object in different frames. Euclidean distance can measure the similarity between two feature vectors. The smaller the distance, the more similar the features represented by the two vectors, that is, they are very likely to belong to the same object. Suppose the feature vector of the mobile phone in the previous key-frame image is V1, and the feature vector of an object in the current frame is V2. By calculating the Euclidean distance between V1 and V2, if the distance is less than a set threshold (for example, the threshold is set to 50, and this value is determined according to experiments and experience), then it is considered that these two feature vectors represent the same mobile phone. Then, according to the position information of the feature vector in the image (such as the center coordinates of the object corresponding to the feature vector), determine the position changes of the mobile phone in different frames. For example, the center coordinates of the mobile phone in the previous frame are (100, 100), and the center coordinates of the object determined to be the same mobile phone through feature matching in the current frame become (120, 110). From this, the displacement of the mobile phone between these two frames can be calculated as (20, 10). As time goes by, continuously compare the feature vectors of the same object in multiple groups of adjacent key-frame images, and the motion trajectory information of the object can be generated. For example, if the mobile phone gradually approaches the screen from a certain direction, its motion trajectory can be described by a series of coordinate changes.

[0050] Compare the generated object motion trajectory information with the preset normal behavior patterns and abnormal behavior patterns to generate a motion state evaluation result. The preset normal behavior patterns and abnormal behavior patterns are set based on a large amount of historical data and business experience. Normal behavior patterns may include the motion trajectories of hands when a person is operating a computer normally, the moving range of a mobile phone in normal usage scenarios, etc.; abnormal behavior patterns target actions such as a mobile phone taking pictures of the screen, such as features like the mobile phone quickly approaching the screen and staying briefly in front of the screen. Compare the generated object motion trajectory information with these preset patterns. For example, if the motion trajectory of the mobile phone shows that it quickly moves from a position far from the screen to in front of the screen within a short period of time and stays still in front of the screen for a while, which matches the preset abnormal behavior pattern of the mobile phone taking pictures of the screen, then it is determined that this motion state is abnormal; on the contrary, if the motion trajectory of the mobile phone conforms to the normal behavior pattern of a person holding the mobile phone for other operations (such as viewing information), that is, the motion range and speed are within the normal range, then it is determined to be in a normal motion state. Through this comparison process, a motion state evaluation result is generated to clarify whether the current object's motion belongs to normal or abnormal behavior.

[0051] Process the motion state evaluation result based on the environmental information of the target key-frame image to generate the screen-peeping risk feature. The environmental information of the target key-frame image includes the screen display content (whether it is sensitive information), the distribution of surrounding people (analyzing the position and number of people through multiple frames of images obtained by the camera), etc. Combining the motion state evaluation result, if the motion state is determined to be abnormal, and the screen display content is sensitive information (such as judging that the confidential case materials are displayed on the screen through OCR text recognition technology), and there are other people approaching the screen around, then it can be determined that there is a relatively high screen-peeping risk. For example, the key-frame image shows a person holding a mobile phone approaching the screen, the motion state evaluation is abnormal, and the screen shows unpublicized case information, and there are other participants around. At this time, according to these environmental information and the motion state evaluation result, screen-peeping risk features are generated, such as the risk level (high, medium, low), risk type (such as external person screen-peeping, internal person abnormal operation, etc.), so as to more accurately evaluate the security risk in the current scenario.

[0052] Process the motion state evaluation result and the screen-peeping risk feature to generate a real-time motion state and screen-peeping risk feature report of the target screen object. Integrate the motion state evaluation result and the screen-peeping risk feature to form a detailed report. The report content includes the current motion state (normal / abnormal) of the object, the motion trajectory, the speed change situation, as well as the level, type and relevant basis of the screen-peeping risk. For example, record in a report: "The motion state of the target screen object (mobile phone) is abnormal, and the motion trajectory shows that it moves quickly from the left side of the meeting room to in front of the screen and stays, and the speed gradually decreases when approaching the screen. The screen-peeping risk level is high, and the risk type is that an internal person may take pictures of confidential information, based on the fact that the screen shows confidential case materials and there are other people around." Such a report provides comprehensive and accurate information support for subsequent abnormal operations and screen-peeping warnings, facilitating the system to respond in a timely manner, such as issuing an alarm, notifying the administrator, etc., and also helps to analyze and process the event subsequently.

[0053] S106, process the real-time motion state of the target screen object based on the preset dynamic threshold and the screen-peeping risk feature to generate a stable frame image.

[0054] In one implementation, the real-time motion state of the target screen object is processed based on a preset dynamic threshold to generate a preliminary screening result. The real-time motion state information of the target screen object (such as the object's motion speed, acceleration, motion direction, etc.) is obtained through previous steps. The preset dynamic threshold is set according to a large amount of historical data, normal behavior patterns in different scenarios, and security policies. For example, based on the analysis of the daily office scenarios of the public security bureau, it is determined that when a person operates a mobile phone under normal circumstances, the motion speed of the mobile phone in front of the screen generally does not exceed 50 pixels per second, and the acceleration fluctuates within a certain range. The real-time motion state of the target screen object is compared with the preset dynamic threshold. If the motion speed of the mobile phone exceeds 50 pixels per second, or the acceleration shows abnormal changes (such as a sharp increase or decrease in speed within a short period of time), then this motion state is marked as abnormal; if it is within the threshold range, it is marked as normal. In this way, a preliminary screening is performed on the motion states of all detected objects, and the motion states are divided into two categories: normal and abnormal, to generate a preliminary screening result. For example, in a video, it is detected that the motion speed of a mobile phone reaches 80 pixels per second at a certain moment, exceeding the preset threshold, and this motion state is marked as abnormal in the preliminary screening result.

[0055] Based on the screen-peeping risk characteristics, the preliminary screening result is processed to generate risk-associated motion state information. The screen-peeping risk characteristics include factors such as the distance between the person and the screen in the target key frame image, the person's posture (such as whether the person is sideways blocking the screen, whether there are abnormal approaching actions, etc.), and the sensitivity of the screen display content (judging whether confidential information is displayed through text recognition, etc.). Combining the preliminary screening result, if the motion state of an object in the preliminary screening result is marked as abnormal, and at this time the screen-peeping risk characteristics show a high risk, such as the person is too close to the screen and the screen displays confidential information, and the person's posture shows obvious peeping actions, then this motion state is marked as a high-risk motion state; if the motion is abnormal but the screen-peeping risk is low, for example, although the person has abnormal actions but the screen displays ordinary office documents, it is marked as a low-risk motion state; for the normal motion state, it is also recorded in combination with the screen-peeping risk characteristics, such as normal motion and the screen display content is not sensitive, and the surrounding personnel distribution is normal, then it is marked as a normal low-risk motion state. Through such processing, motion state information including a risk level is generated, that is, risk-associated motion state information. For example, in a meeting scenario, the preliminary screening result shows that the motion state of the mobile phone is abnormal. At the same time, by analyzing the screen-peeping risk characteristics, it is found that the mobile phone user is very close to the screen, the screen shows unpublicized case materials, and the person leans forward and deliberately blocks the screen. After comprehensive judgment, this motion state is marked as a high-risk motion state.

[0056] Screen the target key-frame images based on risk-related motion state information to generate a set of potentially stable frames. For the key-frame images corresponding to the high-risk motion states, retain them preferentially because these images may record potential abnormal operations or screen-peeping behaviors; for the key-frame images corresponding to the low-risk motion states, selectively retain them according to certain rules, such as retaining them with a certain probability, or deciding whether to retain them based on their association degree with the key-frame images of the high-risk motion states; for the key-frame images with normal motion states and low screen-peeping risks, only retain a small number of representative images. For example, in a series of key-frame images, 10 frames are related to high-risk motion states, 30 frames are related to low-risk motion states, and 50 frames are related to normal motion states. All 10 frames of images with high-risk motion states are retained; for the 30 frames of images with low-risk motion states, randomly retain 15 frames according to the set probability (such as 50%); for the 50 frames of images with normal motion states, select 5 representative images to retain. Through this screening method, a set of potentially stable frames is generated. The images in this set contain both the key frames that may pose risks and some key frames in the normal state, which are used to generate stable-frame images subsequently to ensure that important information is not missed.

[0057] Perform clarity and stability evaluation processing on the images in the set of potentially stable frames to generate candidate stable-frame images. Evaluate the clarity and stability of each image in the set of potentially stable frames. Use a clarity detection algorithm (such as calculating the gradient magnitude of the image by the Laplace algorithm to measure clarity) to evaluate the clarity of the image, and at the same time evaluate the stability of the image by analyzing the differences between adjacent frames (such as the pixel change rate, the smoothness of the object position change, etc.). Select the images with higher clarity and better stability as candidate stable-frame images. For example, there are 10 images in the set of potentially stable frames. Apply the Laplace algorithm to each image to calculate its gradient magnitude. The larger the gradient magnitude, the clearer the image. At the same time, calculate the pixel change rate between adjacent frames. The smaller the pixel change rate, the more stable the image. After evaluation, it is found that the clarity and stability indexes of 3 of these images are relatively high, and these 3 images are used as candidate stable-frame images. This can ensure that the quality of the finally generated stable-frame images is relatively high, which is more conducive to subsequent analysis and judgment.

[0058] Process the stable frame candidate images to generate stable frame images. In the stable frame candidate images, select the optimal image as the final stable frame image according to certain rules. The selection can be based on the risk level of the images. Preferentially select the images corresponding to the high risk level to ensure key monitoring and analysis of possible abnormal situations. If there are multiple candidate images with high risk levels, further compare their clarity and stability, and select the image with the optimal comprehensive index. If there are no candidate images with high risk levels, select from the candidate images with low risk levels. For example, among 3 stable frame candidate images, 1 is of high risk level and the other 2 are of low risk level. At this time, preferentially select the image with high risk level as the stable frame image. If all 3 images are of low risk level, compare their clarity and stability indexes, and select the image with the highest comprehensive score as the stable frame image. In this way, determine a unique image from the stable frame candidate images as the stable frame image, which is used to generate abnormal operation and screen peeping warning information subsequently, providing a reliable basis for the decision-making and processing of the system.

[0059] S107, process the stable frame images to generate abnormal operation and screen peeping warning information and optimize energy consumption management.

[0060] In one implementation, perform object detection processing on the stable frame images to generate detection result information. The stable frame images are obtained after a series of screening and processing, and contain images that may have abnormal operations or screen peeping risks. Use the trained object detection model (such as the improved YOLO model trained previously) to analyze the stable frame images. The model will identify the objects in the images, determine the categories of the objects (such as mobile phones, people), positions (represented by the coordinates of the rectangular frames), and confidence levels (indicating the degree of certainty of the model about the detection results), etc. For example, in a stable frame image, the object detection model identifies that there is an object as a mobile phone, and its position is within the rectangular frame with image coordinates (100, 150)-(200, 250), and the confidence level is 0.9. At the same time, two people are also detected in the image, located at different positions, and the corresponding coordinates and confidence level information are also given. These recognition results constitute the detection result information, providing the basic data for subsequent judgment of whether there are abnormal operations.

[0061] Compare and process the detection result information based on preset rules to generate an abnormal operation judgment result. The preset rules are formulated according to business requirements and security policies and are used to determine whether the detected object behavior belongs to an abnormal operation. For example, the rule may stipulate that if a mobile phone stays within a specific range in front of the screen (such as within 50 pixels from the screen edge) for more than 2 seconds, and the angle between the mobile phone and the screen is within a certain range (such as the angle with the screen plane is less than 30 degrees), it is determined that there may be an abnormal operation of taking a picture of the screen. Compare the result information obtained from object detection with the preset rules. If the information such as the position and stay time of the mobile phone is detected to conform to the above rules, it is determined that there is an abnormal operation; otherwise, it is determined to be a normal operation. Suppose in the detection result information, the mobile phone stays at 30 pixels in front of the screen for 3 seconds, and the angle with the screen plane is 20 degrees. By comparing with the preset rules, the generated abnormal operation judgment result is "There is an abnormal operation, and it may be taking a picture of the screen."

[0062] Process the screen peeping risk features and the abnormal operation judgment result to generate a screen peeping warning information. The screen peeping risk features include the sensitivity of the screen display content (judging whether confidential information is displayed on the screen through OCR text recognition technology), the distribution of surrounding personnel (analyzing the position and number of personnel through multiple frames of images obtained by the camera), etc. Combining the abnormal operation judgment result, if it is judged that there is an abnormal operation, and the screen display content is sensitive information (such as identifying that confidential case materials are displayed on the screen), and there are other people approaching the screen around, then generate the corresponding screen peeping warning information.

[0063] For example, the abnormal operation judgment result is that there is an abnormal operation, and through OCR technology, it is identified that the screen displays confidential information, and at the same time, it is analyzed from the camera image that there are other people close to the screen. At this time, the generated screen peeping warning information may be "High-risk screen peeping warning, detecting possible behavior of taking pictures of confidential information, and there are other people around", and different warning levels (such as high, medium, low) are set according to the risk level, so that the system can take corresponding measures.

[0064] Get the real-time status information and preset energy consumption management information of the target device. The target device refers to a computer device with an anti-screen shooting and anti-peeping system installed. The real-time status information includes the device's power, CPU usage, memory usage, etc. The preset energy consumption management information is a strategy formulated according to the device usage scenario and energy-saving requirements. For example, when the power is less than 20% and there is no operation for a period of time, the screen brightness is automatically reduced; after 6 o'clock in the evening, if no one is detected in front of the screen, the computer is forced to shut down automatically after one hour. The real-time status information of the target device is obtained through the device's built-in sensors and system monitoring tools. For example, the current power of a computer device is 30%, the CPU usage is 20%, and the memory usage is 50%. At the same time, read the preset energy consumption management information from the system configuration file, such as the above-mentioned power and no-person detection shutdown strategy.

[0065] Based on the real-time status information and screen peeping warning information of the target device, the preset energy consumption management information is processed to generate energy consumption management instructions, wherein the energy consumption management instructions are used to control the target device to operate. The preset energy consumption management information is adjusted and the energy consumption management instructions are generated by comprehensively considering the real-time status information and screen peeping warning information of the target device. If the screen peeping warning information is detected, in order to ensure information security, even if the device has sufficient power, energy-saving operations such as automatic shutdown or lowering the screen brightness may be suspended to ensure that relevant personnel can handle abnormal situations in a timely manner; if the screen peeping warning information is not detected, but the device power is low and meets the preset energy-saving conditions, the corresponding energy consumption management instructions are generated according to the preset strategy.

[0066] For example, if a screen peeking warning message is received, the energy consumption management instruction generated may be "maintain the current operating state of the device, prohibit automatic shutdown and reduce screen brightness operations"; if no screen peeking warning message is received, and the current time is after 6 o'clock in the evening, there is no one in front of the screen and no one has been detected for one hour, and the device power is higher than 10%, then the energy consumption management instruction generated is "automatic shutdown" and sent to the target device to control the device to perform corresponding operations, thereby realizing the coordinated work of energy consumption management and security monitoring.

[0067] The present application also provides an implementation. Relying on the computing power of the front-end computer, the application realizes a system that integrates functions of anti-screen capture, anti-peeping, and energy consumption reduction. Without the support of the computing power of the back-end server cluster, through optimizing algorithms and multi-threading technologies, the system is ensured to operate efficiently, protecting user information security and reducing energy consumption. First, various types of pictures containing people, mobile phones, and items that are easily misidentified as mobile phones (such as glasses, epaulets, pens, etc.) in various scenarios are widely collected. The sources of these pictures are as diverse as possible, covering different lighting conditions, shooting angles, personnel postures, and complex backgrounds. For example, in the office scenario, images of employees using mobile phones or placing related items at different positions and times are collected; in the conference room scenario, pictures of personnel interacting with mobile phones during the meeting are captured. For the collected pictures, professional image annotation tools are used to accurately annotate the targets therein. For each person, their outline, position, and key feature points are annotated; for mobile phones, information such as the overall outline and model of the mobile phone is annotated; for items that are easily misidentified, their outlines and categories are also annotated in detail. The annotation information will be saved in a specific data format, such as XML or JSON format, for convenient subsequent training use. The annotated pictures are sorted into a training data set. To improve the generalization ability of the model, the data is reasonably divided. A part is used as the training set for model training, and another part is used as the validation set for evaluating the performance of the model. During the division process, it is ensured that various types of data have a suitable proportional distribution in the training set and the validation set, avoiding the situation of data imbalance.

[0068] Using a pre-trained model of the object detection algorithm as the basis, it is retrained with the newly constructed data set. During the training process, the parameters of the model are adjusted to enable the model to learn the features and rules in the data. For example, for a convolutional neural network (CNN), parameters such as the weights and biases of the convolutional layer are adjusted so that the model can more accurately identify mobile phones and distinguish items that are easily misidentified. By increasing the data of common misdetection categories, when the model faces complex scenarios, it can better distinguish real mobile phone targets from other interfering items, thereby effectively improving the accuracy of mobile phone detection and providing a reliable model support for subsequent object detection tasks.

[0069] The video image frames in front of the screen are collected in real time using a computer camera. The installation position and angle of the camera are carefully adjusted to ensure that the activities of the people in front of the screen and the mobile phone usage can be clearly captured. To ensure the continuity and stability of the collection, an appropriate frame rate is set, such as 25 frames per second or 30 frames per second. In this way, sufficient image information can be obtained for subsequent analysis without occupying too many resources. The video image frames are continuously collected and used as the data to be analyzed. These data will be stored in the local cache in a timely manner and wait for subsequent processing. During the data acquisition process, through multi-threading technology or asynchronous processing mechanisms, it is ensured that the data acquisition does not affect the normal operation of other system functions, and the fluency and stability of the system are guaranteed. A series of preprocessing operations are performed on the collected image frames. First, format conversion is carried out, converting the original image format (such as YUV format) collected by the camera into a format suitable for subsequent processing (such as RGB format). Then, noise reduction processing is carried out, and algorithms such as mean filtering and Gaussian filtering are used to remove the noise in the image, improving the clarity and quality of the image. In addition, the image is subjected to size normalization processing, adjusting images of different sizes to a unified size, which is convenient for the input and processing of subsequent models. Through these preprocessing operations, the quality of the image is enhanced, the influence of noise and interference factors on subsequent object detection is reduced, and a better image data basis is provided for accurate object detection.

[0070] The three-frame difference method determines the motion state of an object based on the pixel differences between three adjacent frames of images. In the video stream, three consecutive frames of images are sequentially obtained and denoted as , , . Calculate the pixel differences between and as well as between and to obtain two difference images and . By setting a connected region or area threshold, the pixel change regions in the difference images are analyzed. If the pixel change amount in a certain region exceeds the threshold and the connectivity characteristics (such as the connection relationship between pixels) of this region meet the preset conditions, it is considered that there is object movement in this region. When movement is detected in a certain region, this region is marked as "1"; if the pixel changes in all regions do not exceed the threshold, indicating that the object is in a stationary state, it is marked as "0". This marking method quantifies the movement situation of the object in each frame of the image and provides a clear data basis for subsequent analysis. For example, in an actual application scenario, when someone holds a mobile phone close to the screen to prepare for shooting, the movement of the mobile phone and the person will cause large pixel changes in the corresponding regions of the image, which are thus marked as "1"; while in the case of no operation, the objects in front of the screen remain stationary, and the markings of each region of the image are "0".

[0071] Analyze the labeling results of consecutive multiple frames to find cases that match specific patterns (such as 10, 110, 1100, 11000, etc.). These patterns are summarized based on the characteristics of mobile phone screen capture behavior. Taking the "10" pattern as an example, it means that the object first moves (labeled as 1), and then stops (labeled as 0), which conforms to the behavior characteristics that the mobile phone first approaches the screen (moves) and then pauses briefly (stops) at the moment of shooting during normal shooting. The "110" pattern indicates that the object has a continuous movement first and then stops, which is applicable to the situation where the mobile phone has some adjustment actions during the process of approaching the screen.

[0072] When multiple frames that match a specific pattern are recognized, key frames need to be selected from these frames. The Laplace algorithm is used to evaluate the sharpness of the static frames (frames labeled as "0"). The Laplace algorithm highlights the edges and detail information in the image by calculating the second derivative of the image. The larger the Laplace value of the image, the higher the sharpness of the image. Among the static frames that match the pattern, select the frame with the largest Laplace value as the key frame. For example, in a sequence of frames that match the "110" pattern, there are multiple static frames. Calculate the Laplace values of each static frame through the Laplace algorithm and select the frame with the highest Laplace value as the key frame. The key frame selected in this way contains richer detail information, which is beneficial for subsequent object detection algorithms to more accurately identify target objects such as mobile phones.

[0073] In actual situations, there is the behavior of secretly taking pictures with a mobile phone moving slowly deliberately. In this case, it may cause multiple consecutive frames to be judged as static (labeled as "0"). If only relying on the previous pattern matching method, these screen capture behaviors with small movements may be missed. Therefore, a continuous static frame processing mechanism is set up. When the system detects that more than 15 consecutive frames are judged as static, even if there is no situation that matches a specific pattern, one of the frames will be extracted as the key frame. The extraction of this frame can be randomly selected, or according to certain rules, such as selecting the frame in the middle of the sequence or the latest frame according to the timestamp. In this way, it is ensured that even in the case of secretly taking pictures with a mobile phone moving slowly, key frames can be obtained in time for object detection, improving the comprehensiveness and accuracy of system detection and effectively avoiding the detection omission problem caused by the inconspicuous shooting action.

[0074] Adopt Method 2 for multi-threaded operation. The frame reading part is responsible for a separate thread to continuously collect image frames; the object detection part uses a single thread to ensure the stability of the detection process; the background interaction part uses a separate thread to process in parallel with other threads. The key frame extraction thread and the object detection thread transfer key frame data through a specific communication mechanism, and use the trained model to perform object detection on the key frame, identify the detection frames containing targets such as people and mobile phones, and obtain the detection results.

[0075] The improved non-maximum suppression method is used to filter out detection frames that are not mobile phones and people, thus narrowing the processing scope. In the same type of detection frames, the original NMS algorithm is used to remove the detection frames that meet IoU>eta and have low confidence; in different types of detection frames, according to the composite condition (detection frame A is completely contained in detection frame B and 0 <IoU<0.2,或者检测框A和B的IoU> 0.77 and the center point distance is less than dis) to delete low confidence detection frames. If there are detection frames of human category in the remaining detection frames, group them according to the condition that IoU is not equal to 0, and find the detection frames with the highest confidence in each group to form a sequence R. If sequence R is not empty and there are detection frames of mobile phone category in the remaining detection frames, adjust the confidence conf1 of the mobile phone detection frame according to the distance between the mobile phone detection frame and the center point of the nearest detection frame in sequence R (conf1=λ*conf, λ is related to the center point distance). According to the adjusted confidence threshold conf1, keep or delete the mobile phone detection frame to reduce the false detection rate and improve the detection accuracy.

[0076] If a mobile phone detection frame is detected, the system automatically captures the screen image and camera photo, and displays a screen capture warning message on the screen. When it is detected that there is no one in front of the screen, the computer automatically locks the screen (currently replaced by a pop-up window covering the screen) to prevent the computer desktop information from being exposed and unauthorized use. After 6 o'clock in the evening, continue to detect whether there is anyone in front of the screen. If no one is detected for an hour, and more than half an hour has passed since the last detection of a person, the computer automatically shuts down. Before shutting down, automatically save documents of office software such as office and WPS to reduce the risk of hardware failure.

[0077] In view of the differences in terminal computer configurations, a lightweight version and a normal version of the system are provided. Terminals with poor configurations (such as 4G memory) install the lightweight version, and terminals with better configurations install the normal version to ensure that the system runs stably in different hardware environments. The system runs on local computers in a decentralized manner, without the need for GPUs and server clusters. When the lightweight version runs on a computer with 2.4GHz, 4G memory, and a 32-bit operating system configuration, the CPU consumption is within an average of less than 10%, and the memory consumption is within an average of less than 100M, which does not affect other office business. Using multi-threaded concurrent execution technology, functions such as sending heartbeat packets, taking screenshots, and uploading pictures are separated into separate threads, and processed in parallel with core functions such as target detection to avoid blocking the main process and improve system response speed.

[0078] In addition, an uninstall password can be added to prevent unauthorized uninstallation; set the system to start automatically when the computer boots up and send heartbeat packets to the local server, and the administrator can master the installation status through IP comparison; set up a daemon process to prevent the process from being manually terminated by the user. When no one is detected, the screen will be automatically locked or a prompt message will pop up to cover the screen, preventing the user from deliberately blocking the camera or adjusting the camera direction. After multiple rounds of optimization and testing, the system is compatible with 32-bit, 64-bit and above operating systems, and does not require additional operations by the user, covering most office computers. When a specific system requests to use the camera, the system releases the camera's right of use; after the use is over, the right of use is immediately recovered to ensure the security and rationality of the camera's use.

[0079] In this application, the server obtains the target screen information, the preset model, the training sample set and the validation set. By processing the sample set, a target multi-threaded concurrent image detection model is trained, laying a foundation for subsequent detection. Then, the key frame extraction algorithm is used to process the target screen information. First, the model preprocesses to generate image data blocks, and then through processing such as three-frame difference method, continuous multi-frame data is generated. The object motion state information is analyzed, potential key frames are screened out, and then continuous three-frame images of the target screen object are generated. After that, the non-maximum suppression method is used to filter these images to obtain the target key frame images.

[0080] Feature extraction is performed on the target key frame images. By comparing the feature vectors of adjacent frames, the object motion trajectory is determined. The motion state is evaluated in combination with the preset behavior pattern. At the same time, the screen peeping risk features are generated based on the environmental information, forming a real-time motion state and screen peeping risk feature report. Based on this, the preset dynamic threshold and the screen peeping risk features are used to process the real-time motion state, and the stable frame images are screened out. Finally, target detection is performed on the stable frame images, and it is judged whether there are abnormal operations according to the preset rules. Combining the screen peeping risk features, warning information is generated. At the same time, the real-time state of the device and the preset energy consumption management information are obtained, and an energy consumption management instruction is comprehensively generated to control the device operation, realizing early warning of abnormal operations and energy consumption optimization. This technology not only ensures the information security of the public security bureau, but also responds to the requirement of low energy consumption and improves the intelligent level of device management. Centering on the needs of the public security bureau, integrating technologies such as multi-threaded concurrency and key frame extraction, it realizes the functions of anti-screen capture, anti-peeping and energy consumption reduction, ensuring information security and reducing energy consumption.

[0081] In one implementation, as Figure 2 shown, this application also provides a low-energy intelligent anti-screen capture device based on multi-threaded concurrency and key frame extraction, including:

[0082] An acquisition module 201, configured to acquire target screen information, a preset multi-threaded concurrent image detection model, a training sample set and a validation set;

[0083] The processing module 202 is configured to train a preset multi-threaded concurrent image detection model based on a training sample set and a validation set to generate a target multi-threaded concurrent image detection model; extract key frames from the target screen information based on the target multi-threaded concurrent image detection model and a key frame extraction algorithm to generate three consecutive frames of images of the target screen object; filter the three consecutive frames of images of the target screen object based on a non-maximum suppression method to generate target key frame images; process the target key frame images to generate the real-time motion state and screen peeping risk features of the target screen object; process the real-time motion state of the target screen object based on a preset dynamic threshold and the screen peeping risk features to generate stable frame images; process the stable frame images to generate abnormal operation and screen peeping warning information and optimize energy consumption management.

[0084] The computer-readable storage medium provided in the above embodiments of the present application and the low-energy intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0085] Each embodiment in the present application is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the method, electronic device, electronic equipment, and readable storage medium for evaluating the low-energy intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction, since they are basically similar to the embodiments of the above-mentioned low-energy intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction, the description is relatively simple, and the relevant parts can be referred to the partial description of the embodiments of the above-mentioned low-energy intelligent anti-screen capture method based on multi-threaded concurrency and key frame extraction.

Claims

1. A low-energy intelligent anti-screening method based on multi-threaded concurrency and key frame extraction, characterized in that: include: Obtain target screen information, preset multi-threaded concurrent image detection model, training sample set and verification set; Based on the training sample set and the verification set, the preset multi-threaded concurrent image detection model is trained and processed to generate a target multi-threaded concurrent image detection model; Based on the target multi-threaded concurrent image detection model and key frame extraction algorithm, the target screen information is extracted with key frames to generate three consecutive frames of images of the target screen object; Based on the non-maximum suppression method, three consecutive frames of images of the target screen object are filtered to generate a target key frame image; Process the target key frame image to generate the real-time motion state and peeping risk characteristics of the target screen object; The real-time motion state of the target screen object is processed based on the preset dynamic threshold and the peeping risk characteristics to generate a stable frame image; Process stable frame images, generate abnormal operation and screen peeking warning information, and optimize energy consumption management.

2. The method according to claim 1, characterized in that The preset multi-threaded concurrent image detection model is trained based on the training sample set and the verification set to generate a target multi-threaded concurrent image detection model, including: Count the number of image features in the training sample set and generate the sampling ratio based on the balance of feature distribution: Based on the sampling ratio, the training sample set is subjected to multimodal stratified sampling processing to generate a preset number of sampling feature combinations: Based on statistical tests of any key image feature combined with each sampling feature, the data is divided into a group with good detection effect and a group with poor detection effect, where each group contains a preset number of data samples and at least one sample has error identification information: The preset multi-threaded concurrent image detection model is iteratively trained based on the good detection effect group and the poor detection effect group. The model parameters are optimized through cross-validation to generate the trained model and performance indicators: If the erroneous identification information in the training results is recognized by the model as a key factor affecting the evaluation of the detection effect, the trained model will be used as the target multi-threaded concurrent image detection model.

3. The method according to claim 1, characterized in that Based on the target multi-threaded concurrent image detection model and key frame extraction algorithm, the target screen information is extracted by key frame, and three consecutive frames of images of the target screen object are generated, including: Preprocess the target screen information based on the target multi-threaded concurrent image detection model to generate several image data blocks; Processing several image data blocks based on key frame extraction algorithm to generate continuous multi-frame data; Processing multiple frames of continuous data to generate object motion state information; Processing the motion state of the object to generate a potential key frame, wherein the potential key frame is a frame in a static state with the highest definition; The potential key frames are processed to generate three consecutive frames of images of the target screen object.

4. The method according to claim 3, characterized in that Based on the non-maximum suppression method, three consecutive frames of images of the target screen object are filtered to generate a target key frame image, including: Based on the non-maximum suppression method, three consecutive frames of images of the target screen object are preliminarily screened for categories to generate a preliminarily screened image set; Process similar detection frames in the preliminary screened image set to generate an optimized image set; Process the different types of detection frames in the first optimized image set to generate a second optimized image set; Performing person detection frame grouping processing on the secondary optimized image set to generate a target sequence and an intermediate image set; The confidence of the mobile phone detection frame is adjusted based on the target sequence to generate the target key frame image.

5. The method according to claim 1, characterized in that The target key frame image is processed to generate the real-time motion status of the target screen object and the risk characteristics of screen peeping, including: Perform feature extraction based on the target key frame image and generate a feature vector set; Compare the feature vector sets of adjacent key frame images to determine the position change of the same object in different frames and generate the object motion trajectory information; Compare the generated object motion trajectory information with the preset normal behavior pattern and abnormal behavior pattern to generate a motion state assessment result; The motion state assessment results are processed based on the environmental information of the target key frame image to generate the screen peeping risk feature; The motion state assessment results and the screen peeping risk characteristics are processed to generate a real-time motion state and screen peeping risk characteristic report of the target screen object.

6. The method according to claim 5, characterized in that The real-time motion state of the target screen object is processed based on the preset dynamic threshold and the peeping risk characteristics to generate a stable frame image, including: The real-time motion state of the target screen object is processed based on the preset dynamic threshold to generate preliminary screening results: Process the preliminary screening results based on the risk characteristics of peeping screens to generate risk-related motion status information; The target key frame images are screened based on the risk-related motion state information to generate a set of potential stable frames: Perform clarity and stability evaluation on the images in the potential stable frame set to generate stable frame candidate images: The stable frame candidate images are processed to generate stable frame images.

7. The method according to claim 6, characterized in that Process stable frame images, generate abnormal operation and screen peeking warning information, and optimize energy consumption management, including: Performing target detection processing on the stable frame image to generate detection result information; Compare and process the detection result information based on preset rules to generate abnormal operation judgment results; Process the screen peeping risk characteristics and abnormal operation judgment results to generate screen peeping warning information; Obtain real-time status information and preset energy consumption management information of target devices; The preset energy consumption management information is processed based on the real-time status information and the screen peeking warning information of the target device to generate an energy consumption management instruction, wherein the energy consumption management instruction is used to control the target device to operate.

8. A low-energy intelligent anti-filming screen device based on multi-threaded concurrency and key frame extraction, characterized in that: The device comprises: An acquisition module is used to obtain target screen information, preset multi-threaded concurrent image detection models, training sample sets and verification sets; The processing module is used to train and process the preset multi-threaded concurrent image detection model based on the training sample set and the verification set to generate a target multi-threaded concurrent image detection model; perform key frame extraction on the target screen information based on the target multi-threaded concurrent image detection model and the key frame extraction algorithm to generate three consecutive frames of images of the target screen object; filter and process the three consecutive frames of images of the target screen object based on the non-maximum suppression method to generate the target key frame image; process the target key frame image to generate the real-time motion state and screen peeping risk characteristics of the target screen object; process the real-time motion state of the target screen object based on the preset dynamic threshold and the screen peeping risk characteristics to generate a stable frame image; process the stable frame image to generate abnormal operation and screen peeping warning information and optimize energy consumption management.

9. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the low-energy intelligent anti-screen shooting method based on multi-thread concurrency and key frame extraction as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the low-energy intelligent anti-screen shooting method based on multi-thread concurrency and key frame extraction described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Indoor monitoring video key frame real-time extraction method based on machine learning

    CN110096945A

  • System for preventing full-intelligent computer display screen from being illegally shot by mobile phone

    CN110443136A

  • Method and system for quickly extracting key frame based on image features

    CN111597911A

  • Anti-candid image processing method and device, terminal and storage medium

    CN111711794A

  • Shooting behavior detection method and device, equipment and storage medium

    CN114743264A