Screen privacy dynamic protection method and system, terminal and medium

By comprehensively assessing personnel threats and interface sensitivity, and adopting a focus-following sanitization strategy, screen protection is automatically triggered, solving the problems of false protection and slow response in traditional methods, and achieving efficient data security protection and business continuity.

CN121765780APending Publication Date: 2026-03-31INSPUR FINANCIAL INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

When preventing malicious eavesdropping, existing public terminals often struggle to distinguish between operators and bystanders using traditional methods, leading to false alarms or slow responses that impact business processing efficiency and user experience.

Method used

By comprehensively assessing personnel threats, user operations, and interface content sensitivity, a focus-following purification strategy and business context awareness are adopted to automatically trigger protective measures, including overlaying a dynamic visual interference layer on the screen and generating a local purification area to ensure that the operator's line of sight is clearly visible.

Benefits of technology

Effectively prevent critical information from being maliciously spied on, improve data security protection level, ensure the smoothness and accuracy of business processing, achieve dynamic matching between protection strength and business risks, and balance security and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765780A_ABST
    Figure CN121765780A_ABST
Patent Text Reader

Abstract

The invention relates to the field of privacy protection of electronic equipment, and particularly provides a screen privacy dynamic protection method and system, a terminal and a medium, and the method comprises the steps: obtaining environment perception data, user state data and service context data, and determining a current comprehensive privacy risk level according to the obtained data; judging whether the comprehensive privacy risk level is higher than a preset risk level threshold value or not; if yes, starting a focus following purification strategy, and controlling a graphic rendering engine to execute privacy protection rendering processing, including: superposing a dynamic visual interference layer on original content of a screen, generating a local purification area on the dynamic visual interference layer based on user real-time sight focus coordinates in acquired user state data, and performing privacy protection rendering processing on the local purification area; a screen content area corresponding to the real-time sight focus coordinate of the user is kept, and other screen contents form visual interference. The key information is effectively prevented from being peeped maliciously, the data security protection level of the public terminal is improved, and security and convenience are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of privacy protection for electronic devices, specifically to a method, system, terminal, and medium for dynamic screen privacy protection. Background Technology

[0002] In public places such as government service halls and bank branches, a large number of government service kiosks and bank ATMs are deployed for public use, displaying users' personal information when conducting business. Currently, privacy protection for these public terminals mainly relies on two methods: one is to install wide-viewing-angle privacy films, but this reduces the overall brightness and clarity of the screen, affecting the user's viewing experience, and the effect is even worse in strong light environments; the other is to use simple distance sensors or area warning signs to remind users to pay attention to their surroundings, but lacks proactive and real-time protection capabilities. Especially in crowded and complex public environments, traditional methods are difficult to effectively deal with malicious peeping behaviors such as shoulder peeping. Related solutions attempt to detect bystanders and blur the screen using cameras, but often fail to distinguish between the user and the bystander, cannot determine the bystander's true intentions, or fail to differentiate the user's own line of sight, resulting in false alarms or slow responses, seriously affecting business processing efficiency and user experience. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides a method, system, terminal, and medium for dynamic screen privacy protection. By comprehensively assessing personnel threats, user operations, and interface content sensitivity, and employing a focus-following purification strategy and business context awareness, protection is automatically triggered when high-risk prying is detected. This avoids interference with user operations caused by traditional protection methods, ensures smooth and accurate business processing, and achieves dynamic matching between protection strength and business risks. It effectively prevents critical information from being maliciously spied on, improves the level of data security protection for public terminals, and balances security and convenience.

[0004] In a first aspect, the technical solution of the present invention provides a method for dynamic protection of screen privacy, comprising the following steps: Acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data; Determine whether the overall privacy risk level is higher than the preset risk level threshold; If not, then control the graphics rendering engine to render the screen content normally; If so, the focus-following purification strategy is enabled, and the graphics rendering engine is controlled to perform privacy-preserving rendering processing, including: overlaying a dynamic visual interference layer on the original screen content, and generating a local purification area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is preserved, while the rest of the screen content forms visual interference.

[0005] Secondly, the technical solution of the present invention provides a screen privacy dynamic protection system, comprising: The privacy risk level determination module is used to acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data. The rendering strategy switching condition judgment module is used to determine whether the overall privacy risk level is higher than the preset risk level threshold; The screen rendering module is used to control the graphics rendering engine to render the screen content normally when the overall privacy risk level is not higher than the preset risk level threshold; when the overall privacy risk level is higher than the preset risk level threshold, it controls the graphics rendering engine to perform privacy protection rendering processing, including: superimposing a dynamic visual interference layer on the original screen content, and generating a local cleanup area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is preserved, while the rest of the screen content forms visual interference.

[0006] Thirdly, the technical solution of the present invention provides a terminal, including: Storage device used to store dynamic screen privacy protection programs; A processor is used to implement the steps of the screen privacy dynamic protection method described above when executing the screen privacy dynamic protection program.

[0007] Fourthly, the present invention provides a computer-readable storage medium storing a screen privacy dynamic protection program, wherein the screen privacy dynamic protection program, when executed by a processor, implements the steps of the screen privacy dynamic protection method described above.

[0008] As can be seen from the above technical solutions, this application has the following advantages: By comprehensively assessing the threat level of surrounding personnel, user operation status, and the sensitivity of the current interface content, the system can automatically trigger protection when high-risk spying intentions are detected, effectively preventing malicious spying of key information during business processing and improving the data security protection level of public terminals; by adopting a focus-following purification strategy, it ensures that the screen content in the user's field of vision remains clear and visible when the user enters a password or verifies sensitive information, avoiding the interference caused by traditional overall blurring or dimming of the screen to the user's operation, and ensuring the smoothness and accuracy of business processing; at the same time, through business context awareness, it can identify the sensitivity differences of different processing links, provide protection for high-sensitivity links, and give lower alerts for links where information can be disclosed, achieving a match between protection strength and business risk, and ensuring a balance between security and convenience. Attached Figure Description

[0009] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of a screen privacy dynamic protection method provided in an embodiment of the present invention.

[0011] Figure 2 This is a schematic block diagram of a screen privacy dynamic protection system provided in an embodiment of the present invention.

[0012] Figure 3 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present invention. Detailed Implementation

[0013] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0015] Figure 1 This is a schematic flowchart of a dynamic screen privacy protection method provided in an embodiment of the present invention. Figure 1 The executing entity can be a dynamic screen privacy protection system. The dynamic screen privacy protection method provided in this embodiment is executed by a computer device, and correspondingly, the dynamic screen privacy protection system runs on the computer device. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0016] like Figure 1 As shown, the method includes the following steps.

[0017] S1 acquires environmental awareness data, user status data, and business context data, and determines the current comprehensive privacy risk level based on the acquired data.

[0018] S2 determines whether the overall privacy risk level is higher than the preset risk level threshold.

[0019] S3, if not, controls the graphics rendering engine to render the screen content normally.

[0020] S4, if so, then enable the focus-following purification strategy and control the graphics rendering engine to perform privacy-preserving rendering processing, including: overlaying a dynamic visual interference layer on the original screen content, and generating a local purification area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is maintained, while the rest of the screen content forms visual interference.

[0021] As a refinement and extension of the specific implementation of the above embodiments, in order to fully explain the specific implementation process of this embodiment, the following will provide possible embodiments to describe the specific implementation of the above steps in a non-limiting manner.

[0022] Considering that privacy threats are not determined by a single factor, but rather by the dynamic interaction of the environment, the user's own state, and the screen content, this embodiment constructs a comprehensive privacy risk level assessment model based on temporal inference through multi-source heterogeneous data fusion. It achieves risk level determination by parallel collection and structuring of three types of feature vectors, followed by temporal fusion and lightweight inference. Specifically, step S1 acquires environmental awareness data, user state data, and business context data, and determines the current comprehensive privacy risk level based on the acquired data, followed by steps S101 to S104.

[0023] S101, Collect perception data of the surrounding area and extract environmental feature vectors. The environmental feature vectors include at least a proximity score calculated based on the distance of the non-operating user relative to the device and their movement speed, and a spying intent score calculated based on their body and line of sight orientation.

[0024] Through sensors such as depth cameras, the system continuously captures the dynamics of multiple targets in the area surrounding the device. It calculates the physical indicators (distance, speed) of non-operating users and analyzes the relative relationship between their body orientation, line of sight, and the device screen to synthesize proximity scores and spying intent scores. This transforms "someone is approaching" into a quantifiable, semantically meaningful threat level. Furthermore, it extracts statistics on the number of threats and their trends over time to form an environmental feature vector reflecting the overall threat situation.

[0025] S102: Collect facial image data of the user, extract gaze tracking features, and construct user state feature vector.

[0026] By using a camera facing the user, the system tracks their gaze focus and the angle between their gaze and the screen in real time. Combined with gaze dwell time and the flow of user actions, the system calculates attentional distraction and operational activity, thus constructing a user state feature vector to determine whether the user is focused on the task at hand. When a user's attention is distracted, their ability to perceive environmental risks decreases, and the system needs to increase its alertness; conversely, when the user is focused, their own alertness can be relied upon appropriately.

[0027] S103, obtain the sensitivity information of the current business interface and construct the business context feature vector.

[0028] The screen UI element tree is parsed, and each visible element is labeled with a sensitivity level according to predefined rules. A business context feature vector is generated by calculating the sensitive area coverage, identifying the highest sensitivity level, and determining the sensitivity of the element with input focus. This vector reflects the inherent risk value of the information carried by the screen at the current moment.

[0029] S104 performs time alignment and fusion of environmental feature vectors, user state feature vectors, and business context feature vectors, and inputs them into a pre-trained lightweight temporal inference model, which outputs a comprehensive privacy risk level.

[0030] A pre-trained lightweight temporal reasoning model learns the intrinsic relationships between three types of features, ultimately outputting a comprehensive privacy risk level. For example, when environmental features indicate "the presence of a high-threat spy" (high spying intent score), user state features indicate "the user is staring intently at a certain part of the screen" (high focus), and business context features indicate "a password is being entered at this location" (extremely high sensitivity), the model should output a high-risk level. This embodiment quantifies and aligns information from the three dimensions and uses a lightweight temporal model to simulate their comprehensive decision-making process, enabling real-time and accurate assessment of screen privacy risks and improving the accuracy of dynamic privacy protection.

[0031] Step S101 collects perception data of the surrounding area and extracts environmental feature vectors, specifically including the following steps S101.1 to S101.5.

[0032] S101.1, Collect image data containing depth information, perform real-time target detection and multi-target tracking on the image data, assign a unique identifier to each detected non-operational user and obtain their motion trajectory.

[0033] The device continuously captures image sequences of the areas in front of and to the sides using a depth camera or a combination of vision sensors. These image sequences contain both color (RGB) and depth information.

[0034] For each frame of the image, a real-time object detection algorithm, such as a lightweight model based on YOLO or SSD, is run to identify all "human" targets in the image. For each detected human target, a multi-target tracking algorithm is further executed. This algorithm can be based on DeepSORT or a related lightweight tracking framework. A unique identifier (ID) is assigned to the target, and the tracking consistency of this ID is maintained in subsequent frames. This allows for the acquisition of the continuous motion trajectory of each non-operating user, ensuring that different individuals can be distinguished and that the behavior of each individual can be independently and continuously observed and analyzed.

[0035] S101.2, For each tracked non-operational user, calculate their basic physical features and high-level semantic features; the basic physical features include at least the distance relative to the device. radial velocity and azimuth Advanced semantic features include at least proximity scores and spying intent scores.

[0036] For each successfully tracked non-operational user individual (represented by its unique ID), two types of features are computed in parallel: basic physical features and advanced semantic features.

[0037] Basic physical features refer to the calculation of the precise spatial relationship between an individual and the device screen using depth information and camera coordinate system transformation. Distance refers to the straight-line distance from the individual to the center of the device screen. Radial velocity refers to the velocity component of the individual along the line of sight (towards or away from the screen); positive values ​​indicate approach, and negative values ​​indicate distance. Azimuth is the horizontal angle of the individual relative to the front of the device screen, used to determine whether it is facing forward, to the side, or behind. This is how features describe the objective motion state of the target in numerical form.

[0038] Advanced semantic features refer to interpreting an individual's behavioral intentions based on physical features and through a predefined computational model.

[0039] The proximity score is a continuous value between 0 and 1, determined by both distance and radial velocity. The closer the distance and the faster the radial velocity (approach), the higher the score. This score is used to quantify the probability and urgency of an individual "approaching the screen".

[0040] Proximity score The calculation formula is:

[0041] In the formula, The distance between the non-operating user and the device. Its radial velocity, and As a preset constant, and These are the weighting coefficients.

[0042] The voyeurism intention score is a continuous value between 0 and 1, calculated by analyzing an individual's 3D body orientation, estimated head posture, and the angle between the estimated gaze direction and the screen. The score increases when the individual's body and / or gaze is explicitly directed towards the screen. This score quantifies the likelihood that an individual "intentionally views screen content."

[0043] The formula for calculating the voyeurism score is:

[0044] In the formula, The angle between the non-operating user's body orientation and the normal to the device screen. This is the estimated angle between the non-operating user's line of sight and the normal to the terminal screen. and These are the weighting coefficients.

[0045] S101.3, based on proximity score and spying intent score, classifies each non-operational user into low threat, medium threat, or high threat level through predefined two-dimensional decision rules.

[0046] To transform continuous semantic scores into discrete decisions, a predefined two-dimensional decision rule is employed, such as a decision matrix or decision tree with proximity score and spying intent score as axes. Based on each individual's two scores, it is categorized into one of three threat levels.

[0047] Low threat: Both proximity score and spying intent score are low. For example, a person walking by at a distance or with their back to the screen.

[0048] Medium threat: One score is average, or both scores are at an average level. For example, a person standing at a moderate distance to the side with an uncertain line of sight.

[0049] High threat: High proximity score and / or spying intent score. For example, a person approaching the screen from the side or rear and clearly looking at the screen.

[0050] S101.4 generates statistical features and time-series features. Statistical features include the current total number of people, the number of high-threat people, the number of medium-threat people, the nearest threat distance, and the maximum proximity score. Time-series features include the changing trend of the number of high-threat people in the most recent K time windows.

[0051] This step further aggregates the threat classification results of individuals to characterize the macro-threat situation and its dynamic changes across the entire monitoring area.

[0052] Statistical features are constructed by extracting global statistical information within the current frame or the current short-term window, including: Current total number of people The total number of people being tracked in the environment; High / medium threat number The number of individuals at the high and medium threat levels, respectively; Closest threat distance The distance to the device among all medium- and high-threat individuals; Maximum proximity score The maximum proximity score among all individuals.

[0053] These static characteristics are used to describe the instantaneous threat density and intensity of the environment.

[0054] Analyzing the changing trends of key threat indicators over the most recent K consecutive time windows constitutes a time-series characteristic. For example, calculating the slope or difference of the number of high-threat individuals over the most recent K windows. If this value remains consistently positive, it indicates that the number of high-threat individuals is increasing, and the environmental risk is rising; conversely, it indicates that the risk is mitigating. Time-series characteristics are used to capture the direction and speed of threat evolution, providing early warning of gradually approaching risks.

[0055] S101.5 concatenates statistical features and temporal features to generate an environmental perception feature vector.

[0056] All the statistical and temporal features calculated above are concatenated in a predetermined order to form a unified, fixed-dimensional environmental perception feature vector. .

[0057] Step S102 involves collecting facial image data of the user, extracting gaze tracking features, and constructing a user state feature vector, specifically including the following steps S102.1 to S102.5.

[0058] S102.1, based on facial images, extract facial key points and construct a 3D eye model, and calculate the gaze direction vector.

[0059] The system continuously captures facial images of the user using the front-facing camera. First, it runs a facial landmark detection model, detecting key feature points such as the center of the pupil, the corners of the eyes, the tip of the nose, and the corners of the mouth. Based on the two-dimensional positions of these landmarks and their correspondence on the three-dimensional face model, combined with the camera's intrinsic parameters, the three-dimensional pose of the user's head, including rotation and translation, is estimated using algorithms such as Perspective-n-Point (PnP).

[0060] Furthermore, a simplified 3D eye model is constructed. This model approximates the eyeball as a sphere, and based on the detected position of the pupil center in the image, combined with prior knowledge or calibration parameters such as the known eyeball radius and optical center offset, the direction of the gaze in the eye coordinate system is calculated. Finally, by fusing head posture and relative eyeball rotation, the user's three-dimensional gaze direction vector at the current moment is calculated. This vector defines the direction of the gaze in the user's head coordinate system.

[0061] S102.2, map the gaze direction vector to the screen coordinate system to obtain the real-time gaze focus coordinates and the angle between the gaze and the screen normal.

[0062] The calculated gaze direction vector is mapped onto the screen's two-dimensional pixel coordinate system using the calibrated spatial transformation relationship between the camera and the screen. This mapping process solves the problem of the intersection point between the gaze and the screen plane, thus obtaining the precise coordinates of the user's gaze point on the screen, i.e., the real-time gaze focus coordinates. This coordinate indicates the specific location on the screen that the user is currently viewing.

[0063] Simultaneously, calculate the angle between the viewing direction vector and the screen plane normal vector. When the angle is close to 0 degrees, it indicates that the user is looking directly at the screen, providing the best viewing experience and maximizing their focus on the screen content. The larger the angle, the further the user's gaze is from the front of the screen, possibly indicating that they are looking to the side, glancing briefly, or about to look away. Their focus on the screen content and efficiency in acquiring information may decrease.

[0064] S102.3 Calculate attention distraction based on the dwell information of the gaze focus coordinates within a preset time window.

[0065] Within a preset sliding time window, a sequence of gaze focus coordinates is continuously recorded. By analyzing this sequence, an attention distraction index is calculated. The calculation formula is:

[0066] in, The total duration of effective fixation, This represents the length of the time window.

[0067] Attention distractibility characterizes the degree of concentration and stability of a user's attention. Highly distracted attention may mean that the user is being disturbed by the environment or is in a non-task state.

[0068] S102.4 Calculate operation activity based on touch or keyboard event stream. .

[0069] Parallel listening to input event streams from the touchscreen, mouse, or keyboard. Calculating an operational activity metric within a time window that coincides with or is associated with the attention assessment. This metric can be defined as: The number of events per unit of time: such as the number of touch clicks or keyboard keystrokes per second; The regularity or suddenness of events: For example, continuous, steady input differs from sporadic, intermittent clicks in terms of activity patterns.

[0070] High user activity is strongly correlated with users actively engaging in tasks such as data input and menu selection, indicating that users are "online" and "participating".

[0071] S102.5, the coordinates of the gaze focus, the angle between the gaze and the screen normal, the degree of attention distraction, and the degree of operation activity are used to form a user state feature vector.

[0072] The four core indicators obtained from the above steps are normalized and concatenated to form a user state feature vector. .

[0073] Step S103 obtains the sensitivity information of the current business interface and constructs a business context feature vector, specifically including the following steps S103.1 to S103.6.

[0074] S103.1, Get the element tree structure of the currently displayed interface.

[0075] The system obtains the element tree structure of the application interface currently displayed on the screen in real time through interfaces provided by the operating system or application framework. This tree structure programmatically describes the hierarchy and containment relationships between all visible and logically existing user interface elements on the screen, and includes attributes such as the type, position, size, and text content of each element. User interface elements include windows, buttons, text boxes, images, layout containers, etc.

[0076] S103.2, based on a predefined sensitivity classification rule base, assigns sensitivity labels to visible interface elements in the element tree. Sensitivity labels include public, low sensitivity, medium sensitivity, high sensitivity, and extremely high sensitivity levels.

[0077] Maintain a predefined sensitivity classification rule base. This rule base contains a series of matching rules that can automatically determine the sensitivity of UI elements based on various attributes.

[0078] Based on element type: For example, "password input box ( <input type="‘password’"> "Amount display text box" is usually marked as "Very high sensitivity"; "Amount display text box" is marked as "High sensitivity"; "Ordinary label text" may be marked as "Public" or "Low sensitivity".

[0079] Based on text content keywords: For example, adjacent input boxes / display areas containing field labels or prompt text such as "ID number", "card number", "mobile phone number", "detailed address" can be marked as "high sensitivity" or "extremely high sensitivity".

[0080] Based on element attributes and context: For example, a button on the "Transaction Confirmation" page may be more sensitive than a button on the "Home" page.

[0081] Traverse all currently visible UI elements in the element tree and assign a sensitivity label to each using a rule base. Label levels can be multi-level, for example: Public (Level 0), Low Sensitivity (Level 1), Medium Sensitivity (Level 2), High Sensitivity (Level 3), and Extremely High Sensitivity (Level 4). This step transforms the screen's visual content into structured data with privacy risk semantics.

[0082] S103.3, a weighting coefficient is preset for each sensitivity level. The weighted risk area of ​​each element is calculated based on the weighting coefficient. The weighted risk areas of all elements are accumulated to obtain the total weighted risk area of ​​the screen. The total weighted risk area is divided by the total pixel area of ​​the screen to obtain the sensitive area coverage.

[0083] To quantify the spatial distribution density of sensitive information on the screen, the system performs weighted area calculations. A weight coefficient is assigned to each sensitivity level, which increases as the level increases. For example: Open: 0, Low Sensitivity: 0.2, Medium Sensitivity: 0.5, High Sensitivity: 0.8, Extremely High Sensitivity: 1.0.

[0084] For each UI element, based on its pixel area and the weight coefficients corresponding to its sensitivity labels Calculate its weighted risk area The total weighted risk area of ​​the screen is obtained by summing the weighted risk areas of all visible elements on the screen. This total value is then divided by the total pixel area of ​​the screen to obtain the sensitive area coverage rate. This value is a floating-point number between 0 and 1, reflecting the weighted spatial proportion of sensitive information on the screen. The higher the value, the denser the privacy risks carried by the current interface as a whole.

[0085] S103.4, Get the highest sensitivity level among all elements on the current screen.

[0086] Iterate through the sensitivity labels of all elements on the current screen and find the highest sensitivity level. For example, if there is even one "Very High Sensitivity" element on the interface, regardless of its size, the highest sensitivity level is "Very High Sensitivity (Level 4)". This indicator identifies the highest potential risk level of the information contained in the current interface.

[0087] S103.5 Identify the interface element that currently has input focus and determine the current input focus state based on its sensitivity label. This state is a discrete value used to indicate the sensitivity level range of the interface element where the input focus is located.

[0088] It monitors the interface elements that currently have input focus in real time, such as text boxes, buttons, etc., where the cursor is located or is accepting keyboard input. After identifying the element, it determines a current input focus state based on its assigned sensitivity label.

[0089] This state can be represented by a discrete enumeration value, such as: FOCUS_PUBLIC, FOCUS_LOW, FOCUS_MEDIUM, FOCUS_HIGH, FOCUS_CRITICAL. This state reflects the sensitivity of the information involved in the user's current interaction. When the focus is on an "extremely sensitive" password field, it means the user is in the most vulnerable "entering critical credentials" stage, and the system should provide the highest level of real-time protection.

[0090] S103.6 constructs a business context feature vector from the sensitive area coverage, the highest sensitivity level, and the current input focus state.

[0091] In determining the overall privacy risk level, this embodiment uses a lightweight temporal reasoning model as the decision engine. This model is designed to run in real time on resource-constrained terminal devices. Its goal is to fuse and analyze three types of asynchronous, heterogeneous, and time-dependent data: environmental awareness feature vectors, user state feature vectors, and business context feature vectors, and output a discrete overall privacy risk level.

[0092] To meet the requirements of real-time performance, low power consumption, and expressive power, this embodiment employs a lightweight temporal inference model, which is a variant of the Temporal Convolutional Network (TCN), a Gated Recurrent Unit (GRU), or a lightweight LSTM. The model's input layer is designed with three independent input branches, receiving aligned environment, user state, and business context feature vectors, respectively. These branches can first undergo feature adaptation and dimensionality reduction through their respective small fully connected networks, followed by concatenation or early fusion via a cross-attention mechanism, before finally being fed into the core temporal processing layer. The model's output layer is a Softmax layer, with the output dimension corresponding to the number of categories representing the overall privacy risk level.

[0093] Specifically, the model input is a time series segment X=[x_t, x_{t-1}, ..., x_{t-T+1}], where x_i represents a fused feature vector at time step i. x_i itself is composed of three parts: an environment-aware feature vector, a user state feature vector, and a business context feature vector. The sequence length T is a hyperparameter, for example, corresponding to data from the past 2-5 seconds, ensuring that the model can capture meaningful dynamic trends.

[0094] The model output is the prediction result y_t for the current time step t, which is a probability distribution vector, for example, y_t = [P(low risk), P(medium risk), P(high risk)]. The final comprehensive privacy risk level is the category with the highest probability.

[0095] Model Training Process: Environmental awareness data, user state data, and business context data (which must be anonymized) are collected synchronously during device operation in various typical scenarios. Security experts, or through a pre-set, more complex rule simulation system, label each data combination at any given time with a "realistic" privacy risk level. Labeling must comprehensively consider factors such as the presence of a definite spy, the user's level of alertness, and the sensitivity of the current screen content. The collected raw features are standardized and normalized, and continuous sequence samples are constructed according to a time window T. The cross-entropy loss function is used to measure the difference between the model's predicted probability distribution and the true label distribution. Adaptive optimizers such as AdamW are employed, combined with learning rate scheduling strategies such as cosine annealing, for training to ensure stable convergence and avoid overfitting. Dropout and weight decay techniques are widely used during training to prevent overfitting.

[0096] This embodiment automatically switches between normal rendering mode and privacy-preserving rendering mode based on the comprehensive privacy risk level output by the lightweight temporal inference model. When the risk level exceeds a preset threshold, a focus-following purification rendering strategy is activated to effectively prevent onlookers from spying while minimizing interference with the user's operations.

[0097] A preset risk level threshold is maintained. When the overall privacy risk level is less than or equal to the risk level threshold, the current environment is deemed safe or the threat is controllable. The system controls the graphics rendering engine to directly send the frame buffer content generated by the application to the display according to the standard process, so that the screen displays clear, unedited, original content.

[0098] When the overall privacy risk level exceeds the risk level threshold, a significant risk of eavesdropping is identified. Immediately switch to protected mode and control the graphics rendering engine to execute post-processing or overlay rendering processes, i.e., privacy-preserving rendering processing. Privacy-preserving rendering processing includes steps one and two.

[0099] Step 1: Generation and overlay of dynamic visual interference layers.

[0100] Later in the graphics rendering pipeline, or via a separate transparent overlay, a dynamic visual interference layer is synthesized and superimposed in real time on top of the original screen content. This interference layer covers the entire screen area and is used to visually degrade the original content below, making it difficult for onlookers at a certain distance and angle to discern details. The interference effect can be, but is not limited to: Controllable blur: such as Gaussian blur and motion blur, the blur intensity can be finely adjusted according to the risk level; Pixelation / Mosaic: Divide the image into blocks and take the average color; Noise overlay: Add dynamic visual noise (such as snowflake noise); Color and contrast distortion: Slightly alters hue, saturation, and contrast.

[0101] Step two: Focus on the generation of the localized purification area.

[0102] The system receives real-time gaze focus coordinates from the user status perception module and dynamically generates a local cleanup area centered on these gaze focus coordinates on the dynamic visual interference layer.

[0103] Specifically, a gradient transparency mask is applied to the pixel location corresponding to the interference layer. This mask is a radially gradient alpha (transparency) map. The transparency is highest at the center point (i.e., the focal point of the view), and gradually decreases (the alpha value increases) towards the edge until it smoothly blends with the completely opaque (alpha=1) area of ​​the surrounding interference layer.

[0104] The gradual change in transparency ensures that there is no abrupt boundary between the purified area and the interference area, avoiding visual jarring and further increasing the difficulty for onlookers to judge the extent of the purified area. The mask shape can be circular or elliptical. As the transparency of the interference layer in this area decreases, or even becomes completely transparent, the original content of the screen below is clearly visible. Therefore, for the user, the content within their field of vision remains clearly visible; while the screen area outside their field of vision becomes difficult to discern due to being covered by the dynamic visual interference layer.

[0105] In this embodiment, the method further includes continuously monitoring the overall privacy risk level. When the overall privacy risk level is detected to be lower than the risk level threshold and remains stable for a preset period of time, the focus-following purification strategy is exited, and normal rendering of the screen content is restored. Through risk monitoring and state self-recovery, the dynamic matching between protective behavior and actual risk is ensured, and unnecessary interference rendering is avoided even after the risk has been eliminated.

[0106] Specifically, during operation, the overall privacy risk level is continuously monitored, with the monitoring cycle synchronized with the running cycle of the risk inference model or the screen refresh rate. In each monitoring cycle, the lightweight temporal inference model outputs the overall privacy risk level for the current moment. When continuous monitoring detects that the latest output overall privacy risk level is again below or equal to the preset risk level threshold, the exit condition is initially triggered. A countdown timer or state maintenance counter is then started, with a timing length equal to a preset stabilization time. Within this stabilization time window, the risk level for consecutive monitoring cycles must remain below the threshold to filter out brief fluctuations or instantaneous decreases in risk. Only when the risk level remains consistently below the threshold throughout the entire preset stabilization time window, and the exit condition is finally confirmed, is an instruction sent to the graphics rendering engine to exit the focus-following cleanup strategy. The graphics rendering engine stops generating and overlaying dynamic visual interference layers, and the logic for generating local cleanup regions based on the gaze focus associated with the interference layer also stops. The graphics rendering engine switches back to standard rendering mode, directly displaying the original content of the application's framebuffer.

[0107] The foregoing has described in detail an embodiment of a dynamic screen privacy protection method. Based on the dynamic screen privacy protection method described in the above embodiment, this invention also provides a dynamic screen privacy protection system corresponding to the method.

[0108] Figure 2 This is a schematic block diagram of a screen privacy dynamic protection system provided in an embodiment of the present invention. In this embodiment, the screen privacy dynamic protection system 200 can be divided into multiple functional modules according to the functions it performs. A module, as referred to in this invention, is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory.

[0109] The privacy risk level determination module 210 is used to acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data.

[0110] The rendering strategy switching condition judgment module 220 is used to determine whether the overall privacy risk level is higher than the preset risk level threshold.

[0111] The screen rendering module 230 is used to control the graphics rendering engine to render the screen content normally when the overall privacy risk level is not higher than the preset risk level threshold; when the overall privacy risk level is higher than the preset risk level threshold, it controls the graphics rendering engine to perform privacy protection rendering processing, including: superimposing a dynamic visual interference layer on the original screen content, and generating a local purification area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is maintained, while the rest of the screen content forms visual interference.

[0112] The screen privacy dynamic protection system of this embodiment is used to implement the aforementioned screen privacy dynamic protection method. Therefore, the specific implementation of this system can be found in the embodiment section of the screen privacy dynamic protection method above. Thus, the specific implementation can be referred to the description of the corresponding embodiments, and will not be elaborated here.

[0113] Furthermore, since the screen privacy dynamic protection system in this embodiment is used to implement the aforementioned screen privacy dynamic protection method, its function corresponds to the function of the above method, and will not be described again here.

[0114] Figure 3 This is a schematic diagram of the structure of a terminal 300 provided in an embodiment of the present invention, including: a processor 310, a memory 320, and a communication unit 330. The processor 310 is used to implement the process steps of the above-described embodiment of the screen privacy dynamic protection method when implementing the screen privacy dynamic protection program stored in the memory 320.

[0115] This invention also provides a computer storage medium, which may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. The computer storage medium stores a dynamic screen privacy protection program, which, when executed by a processor, implements the process steps of the above-described dynamic screen privacy protection method embodiment.

[0116] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for dynamic protection of screen privacy, characterized in that, Includes the following steps: Acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data; Determine whether the overall privacy risk level is higher than the preset risk level threshold; If not, then control the graphics rendering engine to render the screen content normally; If so, the focus-following purification strategy is enabled, and the graphics rendering engine is controlled to perform privacy-preserving rendering processing, including: overlaying a dynamic visual interference layer on the original screen content, and generating a local purification area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is preserved, while the rest of the screen content forms visual interference.

2. The screen privacy dynamic protection method according to claim 1, characterized in that, Acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data, specifically including: Collect perception data of the surrounding area and extract environmental feature vectors, which include at least a proximity score calculated based on the distance of the non-operating user relative to the device and their radial velocity, and a spying intent score calculated based on their body and line of sight orientation. Collect facial image data of the user, extract gaze tracking features, and construct a user state feature vector; Obtain the sensitivity information of the current business interface and construct a business context feature vector; The environmental feature vector, user state feature vector, and business context feature vector are time-aligned and fused, and then input into a pre-trained lightweight temporal inference model, which outputs a comprehensive privacy risk level.

3. The screen privacy dynamic protection method according to claim 2, characterized in that, Collect sensor data from the surrounding area and extract environmental feature vectors, specifically including: Collect image data containing depth information, perform real-time target detection and multi-target tracking on the image data, assign a unique identifier to each detected non-operational user and obtain their motion trajectory; For each tracked non-operational user, calculate their basic physical features and advanced semantic features; the basic physical features include at least distance relative to the device, radial velocity, and azimuth; the advanced semantic features include at least proximity score and spying intent score. Based on proximity score and spying intent score, each non-operational user is classified into low threat, medium threat, or high threat level using predefined two-dimensional decision rules. Generate statistical features and time-series features. The statistical features include the current total number of people, the number of high-threat people, the number of medium-threat people, the nearest threat distance, and the maximum proximity score. The time-series features include the changing trend of the number of high-threat people in the most recent K time windows. Statistical features and time-series features are concatenated to generate an environmental perception feature vector.

4. The screen privacy dynamic protection method according to claim 2, characterized in that, The process involves collecting facial image data of the user, extracting gaze tracking features, and constructing a user state feature vector, specifically including: Based on facial images, key facial points are extracted and a 3D eye model is constructed to calculate the gaze direction vector. Map the gaze direction vector to the screen coordinate system to obtain the real-time gaze focus coordinates and the angle between the gaze and the screen normal. Calculate the degree of attentional distraction based on the dwell time information of the gaze focus coordinates within a preset time window; Calculate the activity level based on touch or keyboard event streams; The user state feature vector is constructed by considering the coordinates of the gaze focus, the angle between the gaze and the screen normal, the degree of attention distraction, and the activity level.

5. The screen privacy dynamic protection method according to claim 4, characterized in that, Obtain the sensitivity information of the current business interface and construct a business context feature vector, specifically including: Get the element tree structure of the currently displayed interface; Based on a predefined sensitivity classification rule base, sensitivity labels are assigned to visible interface elements in the element tree. Sensitivity labels include public, low sensitivity, medium sensitivity, high sensitivity, and extremely high sensitivity levels. A weighting coefficient is preset for each sensitivity level. The weighted risk area of ​​each element is calculated based on the weighting coefficient. The weighted risk areas of all elements are accumulated to obtain the total weighted risk area of ​​the screen. The total weighted risk area is divided by the total pixel area of ​​the screen to obtain the sensitive area coverage. Get the highest sensitivity level among all elements on the current screen; Identify the interface element that currently has input focus and determine the current input focus state based on its sensitivity label. This state is a discrete value used to indicate the range of sensitivity levels of the interface element where the input focus is located. The sensitive area coverage, the highest sensitivity level, and the current input focus state are used to construct the business context feature vector.

6. The screen privacy dynamic protection method according to claim 1, characterized in that, The method for generating local clean areas is as follows: with the user's real-time gaze focus coordinates as the center, a gradient transparency mask is applied to the pixels at the corresponding positions of the dynamic visual interference layer to reduce the transparency of the interference layer in that area.

7. The screen privacy dynamic protection method according to claim 1, characterized in that, The method also includes: Continuously monitor the overall privacy risk level; Once the overall privacy risk level is detected to be below the risk level threshold and remains stable for a preset period of time, the focus-following purification strategy will be exited, and normal rendering of screen content will be restored.

8. A screen privacy dynamic protection system, characterized in that, include: The privacy risk level determination module is used to acquire environmental awareness data, user status data, and business context data, and determine the current comprehensive privacy risk level based on the acquired data. The rendering strategy switching condition judgment module is used to determine whether the overall privacy risk level is higher than the preset risk level threshold; The screen rendering module is used to control the graphics rendering engine to render the screen content normally when the overall privacy risk level is not higher than the preset risk level threshold. When the overall privacy risk level exceeds the preset risk level threshold, the graphics rendering engine is controlled to perform privacy-preserving rendering processing, including: overlaying a dynamic visual interference layer on the original screen content, and generating a local cleanup area on the dynamic visual interference layer based on the user's real-time gaze focus coordinates in the acquired user state data, so that the screen content area corresponding to the user's real-time gaze focus coordinates is preserved, while the rest of the screen content forms visual interference.

9. A terminal, characterized in that, include: Storage device used to store dynamic screen privacy protection programs; A processor, configured to implement the steps of the screen privacy dynamic protection method as described in any one of claims 1 to 7 when executing the screen privacy dynamic protection program.

10. A computer-readable storage medium, characterized in that, The readable storage medium stores a screen privacy dynamic protection program, which, when executed by a processor, implements the steps of the screen privacy dynamic protection method as described in any one of claims 1 to 7.