Picture sensitive information detection method and system
By embedding watermarks in the terminal and web interfaces, combining large language models and multimodal technology to detect image sensitive information, the problem of data leakage cannot be detected in time in traditional methods is solved, active early warning and rapid traceability are achieved, and data security protection capabilities are improved.
Patent Information
- Application Number
- CN202510486815.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
It is difficult for the existing technology to actively discover and timely detect outgoing pictures with sensitive information. Traditional methods conduct post-tracement and responsibilities after data leakage, and cannot promptly warn of potential risks, and their generalization capabilities are limited.
By embedding watermarks on the terminal desktop and web application interfaces, the outgoing behavior of the user's image is detected, global visual and local semantic feature extraction is performed, forward/reverse hypothetical text is generated using a large language model and contradiction verification is performed, and multi-modal large model is combined to evaluate whether the picture carries sensitive information, and watermark detection and traceability are performed.
It realizes active discovery and timely warning of data leakage, improves detection accuracy and generalization capabilities, can quickly locate the source of leakage, build a complete security protection system, and reduces data leakage losses.
Smart Images

Figure CN120012056A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data security technology, and in particular to a method and system for detecting sensitive information in an image. Background Art
[0002] In the era of mobile Internet, sensitive, private or confidential data leakage incidents occur frequently in enterprises and institutions at all levels, which not only bring irreparable economic losses to the relevant units, but also bring great challenges to national defense security and social stability. Data security is increasingly becoming a key issue of concern to enterprises and institutions at all levels.
[0003] Government business systems carry high-value data, concentrate sensitive information, and need to face many application departments and a large number of users. They are exposed to a large number of risks and are difficult to control information security. Especially in the process of using human-computer interaction data, when electronic data is converted into information forms that can be directly perceived by humans, such as through the display screen. After being converted into optical information, it often deviates from the scope of traditional data security control and there is a huge risk of data leakage.
[0004] Screen display is one of the essential core components of human-computer interaction, and the popularity of various smart terminals makes it easier to steal secrets through screen photography. At the same time, it is often difficult to trace the person responsible for leaking or stealing secrets through traditional means. Especially with the widespread use of smart phone terminals, the confidentiality and security of key information faces major challenges, and the difficulty of leak control is increasing. At present, leaks in the form of screen photography and screenshots have become a management blind spot and core pain point for the security and confidentiality work of many units.
[0005] With the popularity of mobile devices and the rapid development of the Internet, data leakage is becoming increasingly serious. Sensitive, private or confidential data of enterprises and institutions are frequently leaked, posing huge challenges to personal privacy and corporate interests. The traditional solution is to audit outbound data by installing DLP data leakage prevention software on network exits and terminal devices, and to prevent data leakage by installing watermark software and anti-leakage software on terminal devices. Third, the traditional image sensitive information recognition technology uses a deep learning-based method to extract the convolution features of images using networks such as CNN. The boundaries of sensitive information judgment are fuzzy, the generalization ability is poor, and it is easily affected by image processing operations (such as compression, filtering, etc.).
[0006] Traditional DLP software mainly relies on predefined rules and patterns, which can respond to known threats, but may have limited effect on unknown threats and variants. It is impossible to accurately trace and locate the source of leaked outbound data.
[0007] Traditional watermarking software can only trace and determine the responsibility of data leakage that has already occurred, and cannot timely and effectively warn of potential data leakage risks. In addition, traditional anti-leakage software uses deep learning-based methods to extract convolutional features of images using networks such as CNN. These features are usually designed for specific steganographic algorithms or specific image types, and have limited generalization capabilities. In addition, the statistical features or underlying visual features of the image are usually extracted, which are often difficult to explain intuitively and are easily affected by image processing operations (such as compression, filtering, etc.). Summary of the invention
[0008] The technical problem to be solved by the present invention is how to actively discover and detect outgoing pictures containing sensitive information.
[0009] The present invention solves the above technical problems through the following technical means: The method for detecting sensitive information in an image includes: S1. Embed watermarks on the terminal desktop and web application interface; S2. Detect the outbound image behavior of the user and capture the outbound image; S3. Detect sensitive information on the outgoing image, specifically: Performing global visual feature extraction and local semantic feature extraction on the external image to generate semantic information with strong correlation between the features of the two images; Recognize and extract text from the outgoing image; The semantic information strongly associated with the features of the two pictures and the text in the external picture are guided by a large language model to generate two sections of positive / reverse hypothesis text and corresponding answers; The two sections of positive / reverse hypothesis text and the corresponding answers are subjected to contradiction verification and evidence synthesis to generate a reinforced synthetic evidence vector; evaluating whether the outgoing image carries sensitive information based on the synthesized evidence vector; S4. Perform watermark detection and traceability on the outbound images carrying sensitive information.
[0010] The present invention combines the existing security construction status, takes advantage of the multimodal large model to strengthen the detection of sensitive information in images, and builds the ability to actively discover and detect leakage behaviors such as screenshots or photos of terminal screens and web application interfaces. It serves the suspicious leakage information found during the use of important data on the government office platform, realizes the active discovery of data leakage, and provides effective early warning before the leakage incident spreads further. Potential data leakage risks can be discovered in a timely manner, so that corresponding measures can be taken to prevent and reduce the losses caused by data leakage.
[0011] Furthermore, the S3 generates semantic information of two segments of image features that are strongly associated and identifies and extracts text from the outgoing image, specifically: The BLIP-2 model is used to obtain the overall semantics and visual information of the outgoing image, and a global visual representation vector is output; the DETR model is used to identify important target objects in the image, and a corresponding local feature vector is output; The text in the image is recognized by OCR, the corresponding text content and its position are output, the global visual representation vector, the local feature vector and the text content and its position information are combined to generate a preliminary text description of the image.
[0012] Furthermore, the two sections of forward / reverse hypothesis text and corresponding answer generation method in S3 are: the preliminary description of the external image and the prompt word are input into the DeepSeek-R1 large language model; after receiving the preliminary description of the external image and the prompt word, the DeepSeek-R1 large language model performs a round of reasoning according to the image content and the prompt requirements, and gives a thinking result; The DeepSeek-R1 large language model outputs two paragraphs of text: The first paragraph of text: Answer the "non-hypothetical" question - "Does this image carry sensitive information?".
[0013] Second paragraph of text: Based on the results of the first paragraph, an opposite hypothesis is proposed and the DeepSeek-R1 large language model is allowed to further reason. If the first paragraph judges that "the image does not contain sensitive information", then the second paragraph will assume that "the image contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph judges that "the image contains sensitive information", then the second paragraph will assume that "the image does not contain sensitive information" and let the DeepSeek-R1 large language model infer possible reasons or context.
[0014] Furthermore, the contradiction verification and evidence synthesis in S3 are specifically as follows: By quantitatively analyzing the contradiction strength between the positive and negative hypothesis texts generated, multimodal joint evidence is constructed; Firstly, semantic embedding is extracted based on the text encoder, combined with the logical contradiction probability output by the natural language inference model, and the comprehensive contradiction coefficient Γ is calculated through the weighted fusion formula to dynamically balance semantic divergence and logical paradox. Subsequently, the global and local visual features are concatenated and input into the cross-modal fusion network together with the Γ value, and the gating mechanism is used to filter effective information to generate an enhanced synthetic evidence vector. The formula for calculating the contradiction intensity is as follows:
[0015] Where: T1, T2 are positive and negative texts, λ is the balance factor between semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model; The evidence synthesis formula is as follows:
[0016] in: is the visual feature fusion weight matrix, k×2d means that this is a matrix with k rows and 2d columns; the association mapping between global features and local features is established through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, so that the high contradiction area produces a significant shift in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the feature splicing operation.
[0017] Furthermore, the specific method of evaluating whether the outgoing image carries sensitive information based on the evidence vector in S3 is: First, the Mahalanobis distance of the synthetic evidence vector is calculated, and the feature legitimacy score ρ is generated through an exponential decay function. The score ρ is directly involved in the sensitive probability calculation to form the anti-disturbance confidence; Subsequently, an entropy-aware dynamic threshold mechanism is introduced. When the score ρ value is lower than the set value, the decision boundary is automatically relaxed and the sample is marked as "pending review"; When making the final decision, the system performs strict threshold judgment on samples with high ρ values, and initiates a defensive decision-making process for samples with medium and low ρ values.
[0018] The present invention also provides a system for detecting sensitive information in an image, comprising: Watermark embedding module: embed watermarks into terminal desktop and web application interfaces; Detection module: detects the outgoing pictures from the user end, wherein the outgoing pictures contain watermarks; Sensitive information detection module: extract global visual features and local semantic features of the outgoing image to generate semantic information with strong correlation between the features of two images; identify and extract text in the outgoing image; use the semantic information with strong correlation between the features of the two images and the text in the outgoing image to generate two forward / reverse hypothesis texts and corresponding answers through a large language model; perform contradiction verification and evidence synthesis on the two forward / reverse hypothesis texts and the corresponding answers to generate a reinforced synthetic evidence vector; evaluate whether the outgoing image carries sensitive information based on the synthetic evidence vector; Tracing module: performs watermark detection and tracing on the outbound images that carry sensitive information.
[0019] Furthermore, the semantic information and text recognition in the sensitive information detection module are specifically as follows: the external picture obtains the overall semantics and visual information of the image through the BLIP-2 model, and outputs the global visual representation vector; the important target objects in the image are identified through the DETR model, and the corresponding local feature vectors are output; the text in the image is identified through OCR, and the corresponding text content and its position are output; the global visual representation vector, the local feature vector and the text content and its position information are combined to generate a preliminary text description of the image.
[0020] Furthermore, the two forward / reverse hypothesis texts and corresponding answer generation methods in the sensitive information detection module are as follows: the preliminary description of the external image and the prompt word are input into the DeepSeek-R1 large language model; after receiving the preliminary description of the image and the prompt word, the DeepSeek-R1 large language model performs a round of reasoning based on the image content and the prompt requirements, and gives a thinking result; The DeepSeek-R1 large language model outputs two paragraphs of text: The first paragraph of text: Answer the "non-hypothetical" question - "Does this image carry sensitive information?".
[0021] Second paragraph of text: Based on the results of the first paragraph, an opposite hypothesis is proposed and the DeepSeek-R1 large language model is allowed to further reason. If the first paragraph judges that "the image does not contain sensitive information", then the second paragraph will assume that "the image contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph judges that "the image contains sensitive information", then the second paragraph will assume that "the image does not contain sensitive information" and let the DeepSeek-R1 large language model infer possible reasons or context.
[0022] Furthermore, the contradiction verification and evidence synthesis in the sensitive information detection module are specifically as follows: construct multimodal joint evidence by quantitatively analyzing the contradiction strength between the forward and reverse hypothesis texts generated; first, extract semantic embedding based on the text encoder, combine the logical contradiction probability output by the natural language inference model, calculate the comprehensive contradiction coefficient Γ through the weighted fusion formula, and dynamically balance semantic divergence and logical paradox; then, after splicing the global and local visual features, input them into the cross-modal fusion network together with the Γ value, use the gating mechanism to screen effective information, and generate an enhanced synthetic evidence vector; The formula for calculating the contradiction intensity is as follows:
[0023] Where: T1, T2 are positive and negative texts, λ is the balance factor between semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model; The evidence synthesis formula is as follows:
[0024] in: is the visual feature fusion weight matrix, k×2d means that this is a matrix with k rows and 2d columns; the association mapping between global features and local features is established through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, so that the high contradiction area produces a significant shift in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the feature splicing operation.
[0025] Furthermore, the specific method for outputting whether an image carries sensitive information based on the evidence vector in the sensitive information detection module is as follows: first, the Mahalanobis distance of the synthetic evidence vector is calculated, and a feature legitimacy score ρ is generated through an exponential decay function, and the score ρ is directly involved in the sensitive probability calculation to form an anti-interference confidence level; then, an entropy-aware dynamic threshold mechanism is introduced, and when the ρ value is lower than the set value, the judgment boundary is automatically relaxed and the sample is marked as "pending review"; when making the final decision, the system performs strict threshold judgment on samples with high ρ values, and initiates a defensive decision-making process for samples with medium and low ρ values.
[0026] The advantages of the present invention are: Active discovery and detection capabilities: The present invention can detect and warn of data leakage incidents in a timely manner before they occur by real-time monitoring of terminal screens, screenshots or photos of web application interfaces, etc., while traditional data leakage detection methods can often only detect and handle data leakage after it occurs.
[0027] Image sensitive information detection method based on multimodal large model: Calling multimodal large model, using the semantic representation ability of large-scale pre-trained model to output high-quality analysis results. Combining feature cross-fusion module and decision scoring module, it can accurately capture possible sensitive information clues and judge the results, enhancing the accuracy and generalization ability of image sensitive information detection.
[0028] Watermark information detection and tracing linkage: This invention uses watermark information detection technology to quickly locate the source of the leak, and links with the digital watermark tracing platform to achieve rapid tracing and accurate location of the source of the data leak. This innovative technology can effectively trace the person responsible for the data leak, thereby better preventing similar incidents from happening again.
[0029] Complete security protection system: This solution not only realizes the active discovery and timely warning of data leakage, but also builds a complete data security protection system with the help of DLP technology and OCR content recognition, thereby improving the efficiency and effectiveness of enterprises in dealing with data leakage problems.
[0030] In summary, the present invention solves the problem of sensitive, private or confidential data leakage faced by various units in the mobile Internet era through technical features such as active discovery and detection of data leakage, image sensitive information detection technology based on multimodal large models, watermark information detection and traceability linkage, and a complete security protection system, providing strong protection for personal privacy and corporate interests. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is an architecture diagram of a picture data leakage detection and early warning system in an embodiment of the present invention; Figure 2 This is a flow chart of watermark embedding in an embodiment of the present invention; Figure 3 This is a flow chart of a method for detecting sensitive information in an image in an embodiment of the present invention; Figure 4 This is a flowchart of image data leakage discovery and warning in an embodiment of the present invention; Figure 5 This is a screenshot of the terminal screen in the test case in the embodiment of the present invention, which contains an invisible watermark; Figure 6 For Figure 5 The traceability to the evidence collection result interface; Figure 7 This is a screenshot of a web page in a test case in an embodiment of the present invention; Figure 8 For Figure 5 The traceability to the evidence collection result interface; Fig. 9 In the test case in the embodiment of the present invention, the image leakage discovery warning system is opened, and the interface for setting the interception strategy can be clicked; Figures 10 to 12 They are respectively interception interfaces according to different sensitive words in the test cases in the embodiments of the present invention; Fig.13 This is the image detection record list interface intercepted by the network DLP in the test case in the embodiment of the present invention; Fig.14 This is the outbound image detection record list interface in the test case in the embodiment of the present invention; Fig.15 This is an interface in which the abnormal reporting record list in an embodiment of the present invention can normally generate abnormal records and display them in the list. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] The image sensitive information detection method of this embodiment focuses on improving the active discovery and detection capabilities of the government office platform data traceability security, and combines the recognition and analysis capabilities of the multimodal large model with the image detection and early warning method, which has better generalization and interpretability. Figure 1 As shown, the specific method is as follows: S1: Embed watermarks on the terminal desktop and web application interface based on the digital watermark traceability platform.
[0034] In order to quickly trace the source of data leakage to a specific terminal or responsible person when it occurs, and to improve the accuracy and efficiency of data leakage protection, this embodiment integrates the watermark capability of the digital watermark tracing platform into the government platform and other businesses that require watermark protection.
[0035] In order to ensure the normal traceability of the terminal with integrated watermark in the service, this embodiment will assign different unique bit stream sequences to it, corresponding to the unique traceability identifier, and establish a database mapping relationship between the bit stream and the assigned traceability identifier to maximize the distinction between different terminals.
[0036] The specific process is as follows: Embedding end: The watermark unit to be embedded is encrypted with a random key matrix K, which provides security on the one hand and increases the randomness of the watermark unit on the other hand. For the generated random unit watermark block, this embodiment will flip it to generate a watermark pattern with strong symmetry. Due to its strong symmetry, it has very obvious characteristics in the frequency domain. After superimposed with the material picture, it is easily filtered out by various noise filtering programs. Since the material picture is usually a high-quality picture, the Gaussian noise generated by shooting and the noise generated by compression are relatively low in intensity. Subtracting the pictures before and after the noise filtering program can separate the superimposed watermarks more intact for detection and identification.
[0037] like Figure 2 Shown: Watermark embedding process 1. Extract the brightness component of the image to be embedded Convert the image to be embedded from RGB to YCbCr space and extract the brightness component (Y channel) data.
[0038] 2. Generate basic watermark matrix The watermark unit is encrypted by a random key matrix K, and then flipped to generate a highly symmetrical watermark pattern, which is then tiled horizontally and vertically to the length and width of the image to be embedded, forming a basic watermark matrix.
[0039] 3. Calculate Adaptive Embedding Strength The basic embedding strength of the watermark is 2, and the brightness component (Y channel) is subjected to Wiener filtering. The embedding strength is slightly increased in the image area where the filtering loss exceeds a certain threshold (i.e., the area with relatively large Gaussian noise).
[0040] 4. Brightness component embedding watermark Superimpose the embedded intensity of the watermark matrix* dynamically distributed with the brightness component (Y channel).
[0041] 5. Generate embedded watermark image The superimposed image is converted back to RGB space. Since the watermark disturbance is very small, the overall visual effect is better after embedding the image.
[0042] S2: Detect the outbound image transmission behavior of the user end and intercept the outbound image.
[0043] Detect outbound screenshots or screen shots from the user end using DLP (Data Loss Prevention) image auditing technology.
[0044] Network DLP is deployed at the network exit, mainly to protect data transmitted through the Internet. The protocols supported by Network DLP include TCP session traffic such as email (SMTP), web (HTTP), file transfer (FTP), and file sharing (SMB). It monitors data in network transmission through deep content recognition technology, and performs image audit actions based on security policies defined by intelligent data classification and classification, and promptly transmits outbound images that may carry sensitive data to the image leakage warning and discovery system.
[0045] Terminal DLP mainly protects data on computer terminals. Terminal DLP monitors data transmitted through file sharing, email, Web, applications, etc. on the terminals, and performs image auditing actions based on security policies defined by intelligent data classification and grading. It promptly transmits outbound images that may carry sensitive data to the image leakage warning and detection system.
[0046] S3: Detect sensitive information on screenshots or screen shots sent externally.
[0047] This embodiment detects sensitive information in images based on the DeepSeek-R1 large language model and contradiction drive, as follows: This embodiment proposes a method for analyzing and detecting image sensitive information based on the DeepSeek large language model. Figure 3As shown in the figure, the core of this method is to make full use of the powerful understanding and analysis capabilities of the DeepSeek-R1 large language model, and combine the text and image feature fusion technology to conduct in-depth and detailed analysis of the image. First, the image to be tested is subjected to global visual feature extraction by the BLIP-2 model, and local semantic feature extraction by the DETR model to generate semantic information with strong correlation between the features of the two images. And the image is recognized by OCR with high accuracy to extract the text in the image. Then the two segments of image semantic information and the text recognized by OCR are guided by the DeepSeek-R1 large language model to generate two paragraphs of text: the first paragraph is a direct answer to the non-hypothetical question of "whether the image carries sensitive information", and the second paragraph cleverly proposes the opposite hypothetical scenario based on the conclusion of the first paragraph (for example, if the first paragraph believes that the image carries sensitive information, the second paragraph will be described and analyzed based on the assumption that "the image does not carry sensitive information"), and then input into the large language model again to generate the corresponding text output. Next, the generated forward / reverse hypothesis text pair (T1, T2) is passed through the contradiction verification and evidence synthesis module to generate a reinforced synthetic evidence vector. Finally, the evidence vector is input into the adversarial decision generation module, which outputs the final judgment result of whether the image carries sensitive information, including the probability score and the type of sensitive information. Through this complete process of "positive and negative hypothesis" reasoning, contradiction verification synthesis and adversarial decision generation combined with the DeepSeek-R1 large language model, efficient and robust detection of sensitive information in images is achieved.
[0048] Here is a description of each step: Input image and semantic information extraction: The user submits an image as the detection object. The image obtains the overall semantics and visual information of the image through the BLIP-2 model, and outputs a high-level (global) visual representation vector; through the DETR model, the important target objects in the image are identified and the corresponding local feature vectors are output; at the same time, the text in the image is recognized through high-accuracy OCR, and the corresponding text content and its position in the image are output. The global features, local features and text information extracted by OCR are combined to generate a preliminary text description of the image.
[0049] The above picture description and prompt are input into the DeepSeek-R1 large language model. The prompt is as follows: "You are an expert in detecting sensitive information in pictures: 1) Non-hypothetical prompt: Does this picture contain sensitive information? 2) Answer format: ... 3) Answer requirements: ... 4) Other requirements: ... " After receiving the picture description and prompt, the DeepSeek-R1 large language model performs a round of reasoning based on the image content and prompt requirements and gives the thinking results.
[0050] Output of DeepSeek-R1 large language model (two paragraphs of text): The first paragraph of text: Answers the "non-hypothetical" question - "Does this image carry sensitive information?".
[0051] Second paragraph of text: Based on the result of the first paragraph, an opposite hypothesis is proposed and the DeepSeek-R1 large language model is further reasoned. If the first paragraph judges that "the image does not contain sensitive information", then the second paragraph will assume that "the image contains sensitive information" and require the model to give a corresponding argument or reason; if the first paragraph judges that "the image contains sensitive information", then the second paragraph will assume that "the image does not contain sensitive information" and then let the model infer possible reasons or contexts. Finally, two formatted texts (T1, T2) are output.
[0052] Contradiction verification and evidence synthesis module: This module constructs multimodal joint evidence by quantitatively analyzing the contradiction intensity between the forward and reverse hypothesis texts (T1, T2) generated by LLM. First, semantic embedding is extracted based on the text encoder, and the logical contradiction probability output by the natural language inference model is combined to calculate the comprehensive contradiction coefficient Γ through the weighted fusion formula to dynamically balance semantic divergence and logical paradox (for example, when there are keyword conflicts and logical mutual exclusion in the text pair, the Γ value increases significantly). Subsequently, the global and local visual features are concatenated and input into the cross-modal fusion network together with the Γ value, and the gating mechanism is used to filter effective information to generate an enhanced synthetic evidence vector. This process guides visual features to focus on sensitive areas through contradictory signals. For example, when detecting nude content, even if the local features do not recognize the human body, the high contradiction coefficient will still trigger the deep analysis of the global features.
[0053] The formula for calculating the contradiction intensity is as follows:
[0054] Where: T1, T2 are the above texts, λ is the balance factor between semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model.
[0055] The evidence synthesis formula is as follows:
[0056] Where: E syn is the synthesized feature vector, which represents the final feature representation obtained by fusing global features and local features and combining the conflict strength. syn As the output of the model, it is used for subsequent tasks (such as classification or prediction).
[0057] is the visual feature fusion weight matrix, which is used to map the vector [fglobal;flocal] after the global feature and the local feature are concatenated into a k-dimensional feature space. k×2d means that this is a matrix with k rows and 2d columns, where: 2d is the dimension of the input feature, usually concatenated from two d-dimensional vectors (assuming that the dimensions of fglobal and flocal are both d). k is the dimension of the hidden layer (or the dimension of the "intermediate representation"), which indicates the size of the feature space after linear transformation. W1 maps the input 2d-dimensional vector to the k-dimensional space through linear transformation, which may be used to extract high-order interactive features of the input or reduce dimensionality. is the conflict intensity projection vector, which is used to expand the scalar Γ into a k-dimensional vector for adjusting the offset in the feature space. k×1 indicates that this is a column vector (or a single-column matrix), where: k is consistent with the output dimension of W1, ensuring matrix multiplication with the intermediate result of W1. 1 represents the dimension of the final output (usually a scalar value). W2 further maps the k-dimensional intermediate representation to a scalar output, such as for generating attention weights, classification probabilities (through Sigmoid), or regression values. global: global features, usually representing the overall information of the scene (such as scene understanding). flocal: local features, usually representing the local information of the target object (such as target detection). Through feature concatenation [fglobal;flocal], global and local information are combined to enhance the comprehensiveness of the features. The scalar Γ value is expanded to a k-dimensional weight vector, so that the high-conflict area produces a significant offset in the feature space; GELU is the Gaussian error linear unit activation function, and [;] is the feature concatenation operation.
[0058] [fglobal;flocal] is a feature concatenation operation. Function: concatenates global features and local features into a 2D vector for subsequent feature fusion. W1[fglobal;flocal] means matrix multiplication, mapping the concatenated feature vector to a k-dimensional feature space. Function: By learning the weight matrix W1, establish the association mapping between global features and local features.
[0059] W2· Γ means scalar multiplication, which expands the scalar Γ into a k-dimensional vector. Its function is to convert the conflict intensity Γ into an offset in the feature space through the weight matrix W2.
[0060] In summary, the formula The core idea is: The global features and local features are concatenated and mapped into a common feature space. The offset in the feature space is adjusted by the conflict strength Γ to highlight the features in the high conflict area. The GELU activation function is used to enhance the nonlinear expression ability of the features. Finally, the synthesized feature Esyn Can be used for more complex tasks such as classification or prediction.
[0061] Adversarial decision generation module: This module uses a three-level mechanism of "detection-correction-determination" to achieve robust decision-making. First, the Mahalanobis distance of the synthetic evidence vector is calculated, and the feature legitimacy score ρ is generated through an exponential decay function (when the sample is attacked by an adversarial attack, the ρ value drops sharply due to the distribution offset). This score directly participates in the sensitive probability calculation to form the anti-disturbance confidence. Subsequently, the entropy-aware dynamic threshold mechanism is introduced. When the ρ value is lower than the set value, the judgment boundary is automatically relaxed and the sample is marked as "pending review", which not only prevents adversarial attacks but also avoids the mistaken killing of normal content. When making the final decision, the system performs strict threshold judgment on samples with high ρ values (for example, when ρ>0.85, the confidence level>0.92 is required to be judged as sensitive), and starts a defensive decision-making process for samples with medium and low ρ values, which can significantly reduce the misjudgment rate.
[0062] The robustness verification formula is as follows:
[0063] Where: D M is the Mahalanobis distance, which is used to calculate the degree of deviation of the feature vector from the normal distribution; σ is the standard deviation of the feature distribution; ρ is the feature legitimacy score (ρ∈(0,1]ρ∈(0,1]), and w is the learnable weight vector (dimension is the same as E syn P is the probability value, which indicates the possibility that the image belongs to the sensitive category. It combines the feature legitimacy score ρ and the feature vector E syn The weighted information is finally converted into probability through the activation function σ.
[0064] The adaptive decision formula is as follows:
[0065] Where: θ0 is the basic judgment threshold; η is the sensitivity coefficient.
[0066] Output results: Based on the judgment in the previous step, the judgment results and the corresponding image analysis are output.
[0067] S4: Perform watermark detection on the screenshots or screen shots sent out. (Perform watermark detection on the pictures filtered in step S3) After the sensitive image is uploaded to the leak detection system, the watermark detection interface is called to detect whether the image contains a watermark. If there is no watermark, an alarm is directly reported. If there is a watermark, the watermark extraction interface is continued to be called to extract the watermark traceability identifier, and then the interface of the leak traceability system is called through the extracted traceability identifier to obtain the traceability object information corresponding to the traceability identifier (including ID, user name, etc.), and then this information is reported as an alarm.
[0068] S5: If a watermark is detected, the source is traced based on the watermark information and an alarm is issued.
[0069] This embodiment uses the image reporting function of network DLP and terminal DLP to report outbound images to the leak detection and warning system in real time. The leak detection and warning system screens screenshots and photos that may have data leakage risks in real time. And by detecting the watermark information of the image, it is linked with the digital watermark traceability platform to achieve rapid tracking and tracing, and accurately locate the source of data leakage. All alarm information will be summarized and reported on the security monitoring platform to timely discover and handle data leakage incidents. Among them: Real-time reporting with the help of DLP technology: With the help of network DLP and terminal DLP technology, screenshots, photos, etc. sent to the outside are reported in real time, and the identified images are provided to the data leakage detection and early warning system for content identification and tracing.
[0070] Image sensitive information detection method based on DeepSeek-R1 large language model and contradiction-driven: For images automatically identified by DLP, the system will call a large multimodal model to make full use of the semantic representation capabilities of large-scale pre-trained models and capture potential sensitive signals at the cross-modal level. The information is then input into the feature cross-fusion module to more accurately capture possible clues to sensitive information. Finally, the fused features are input into the decision scoring module to output the final judgment result of whether the image carries sensitive information, thereby enhancing the accuracy and generalization ability of image sensitive information detection.
[0071] Watermark information detection and traceability linkage: For images with sensitive information, the system can automatically identify images of potential data leakage through image sensitive information detection technology based on MLLM (multimodal large model), and quickly locate the source of leakage through watermark information detection technology. At the same time, in conjunction with the digital watermark traceability platform, the source of the leakage incident can be further traced, so as to hold the persons involved accountable and deal with them.
[0072] Alarm information summary and reporting: All audit results and alarm information will be summarized and reported on the security monitoring platform to timely discover and handle data leakage incidents. At the same time, this information can also serve as a basis for subsequent risk assessment and improvement, helping this embodiment to continuously improve the data security protection system.
[0073] Through this series of construction plans, this embodiment will realize active discovery and timely warning of data leakage, ensure data security, and reduce data leakage incidents caused by human factors. At the same time, this will also help improve users' security awareness and prevention capabilities, and promote the attention and maintenance of data security.
[0074] This method is designed based on the construction foundation of the integrated platform in terminal security, network security, and data security, such as Figure 1 As shown, it mainly includes the following three parts: 1. Watermark integration: Integrate watermarks in web applications and terminals that need to be protected to achieve invisible watermark protection for applications and desktops. This solution will help to quickly trace the source of data leakage to specific terminals when a data leakage incident occurs, thereby improving the accuracy and efficiency of data leakage protection.
[0075] 2. Real-time reporting of outbound images: Enable the real-time reporting function of outbound images in the terminal DLP and network DLP systems, monitor all outbound image data in real time according to the intelligent classification and grading rules of the data, retain the outbound images, and synchronize them to the leakage detection and early warning system in real time for leakage risk detection.
[0076] 3. Terminal leakage detection and alarm: The system will conduct in-depth content recognition on the outbound image data, and synchronize the watermark detection and watermark information tracing to the leakage detection and early warning system in real time for the images carrying sensitive information. Then, through the digital watermark tracing platform capabilities, detailed tracing results are obtained. Based on the detection results, the source of the terminal screenshots or photos in the terminal outbound data can be quickly discovered, and when abnormal outbound situations are found, they are immediately reported to the security monitoring platform for alarm.
[0077] like Figure 4 As shown, through the above method, the ability to actively discover and detect the leakage of terminal screens, web application interface screenshots or photos is built, achieving a security effect of multiple protections, security linkage, and active defense from endpoints to gateways.
[0078] The first method: Integrate PC screen watermark SDK through terminal DLP software to realize active discovery and alarm of PC terminal screen screenshots, photos and other leakage behaviors. Through the leakage discovery and early warning system, real-time detection of terminal screen, web application interface screenshots or photos of leakage content, and abnormal alarm. Link the digital watermark traceability platform to quickly track and trace the source, accurately locate the source of data leakage, and report the alarm information to the security monitoring platform for security management personnel to analyze and handle.
[0079] 1. The terminal DLP software client integrates the digital watermark PC screen watermark SDK and enables the PC screen watermark function of the digital watermark traceability platform.
[0080] 2. The terminal DLP software provides the function of reporting outbound images. The terminal DLP monitors all outbound image data in real time according to the intelligent classification and grading rules of the data, retains the outbound images, and synchronizes them to the leakage detection and early warning system in real time for leakage risk detection.
[0081] 3. The image content recognition function of the data leakage detection and early warning system, the system will call a large multimodal model, make full use of the semantic representation ability of the large-scale pre-trained model, and capture potential sensitive signals at the cross-modal level. Then the information is input into the feature cross-fusion module to more accurately capture possible sensitive information clues. Finally, the fused features are input into the decision scoring module, which outputs the final judgment result of whether the image carries sensitive information, enhancing the accuracy and generalization ability of image sensitive information detection.
[0082] 4. For images with sensitive information detected, the watermark detection and traceability evidence module will perform image watermark recognition and extraction. The detection results will be recorded and the alarm information will be reported. If the image that leaked the account information can be successfully traced, the alarm content will be supplemented with the traceability result information; if the image cannot be successfully traced, the alarm content will be supplemented with the specific reason.
[0083] The second type: The network's DLP system is deployed at the network exit and performs real-time monitoring of network exit traffic. The system retrieves and analyzes outbound image data through the image sensitive information recognition module based on MLLM, and transmits the image to the leak detection and early warning system in real time for image content recognition. Once an abnormality is found, the system will report the alarm information to the security monitoring platform through the leak detection engine for security management personnel to analyze and process.
[0084] 1. The network DLP software provides real-time image reporting function, and synchronizes the images with the leak detection and early warning system in real time.
[0085] 2. The image recognition module of the leak detection and warning system uses a multimodal large model and feature cross-fusion to identify the content of external images. Images with no risk of leakage are not processed, and images that may carry sensitive information are retained.
[0086] 3. The watermark detection and traceability evidence module performs image watermark recognition and extraction. The detection results are recorded and the alarm information is reported. If the image that leaked the account information can be successfully traced, the alarm content will supplement the traceability result information; if the image cannot be successfully traced, the alarm content will supplement the specific reason.
[0087] The test results of the above scheme are as follows: 1. PC terminal watermark embedding test Test function: PC terminal watermark embedding function.
[0088] Evaluation criteria: On a PC terminal with a watermark client installed, a watermark can be generated normally to protect the terminal desktop.
[0089] Table 1 PC terminal watermark embedding test
[0090] 2. Web page watermark embedding test Test function: PC terminal web page watermark embedding function.
[0091] Evaluation criteria: In the PC terminal web window, the watermark protection web application can be generated normally.
[0092] Table 2 Web page watermark embedding test
[0093] 3. Terminal DLP outbound image interception test Test function: Terminal DLP outgoing image interception.
[0094] Evaluation criteria: Images sent from terminals with DLP installed can be intercepted normally and sent to the image leakage detection and early warning system.
[0095] Table 3 Terminal DLP outbound image interception test
[0096] 4. Network DLP outbound image interception test Test function: Network DLP outgoing image interception.
[0097] Evaluation criteria: When images are sent via email, network DLP can intercept them and send them to the image leakage detection and early warning system.
[0098] Table 4 Network DLP outbound image interception test
[0099] 5. Keyword interception function Test function: sensitive keyword blocking setting.
[0100] Evaluation criteria: Support adding, deleting and intercepting sensitive keywords.
[0101] Table 5 Keyword interception function
[0102] 6. Content recognition of outgoing images Test function: outbound image content recognition test Evaluation criteria: Support blocking sensitive keywords according to settings Table 6 Content recognition of outgoing images
[0103] 7. Test the source of watermarks on outbound images Test function: watermark traceability test for outbound images.
[0104] Evaluation criteria: Support watermark tracing for intercepted images.
[0105] Table 7 Watermark traceability test for outbound images
[0106] 8. Abnormal reporting record test Test function: abnormal reporting and recording function test.
[0107] Evaluation criteria: Generate abnormal reporting records based on images identified by sensitive keywords and send them to the security monitoring platform.
[0108] Table 8 Abnormal reporting record test
[0109] 9. Test conclusion The system meets the design requirements for sensitive information detection in images, has complete functions, meets performance standards, and is ready for launch. It is recommended to continuously monitor after deployment in a formal environment, and regularly optimize algorithms and expand resources.
[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting sensitive information in an image, characterized in that: include: S1. Embed watermarks on the terminal desktop and web application interface; S2. Detect the outbound image behavior of the user and capture the outbound image; S3. Detect sensitive information on the outgoing image, specifically: Performing global visual feature extraction and local semantic feature extraction on the external image to generate semantic information with strong correlation between the features of the two images; Recognize and extract text from the outgoing image; The semantic information strongly associated with the features of the two pictures and the text in the external picture are guided by a large language model to generate two positive / reverse hypothesis texts and corresponding answers; The two sections of positive / reverse hypothesis text and the corresponding answers are subjected to contradiction verification and evidence synthesis to generate a reinforced synthetic evidence vector; evaluating whether the outgoing image carries sensitive information based on the synthesized evidence vector; S4. Perform watermark detection and traceability on the outbound images carrying sensitive information.
2. The method for detecting sensitive information in an image according to claim 1, characterized in that: In S3, semantic information of strong correlation between two image features is generated and text in the outgoing image is identified and extracted, specifically: The BLIP-2 model is used to obtain the overall semantics and visual information of the outgoing image, and a global visual representation vector is output; the DETR model is used to identify important target objects in the image, and a corresponding local feature vector is output; The text in the image is recognized by OCR, the corresponding text content and its position are output, the global visual representation vector, the local feature vector and the text content and its position information are combined to generate a preliminary text description of the image.
3. The method for detecting sensitive information in an image according to claim 1, characterized in that: The method for generating the two sections of forward / reverse hypothesis text and the corresponding answers in S3 is as follows: the preliminary description of the external image and the prompt word are input into the DeepSeek-R1 large language model; after receiving the preliminary description of the external image and the prompt word, the DeepSeek-R1 large language model performs a round of reasoning based on the image content and the prompt requirements, and gives a thinking result; The DeepSeek-R1 large language model outputs two paragraphs of text: The first paragraph of text: Answer the "non-hypothetical" question - "Does this image carry sensitive information?"; Second paragraph of text: Based on the results of the first paragraph, an opposite hypothesis is proposed and the DeepSeek-R1 large language model is asked to further reason. If the first paragraph judges that "the image does not contain sensitive information", then the second paragraph will assume that "the image contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph judges that "the image contains sensitive information", then the second paragraph will assume that "the image does not contain sensitive information" and then ask the DeepSeek-R1 large language model to infer possible reasons or contexts.
4. The method for detecting sensitive information in an image according to any one of claims 1 to 3, characterized in that: The contradiction verification and evidence synthesis in S3 are specifically as follows: By quantitatively analyzing the contradiction strength between the positive and negative hypothesis texts generated, multimodal joint evidence is constructed; Firstly, semantic embedding is extracted based on the text encoder, combined with the logical contradiction probability output by the natural language inference model, and the comprehensive contradiction coefficient Γ is calculated through the weighted fusion formula to dynamically balance semantic divergence and logical paradox. Subsequently, the global and local visual features are concatenated and input into the cross-modal fusion network together with the Γ value, and the gating mechanism is used to filter effective information to generate an enhanced synthetic evidence vector. The formula for calculating the contradiction intensity is as follows: Where: T1, T2 are positive and negative texts, λ is the balance factor between semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model; The evidence synthesis formula is as follows: in: is the visual feature fusion weight matrix, k×2d means that this is a matrix with k rows and 2d columns; the association mapping between global features and local features is established through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, so that the high contradiction area produces a significant shift in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the feature splicing operation.
5. The method for detecting sensitive information in an image according to claim 4, characterized in that: The specific method of evaluating whether the outgoing image carries sensitive information based on the evidence vector in S3 is: First, the Mahalanobis distance of the synthetic evidence vector is calculated, and the feature legitimacy score ρ is generated through an exponential decay function. The score ρ is directly involved in the sensitive probability calculation to form the anti-disturbance confidence; Subsequently, an entropy-aware dynamic threshold mechanism is introduced. When the score ρ value is lower than the set value, the decision boundary is automatically relaxed and the sample is marked as "pending review"; When making the final decision, the system performs strict threshold judgment on samples with high ρ values, and initiates a defensive decision-making process for samples with medium and low scoring ρ values.
6. Image sensitive information detection system, characterized in that: include: Watermark embedding module: embed watermarks into terminal desktop and web application interfaces; Detection module: detects the outgoing pictures from the user end, wherein the outgoing pictures contain watermarks; Sensitive information detection module: extract global visual features and local semantic features of the outgoing image to generate semantic information with strong correlation between the features of the two images; identify and extract text in the outgoing image; use the semantic information with strong correlation between the features of the two images and the text in the outgoing image to generate two forward / reverse hypothesis texts and corresponding answers through the guidance of a large language model; The two sections of positive / reverse hypothesis text and the corresponding answers are subjected to contradiction verification and evidence synthesis to generate a reinforced synthetic evidence vector; based on the synthetic evidence vector, whether the external image carries sensitive information is evaluated; Tracing module: performs watermark detection and tracing on the outbound images that carry sensitive information.
7. The image sensitive information detection system according to claim 6, characterized in that: The semantic information and text recognition in the sensitive information detection module are specifically as follows: the BLIP-2 model is used to obtain the overall semantics and visual information of the external image, and a global visual representation vector is output; the important target objects in the image are identified through the DETR model, and the corresponding local feature vectors are output; the text in the image is identified through OCR, and the corresponding text content and its position are output; the global visual representation vector, the local feature vector and the text content and its position information are combined to generate a preliminary text description of the image.
8. The image sensitive information detection system according to claim 6, characterized in that: The method for generating two sections of forward / reverse hypothesis text and corresponding answers in the sensitive information detection module is as follows: the preliminary description and prompt words of the external picture are input into the DeepSeek-R1 large language model; after receiving the preliminary description and prompt words of the picture, the DeepSeek-R1 large language model performs a round of reasoning according to the image content and prompt requirements, and gives the thinking result; The DeepSeek-R1 large language model outputs two paragraphs of text: The first paragraph of text: Answer the "non-hypothetical" question - "Does this image carry sensitive information?"; Second paragraph of text: Based on the results of the first paragraph, an opposite hypothesis is proposed and the DeepSeek-R1 large language model is asked to further reason. If the first paragraph judges that "the image does not contain sensitive information", then the second paragraph will assume that "the image contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph judges that "the image contains sensitive information", then the second paragraph will assume that "the image does not contain sensitive information" and then ask the DeepSeek-R1 large language model to infer possible reasons or contexts.
9. The image sensitive information detection system according to any one of claims 6 to 8, characterized in that: The contradiction verification and evidence synthesis in the sensitive information detection module are as follows: construct multimodal joint evidence by quantitatively analyzing the contradiction strength between the forward and reverse hypothesis texts generated; first, extract semantic embedding based on the text encoder, combine the logical contradiction probability output by the natural language inference model, calculate the comprehensive contradiction coefficient Γ through the weighted fusion formula, and dynamically balance semantic divergence and logical paradox; Subsequently, the global and local visual features are concatenated and input into the cross-modal fusion network together with the Γ value, and the gating mechanism is used to filter effective information to generate an enhanced synthetic evidence vector. The formula for calculating the contradiction intensity is as follows: Where: T1, T2 are positive and negative texts, λ is the balance factor between semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model; The evidence synthesis formula is as follows: in: is the visual feature fusion weight matrix, k×2d means that this is a matrix with k rows and 2d columns; the association mapping between global features and local features is established through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, so that the high contradiction area produces a significant shift in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the feature splicing operation.
10. The image sensitive information detection system according to claim 9, characterized in that: The specific method for determining whether an image carries sensitive information based on the evidence vector output in the sensitive information detection module is as follows: first, the Mahalanobis distance of the synthetic evidence vector is calculated, and a feature legitimacy score ρ is generated through an exponential decay function. The score ρ is directly involved in the sensitive probability calculation to form an anti-interference confidence level; then, an entropy-aware dynamic threshold mechanism is introduced. When the ρ value is lower than the set value, the judgment boundary is automatically relaxed and the sample is marked as "pending review"; when making the final decision, the system performs strict threshold judgment on samples with high ρ values, and initiates a defensive decision-making process for samples with medium and low ρ values.
Citation Information
Patent Citations
Fine-grained multi-modal false news detection method
CN113934882A
Law enforcement quality supervision system based on computer vision and semantic analysis
CN118898841A
DNS (Domain Name Server) data leakage detection method and system based on multi-modal time-space characteristics
CN119254524A
Multi-dimensional data authority management and privacy protection method for electric power information network
CN119538276A
Software operation detection method and device, equipment, medium and product
CN119759730A
Cited By
Web page automatic screenshot and anomaly recognition system, method and equipment based on vision-language multi-mode large model and storage medium
CN121170431A