Method and System for Detecting Sensitive Information in Images

By embedding watermarks in the terminal and web application interfaces, combining large language models and multimodal technology to detect sensitive image information, the problem of difficulty in proactively discovering and timely warning in traditional methods is solved, and the active discovery and rapid traceability of data leakage is achieved, and data security protection capabilities are improved.

CN120012056BActive Publication Date: 2025-07-08安徽省大数据中心
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510486815.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-08
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

It is difficult for the existing technology to actively discover and timely detect outgoing pictures with sensitive information. Traditional methods conduct post-tracement and responsibilities after data leakage, and their generalization capabilities are limited, so they cannot effectively warn of potential data leakage risks.

Method used

By embedding watermarks on the terminal desktop and web application interfaces, the outgoing behavior of the user's image is detected, global visual and local semantic feature extraction is performed, forward/reverse hypothetical text and contradiction verification are generated using a large language model, and multi-modal large model is used to evaluate whether the picture carries sensitive information, and watermark detection and traceability are carried out.

Benefits of technology

It realizes active discovery and timely warning of data leakage, improves detection accuracy and generalization capabilities, and can quickly locate the source of leakage and reduces losses caused by data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012056B_ABST
    Figure CN120012056B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for detecting sensitive information in pictures. By combining the existing security construction status, this method utilizes the advantages of multimodal large models to strengthen the detection of sensitive information in pictures, and constructs the ability to actively discover and detect leakage behaviors of terminal screens, screenshots of web application interfaces, or photos. It serves the suspicious leakage information discovered during the use of important data in the government affairs office platform, realizes the active discovery of data leakage, and conducts effective early warning before the leakage incident further spreads. Potential data leakage risks can be discovered in a timely manner, so as to take corresponding measures for prevention and reduce the losses caused by data leakage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security, and in particular, to a method and system for detecting sensitive information in pictures. Background Art

[0002] In the era of mobile Internet, leakage incidents of sensitive, private or confidential data of enterprises and institutions at all levels occur frequently and at a high rate, which not only brings irreparable economic losses to relevant units, but also poses great challenges to national defense security and social stability. Data security has increasingly become a key issue that enterprises and institutions at all levels are concerned about.

[0003] Government affairs business systems carry high data value and concentrated sensitive information, and need to face numerous application departments and a large number of users. The risk exposure surface is large, and the information security control is difficult. Especially in the process of using human-computer interaction data, when electronic data is converted into an information form that can be directly perceived by humans, such as converted into optical information through a display screen, it often gets out of the scope of traditional data security control, and there is a huge risk of data leakage.

[0004] Screen display is one of the essential core components of human-computer interaction. The popularization of various intelligent terminals has made secret stealing behaviors such as screen photography more convenient. At the same time, it is often difficult to trace the responsible persons for photographing and stealing secrets through traditional means. Especially in the current situation where smart phone terminals are widely used, there are major challenges in the security and confidentiality prevention of key information, and the difficulty of leakage control is increasing day by day. Currently, leakage behaviors in the form of screen shooting and screenshotting have become the management blind spots and core pain points of the security and confidentiality work of many units.

[0005] With the popularization of mobile devices and the rapid development of the Internet, the problem of data leakage has become increasingly serious. Leakage incidents of sensitive, private or confidential data of enterprises and institutions occur frequently, bringing huge challenges to personal privacy and corporate interests. The traditional solutions are, first, to install DLP data leakage prevention software at the network exit and terminal devices to audit the externally transmitted data, second, to install watermark software and anti-leakage software on the terminal devices to protect against data leakage. Third, the traditional picture sensitive information recognition technology uses methods based on deep learning to extract the convolutional features of images using networks such as CNN. The boundary of sensitive information judgment is blurred, the generalization ability is poor, and it is easily affected by image processing operations (such as compression, filtering, etc.).

[0006] Traditional DLP software mainly relies on predefined rules and patterns, which can respond to known threats, but may have limited effects on unknown threats and variants. It is impossible to accurately trace and locate the source of leakage for the leaked externally transmitted data.

[0007] Traditional watermarking product software can only trace and determine responsibilities after data leakage has occurred, and cannot effectively warn of potential data leakage risks in a timely manner. Moreover, traditional anti-disclosure software uses methods based on deep learning to extract convolutional features of images using networks such as CNN. These features are usually designed for specific steganography algorithms or specific image types, and have limited generalization ability. And usually the statistical features or underlying visual features of the image are extracted, which are often difficult to intuitively interpret and are easily affected by image processing operations (such as compression, filtering, etc.). Summary of the Invention

[0008] The technical problem to be solved by the present invention is how to actively discover and detect externally sent pictures with sensitive information.

[0009] The present invention solves the above technical problems through the following technical means:

[0010] A method for detecting sensitive information in pictures, including:

[0011] S1. Embed watermarks on the terminal desktop and web application interfaces;

[0012] S2. Detect the behavior of externally sending pictures by the user side and intercept the externally sent pictures;

[0013] S3. Detect sensitive information in the externally sent pictures, specifically:

[0014] Extract global visual features and local semantic features from the externally sent pictures to generate two segments of semantic information with strong correlation of picture features;

[0015] Identify and extract the text in the externally sent pictures;

[0016] Generate two segments of positive / negative hypothesis texts and corresponding answers by guiding the two segments of semantic information with strong correlation of picture features and the text in the externally sent pictures through a large language model;

[0017] Generate an enhanced synthetic evidence vector through contradiction verification and evidence synthesis of the two segments of positive / negative hypothesis texts and corresponding answers;

[0018] Evaluate whether the externally sent pictures carry sensitive information based on the synthetic evidence vector;

[0019] S4. Detect watermarks and trace the origin of the externally sent pictures carrying sensitive information.

[0020] By combining the existing security construction status quo, this invention utilizes the advantages of multimodal large models to strengthen the detection of sensitive information in pictures, and constructs the ability to actively discover and detect leakage behaviors of terminal screens, screenshots of web application interfaces, or photos. It serves the suspicious leakage information discovered during the use of important data on the government affairs office platform, realizes the active discovery of data leakage, and issues an effective warning before the leakage incident further spreads. Potential data leakage risks can be discovered in a timely manner, corresponding measures can be taken to prevent them, and the losses caused by data leakage can be reduced.

[0021] Further, in step S3, generating two pieces of semantic information with strong correlation of picture features and identifying and extracting the text in the externally sent picture specifically includes:

[0022] Pass the externally sent picture through the BLIP-2 model to obtain the overall semantic and visual information of the image, output the global visual representation vector, and identify the important target objects in the image through the DETR model to output the corresponding local feature vectors;

[0023] Identify the text in the image through OCR, output the corresponding text content and its location, and combine the global visual representation vector, local feature vectors, and the text content and its location information to generate a preliminary text description of the image.

[0024] Further, the method for generating two pieces of forward / backward hypothesis texts and corresponding answers in step S3 is: input the preliminary description of the externally sent picture and the prompt words into the DeepSeek-R1 large language model; after receiving the preliminary description of the externally sent picture and the prompt words, the DeepSeek-R1 large language model makes a round of reasoning according to the image content and prompt requirements and gives the thinking result;

[0025] Among them, the two pieces of text output by the DeepSeek-R1 large language model are:

[0026] The first piece of text: Answer the "non-hypothetical" question - "Does this picture carry sensitive information?"

[0027] The second piece of text: Based on the result of the first piece, put forward the opposite hypothesis and let the DeepSeek-R1 large language model make further reasoning; if the first piece judges that "this picture does not contain sensitive information", then the second piece will assume that "the picture contains sensitive information" and require the model to give the corresponding argument or reason; if the first piece judges that "this picture contains sensitive information", then the second piece will assume that "the picture does not contain sensitive information" and then let the DeepSeek-R1 large language model reason about the possible reasons or context.

[0028] Further, the contradiction verification and evidence synthesis in step S3 specifically include:

[0029] Construct multimodal joint evidence based on the contradiction intensity between the positive and negative hypothesis texts generated by quantitative analysis;

[0030] First, extract semantic embeddings based on the text encoder, combine with the logical contradiction probability output by the natural language inference model, and calculate the comprehensive contradiction coefficient Γ through the weighted fusion formula to dynamically balance semantic divergence and logical paradox;

[0031] Subsequently, after splicing the global and local visual features, input them into the cross-modal fusion network together with the Γ value, and use the gating mechanism to filter out effective information to generate an enhanced synthetic evidence vector;

[0032] The contradiction intensity calculation formula is as follows:

[0033]

[0034] Where: T1 and T2 are positive and negative texts, λ is the balance factor of semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model;

[0035] The evidence synthesis formula is as follows:

[0036]

[0037] Where: is the visual feature fusion weight matrix, k×2d indicates that this is a matrix with k rows and 2d columns; establish an association mapping between global features and local features through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector to cause a significant shift in the high-contradiction region in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal;flocal] is the feature splicing operation.

[0038] Furthermore, the specific method for evaluating whether the outgoing picture carries sensitive information based on the evidence vector in S3 is as follows:

[0039] First, calculate the Mahalanobis distance of the synthetic evidence vector, and generate a feature legality score ρ through an exponential decay function. The score ρ directly participates in the sensitive probability calculation to form an anti-interference confidence level;

[0040] Subsequently, introduce an entropy-aware dynamic threshold mechanism. When the score ρ value is lower than the set value, automatically relax the decision boundary and mark the sample as "pending review";

[0041] At the final decision, the system performs a strict threshold determination on samples with high ρ values, while initiating a defensive decision-making process for samples with medium and low ρ values.

[0042] The present invention also provides a picture sensitive information detection system, including:

[0043] A watermark embedding module: embedding watermarks on the terminal desktop and web application interfaces;

[0044] A detection module: detecting the externally transmitted pictures of the user side, where the externally transmitted pictures contain watermarks;

[0045] A sensitive information detection module: extracting global visual features and local semantic features from the externally transmitted pictures to generate two segments of semantic information with strong picture feature correlation; identifying and extracting the text in the externally transmitted pictures; generating two segments of positive / negative hypothesis texts and corresponding answers from the two segments of semantic information with strong picture feature correlation and the text in the externally transmitted pictures through a large language model; generating an enhanced synthetic evidence vector through contradiction verification and evidence synthesis for the two segments of positive / negative hypothesis texts and corresponding answers; evaluating whether the externally transmitted pictures carry sensitive information based on the synthetic evidence vector;

[0046] A traceability module: performing watermark detection and traceability on the externally transmitted pictures carrying sensitive information.

[0047] Further, the semantic information and text recognition in the sensitive information detection module are specifically as follows: the externally transmitted pictures obtain the overall semantics and visual information of the images through the BLIP-2 model, and output global visual representation vectors; identify important target objects in the images through the DETR model, and output corresponding local feature vectors; recognize the text in the images through OCR, and output the corresponding text content and its location; combine the global visual representation vectors, local feature vectors, and the text content and its location information to generate a preliminary text description of the images.

[0048] Further, the method for generating the two segments of positive / negative hypothesis texts and corresponding answers in the sensitive information detection module is: inputting the preliminary description of the externally transmitted pictures and the prompt words into the DeepSeek-R1 large language model; after receiving the preliminary description of the pictures and the prompt words, the DeepSeek-R1 large language model makes a round of reasoning according to the image content and prompt requirements, and gives the thinking result;

[0049] Among them, the two segments of text output by the DeepSeek-R1 large language model are:

[0050] The first segment of text: answering the "non-hypothetical" question - "Does this picture carry sensitive information?"

[0051] Second paragraph of text: Based on the results of the first paragraph, propose the opposite hypothesis and let the DeepSeek-R1 large language model make further inferences. If the first paragraph determines that "the picture does not contain sensitive information", then the second paragraph will assume that "the picture contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph determines that "the picture contains sensitive information", then the second paragraph will assume that "the picture does not contain sensitive information" and then let the DeepSeek-R1 large language model reason about possible reasons or contexts.

[0052] Furthermore, the contradiction verification and evidence synthesis in the sensitive information detection module are specifically as follows: By quantifying the contradiction intensity between the generated positive and negative hypothesis texts, a multi-modal joint evidence is constructed. First, semantic embeddings are extracted based on the text encoder, and combined with the logical contradiction probability output by the natural language inference model. The comprehensive contradiction coefficient Γ is calculated through a weighted fusion formula to dynamically balance semantic divergence and logical paradox. Subsequently, after concatenating the global and local visual features, they are jointly input into the cross-modal fusion network, and a gating mechanism is used to filter out valid information to generate an enhanced synthetic evidence vector.

[0053] The formula for calculating the contradiction intensity is as follows:

[0054]

[0055] where: T1, T2 are positive and negative texts, λ is the balance factor of semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model;

[0056] The formula for evidence synthesis is as follows:

[0057]

[0058] where: is the visual feature fusion weight matrix, k×2d indicates that this is a matrix with k rows and 2d columns; the association mapping between global features and local features is established through learnable parameters; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, causing significant offsets in the high-contradiction regions in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal;flocal] is the feature concatenation operation.

[0059] Furthermore, the specific method for the sensitive information detection module to output whether a picture carries sensitive information based on the evidence vector is as follows: First, calculate the Mahalanobis distance of the synthesized evidence vector, generate a feature legality score ρ through an exponential decay function, and this score ρ directly participates in the sensitive probability calculation to form an anti-interference confidence level; Subsequently, introduce an entropy-aware dynamic threshold mechanism. When the ρ value is lower than the set value, automatically relax the decision boundary and mark the sample as "pending review"; During the final decision-making, the system performs strict threshold determination on samples with high ρ values, while initiating a defensive decision-making process for samples with medium and low ρ values.

[0060] The advantages of the present invention are as follows:

[0061] Active discovery and detection capabilities: By real-time monitoring behaviors such as terminal screens, screenshots of web application interfaces, or taking pictures, the present invention can timely discover and give early warnings before data leakage events occur, while traditional data leakage detection methods often can only discover and process after the data leakage has occurred.

[0062] Picture sensitive information detection method based on a multi-modal large model: Call a multi-modal large model and utilize the semantic representation ability of the large-scale pre-trained model to output high-quality analysis results. Combine the feature cross-fusion module and the decision scoring module to accurately capture possible sensitive information clues and judge the results, enhancing the accuracy and generalization ability of image sensitive information detection.

[0063] Watermark information detection and traceability linkage: Through watermark information detection technology, the present invention can quickly locate the leakage source and link it with the digital watermark traceability platform to achieve fast tracking and tracing, accurately positioning the data leakage source. This innovative technology can effectively trace the responsible person for data leakage, thus better preventing the recurrence of similar events.

[0064] Improved security protection system: This solution not only realizes the active discovery and timely early warning of data leakage, but also constructs a complete data security protection system with the help of DLP technology and OCR content recognition, improving the response efficiency and effect of enterprise units in the face of data leakage problems.

[0065] In summary, through technical features such as active discovery and detection of data leakage, picture sensitive information detection technology based on a multi-modal large model, watermark information detection and traceability linkage, and an improved security protection system, the present invention solves the problems of leakage of sensitive, private, or classified data faced by various units in the mobile Internet era, providing strong protection for personal privacy and corporate interests. Brief Description of the Drawings

[0066] Figure 1 It is the architecture diagram of the picture data leakage discovery and early warning system in the embodiment of the present invention;

[0067] Figure 2It is the flowchart of watermark embedding in the embodiments of the present invention;

[0068] Figure 3 It is the flowchart of the method for detecting sensitive information in pictures in the embodiments of the present invention;

[0069] Figure 4 It is the flowchart of the discovery and early warning of picture data leakage in the embodiments of the present invention;

[0070] Figure 5 It is the screenshot of the terminal screen in the test case of the embodiments of the present invention, which contains a hidden watermark;

[0071] Figure 6 For Figure 5 the interface of the tracing to the forensics result;

[0072] Figure 7 It is the screenshot of the web page in the test case of the embodiments of the present invention;

[0073] Figure 8 For Figure 5 the interface of the tracing to the forensics result;

[0074] Figure 9 It is the interface where the picture leakage discovery and early warning system can be opened and the interception policy can be set in the test case of the embodiments of the present invention;

[0075] Figures 10 to 12 They are respectively the interception interfaces according to different sensitive words in the test case of the embodiments of the present invention;

[0076] Figure 13 It is the interface of the list of picture detection records intercepted by the network DLP in the test case of the embodiments of the present invention;

[0077] Figure 14 It is the interface of the list of detected records of externally sent pictures in the test case of the embodiments of the present invention;

[0078] Figure 15 It is the interface where the abnormal reporting record list can normally generate abnormal records and display them in the list in the embodiments of the present invention. Detailed implementation manners

[0079] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0080] The key point of the picture sensitive information detection method in this embodiment is to enhance the active discovery and detection capabilities of data traceability security in the government affairs office platform, and to combine the recognition and analysis capabilities of multi-modal large models in the picture detection and warning method, which has better generalization and interpretability. As Figure 1 shown, the specific method is as follows:

[0081] S1: Embed watermarks on the terminal desktop and web application interfaces according to the digital watermark traceability platform.

[0082] In order to quickly trace back to a specific terminal or responsible person in the event of a data leakage incident and improve the accuracy and efficiency of data leakage protection, this embodiment integrates the watermark capabilities of the digital watermark traceability platform for the government affairs platform and other services that require watermark protection.

[0083] In order to ensure the normal traceability of the terminals with integrated watermarks in the service, this embodiment will assign different unique bitstream sequences to them, corresponding to unique traceability identifiers, and establish a database mapping relationship between the bitstreams and the assigned traceability identifiers to ensure the maximum distinction between different terminals.

[0084] The specific process is as follows:

[0085] Embedding end: The watermark unit to be embedded is encrypted with a random key matrix K, which provides security on the one hand and increases the randomness of the watermark unit on the other hand. For the generated random unit watermark block, this embodiment will flip it to generate a watermark pattern with strong symmetry. Due to the strong symmetry, it has very obvious characteristics in the frequency domain and is easily filtered out by various noise filtering programs after being superimposed on the source picture. Since the source picture is usually a high-quality picture and the Gaussian noise generated by shooting and the noise intensity generated by compression are low, subtracting the pictures before and after the noise filtering program processing can separate the superimposed watermark relatively intact for detection and recognition.

[0086] As Figure 2 shown: Watermark embedding process

[0087] 1. Extract the luminance component of the picture to be embedded

[0088] Convert the picture to be embedded from RGB to the YCbCr space and extract the luminance component (Y channel) data.

[0089] 2. Generate the basic watermark matrix

[0090] After the watermark unit is encrypted with the random key matrix K and flipped, a watermark pattern with strong symmetry is generated, and then it is tiled horizontally and vertically to the length and width of the picture to be embedded to form the basic watermark matrix.

[0091] 3. Calculate the adaptive embedding strength

[0092] The basic embedding strength of the watermark is 2. Wiener filtering is performed on the luminance component (Y channel), and the embedding strength is slightly increased in the picture areas where the filtering loss exceeds a certain threshold (i.e., the areas with relatively large Gaussian noise).

[0093] 4. Embed the watermark in the luminance component

[0094] Overlay the watermark matrix * with the dynamically distributed embedding strength on the luminance component (Y channel).

[0095] 5. Generate the picture with the embedded watermark

[0096] Convert the overlaid picture back to the RGB space. Since the watermark perturbation is very small, the overall visual effect is good after embedding the watermark into the picture.

[0097] S2: Detect the picture forwarding behavior of the user side and intercept the forwarded picture.

[0098] Detect the forwarded screenshot or screen capture picture of the user side, which is implemented using DLP (Data Loss Prevention) picture auditing technology.

[0099] Network DLP is deployed at the network exit and mainly protects the data transmitted through the Internet. The protocols supported by network DLP include TCP session traffic such as email (SMTP), web (HTTP), file transfer (FTP), file sharing (SMB), etc. Monitor the data in network transmission through deep content recognition technology, and perform picture auditing actions according to the security policies defined by intelligent data classification and grading, and timely transmit the forwarded pictures that may carry sensitive data to the picture leakage warning discovery system.

[0100] Terminal DLP mainly protects the data of the computer terminal. Terminal DLP monitors the data transmitted by file sharing, email, web, application programs, etc. on the terminal, and performs picture auditing actions according to the security policies defined by intelligent data classification and grading, and timely transmits the forwarded pictures that may carry sensitive data to the picture leakage warning discovery system.

[0101] S3: Detect sensitive information in the forwarded screenshot or screen capture picture.

[0102] In this embodiment, picture sensitive information detection is performed based on the DeepSeek-R1 large language model and contradiction drive, as follows:

[0103] This embodiment proposes a method for analyzing and detecting picture sensitive information based on the DeepSeek large language model, such as Figure 3As shown, the core of this method lies in making full use of the powerful understanding and analysis capabilities of the DeepSeek-R1 large language model, combining text and image feature fusion technology to conduct in-depth and detailed analysis of pictures. First, the picture to be tested is passed through the BLIP-2 model for global visual feature extraction, and at the same time through the DETR model for local semantic feature extraction, generating two pieces of semantic information with strong correlation of picture features. And the picture is recognized by highly accurate OCR to extract the text in the picture. Then, the two pieces of picture semantic information and the text after OCR recognition are guided by the DeepSeek-R1 large language model to generate two pieces of text: the first paragraph is a direct answer based on the non-hypothetical question of "whether the picture carries sensitive information", and the second paragraph cleverly proposes the opposite hypothetical situation according to the conclusion of the first paragraph (for example, if the first paragraph believes that the picture carries sensitive information, the second paragraph will be described and analyzed based on the assumption of "the picture does not carry sensitive information"), and is input into the large language model again to generate the corresponding text output. Next, the generated positive / negative hypothetical text pair (T1, T2) is passed through the contradiction verification and evidence synthesis module to generate an enhanced synthetic evidence vector. Finally, the evidence vector is input into the adversarial decision-making generation module to output the final judgment result of whether the picture carries sensitive information, including the possibility score and the type of sensitive information. Through this complete process of "positive and negative hypothesis" reasoning, contradiction verification synthesis, and adversarial decision-making generation combined with the DeepSeek-R1 large language model, the efficient and robust detection of sensitive information in pictures is achieved.

[0104] The following is an introduction to each step:

[0105] Input picture and semantic information extraction: The user submits an image as the detection object. This image obtains the overall semantic and visual information of the image through the BLIP-2 model, and outputs a high-level (global) visual representation vector; through the DETR model, important target objects in the image are recognized, and the corresponding local feature vectors are output; at the same time, the text in the image is recognized by high-accuracy OCR, and the corresponding text content and its position in the image are output. The global features, local features, and text information extracted by OCR are combined to generate a preliminary text description of the image.

[0106] Input the above picture description and prompt into the DeepSeek-R1 large language model. The prompt is as follows: "You are an expert in detecting sensitive information in pictures: 1) Non-hypothetical prompt: Does this picture carry sensitive information? 2) Answer format:...... 3) Answer requirements:...... 4) Other requirements:......"

[0107] After receiving the picture description and prompt, the DeepSeek-R1 large language model makes a round of reasoning based on the image content and prompt requirements and gives the thinking result.

[0108] Output of the DeepSeek-R1 large language model (two texts): The first text: Answer a "non-hypothetical" question - "Does the picture carry sensitive information?".

[0109] The second text: Based on the result of the first paragraph, propose the opposite hypothesis and let the DeepSeek-R1 large language model make further inferences. If the first paragraph determines that "the picture does not contain sensitive information", then the second paragraph will assume that "the picture contains sensitive information" and require the model to give corresponding arguments or reasons; if the first paragraph determines that "the picture contains sensitive information", then the second paragraph will assume that "the picture does not contain sensitive information" and then let the model reason about possible reasons or contexts. Finally, two formatted texts (T1, T2) are output.

[0110] Contradiction verification and evidence synthesis module: This module constructs multi-modal joint evidence by quantitatively analyzing the contradiction intensity between the positive and negative hypothesis texts (T1, T2) generated by the LLM. First, based on the text encoder, semantic embeddings are extracted, and combined with the logical contradiction probability output by the natural language inference model, the comprehensive contradiction coefficient Γ is calculated through a weighted fusion formula to dynamically balance semantic divergence and logical paradox (for example, when there are both keyword conflicts and logical exclusions in the text pair, the value of Γ increases significantly). Subsequently, after splicing the global and local visual features, they are jointly input into the cross-modal fusion network, and the gating mechanism is used to screen out effective information to generate an enhanced synthetic evidence vector. This process guides the visual features to focus on sensitive areas through contradiction signals. For example, when detecting nude content, even if the local features do not recognize a human body, a high contradiction coefficient will still trigger in-depth analysis of the global features.

[0111] The formula for calculating the contradiction intensity is as follows:

[0112]

[0113] Where: T1, T2 are the above texts, λ is the balance factor of semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model.

[0114] The formula for evidence synthesis is as follows:

[0115]

[0116] Where: E syn is the synthesized feature vector, representing the final feature representation obtained by fusing the global and local features and combining the conflict intensity. E syn As the output of the model, it is used for subsequent tasks (such as classification or prediction).

[0117] is the visual feature fusion weight matrix, which is used to map the vector [fglobal; flocal] after concatenating the global feature and the local feature to a k-dimensional feature space. k×2d indicates that this is a matrix with k rows and 2d columns, where: 2d is the dimension of the input feature, usually formed by concatenating two d-dimensional vectors (assuming that the dimensions of fglobal and flocal are both d). k is the dimension of the hidden layer (or the dimension of the "intermediate representation"), representing the size of the feature space after linear transformation. W1 maps the input 2d-dimensional vector to the k-dimensional space through linear transformation, which may be used to extract the high-order interaction features of the input or for dimensionality reduction. is the conflict intensity projection vector, which is used to expand the scalar Γ into a k-dimensional vector for adjusting the offset in the feature space. k×1 indicates that this is a column vector (or a single-column matrix), where: k is consistent with the output dimension of W1 to ensure matrix multiplication with the intermediate result of W1. 1 represents the dimension of the final output (usually a scalar value). W2 further maps the k-dimensional intermediate representation to a scalar output, such as for generating attention weights, classification probabilities (through Sigmoid), or regression values. global: global feature, usually representing the overall information of the scene (such as scene understanding). flocal: local feature, usually representing the local information of the target object (such as object detection). By concatenating the features [fglobal; flocal], the global and local information are combined to enhance the comprehensiveness of the features. Expanding the scalar Γ value into a k-dimensional weight vector causes significant offsets in the high-conflict regions in the feature space; GELU is the Gaussian error linear unit activation function, and [;] is the feature concatenation operation.

[0118] [fglobal; flocal] is the feature concatenation operation. Function: Concatenate the global feature and the local feature into a 2d-dimensional vector for subsequent feature fusion. The meaning of W1[fglobal; flocal] is matrix multiplication, which maps the concatenated feature vector to the k-dimensional feature space. Function: Establish the correlation mapping between the global feature and the local feature by learning the weight matrix W1.

[0119] The meaning of W2· Γ is scalar multiplication, which expands the scalar Γ into a k-dimensional vector. Its function is to convert the conflict intensity Γ into an offset in the feature space through the weight matrix W2.

[0120] In summary, the formula has the core idea that:

[0121] Concatenate the global features and local features and map them to a common feature space. Adjust the offset in the feature space through the conflict intensity Γ to highlight the features in the high-conflict regions. Use the GELU activation function to enhance the non-linear expression ability of the features. Finally, the synthesized feature E syn can be used for more complex tasks such as classification or prediction.

[0122] Adversarial decision-making generation module: This module adopts a three-level mechanism of "detection - correction - determination" to achieve robust decision-making. First, calculate the Mahalanobis distance of the synthesized evidence vector, and generate a feature legality score ρ through an exponential decay function (when the sample is under adversarial attack, the ρ value drops sharply due to distribution shift), and this score directly participates in the sensitive probability calculation to form the anti-interference confidence. Subsequently, introduce an entropy-aware dynamic threshold mechanism. When the ρ value is lower than the set value, automatically relax the determination boundary and mark the sample as "pending review", which not only guards against adversarial attacks but also avoids misclassifying normal content. During the final decision-making, the system performs a strict threshold determination on samples with high ρ values (for example, when ρ > 0.85, the confidence level needs to be > 0.92 to be judged as sensitive), while for samples with medium and low ρ values, start the defensive decision-making process, which can significantly reduce the misjudgment rate.

[0123] The robustness verification formula is as follows:

[0124]

[0125] where: D M is the Mahalanobis distance, which is used to calculate the deviation degree of the feature vector from the normal distribution; σ is the standard deviation of the feature distribution; ρ is the feature legality score (ρ ∈ (0, 1]), and w is the learnable weight vector (with the same dimension as E syn ). P is the probability value, indicating the possibility that the picture belongs to the sensitive category. It combines the weighted information of the feature legality score ρ and the feature vector E syn and is finally converted into a probability through the activation function σ.

[0126] The adaptive decision formula is as follows:

[0127]

[0128] where: θ0 is the basic determination threshold; η is the sensitivity coefficient.

[0129] Output result: According to the judgment of the previous step, output the judgment result and the corresponding picture analysis.

[0130] S4: Perform watermark detection on the externally sent screenshots or screen-captured pictures. (Perform watermark detection on the pictures screened in step S3)

[0131] After the sensitive picture is uploaded to the leakage detection system, the watermark detection interface is called to detect whether the picture contains a watermark. If there is no watermark, an alarm is directly reported. If there is a watermark, the watermark extraction interface is continuously called to extract the watermark traceability identifier. Then, through the extracted traceability identifier, the interface of the leakage traceability system is called to obtain the traceability object information corresponding to the traceability identifier (including information such as ID and username), and then these information are reported as an alarm.

[0132] S5: If a watermark is detected, traceability is performed based on the watermark information, and alarm processing is carried out.

[0133] In this embodiment, the picture outbound reporting function of network DLP and terminal DLP is used to report outbound pictures to the leakage discovery and warning system in real time. The leakage discovery and warning system screens screenshots and photographed pictures that may have the risk of data leakage in real time. And through the detection of the picture watermark information, it is linked with the digital watermark traceability platform to achieve fast traceability and accurately locate the source of data leakage. All alarm information will be summarized and reported on the security monitoring platform to detect and handle data leakage incidents in a timely manner. Among them:

[0134] Real-time reporting with DLP technology: With network DLP and terminal DLP technologies, outbound screenshots, photographed pictures, etc. are reported in real time, and the identified pictures are provided to the data leakage discovery and warning system for content identification and traceability.

[0135] Picture sensitive information detection method based on DeepSeek-R1 large language model and contradiction-driven: For the pictures automatically identified by DLP, the system will call a multi-modal large model, make full use of the semantic representation ability of the large-scale pre-trained model, and capture potential sensitive signals at the cross-modal level. Then the information is input into the feature cross-fusion module to more accurately capture possible sensitive information clues. Finally, the fused features are input into the decision scoring module to output the final judgment result of whether the picture carries sensitive information, enhancing the accuracy and generalization ability of image sensitive information detection.

[0136] Linkage between watermark information detection and traceability: For the identified pictures containing sensitive information, through the picture sensitive information detection technology based on MLLM (multi-modal large model), the system can automatically identify pictures of potential data leakage behaviors. Through the watermark information detection technology, the leakage source can be quickly located. At the same time, by linking with the digital watermark traceability platform, the source of the leakage event can be further traced, so as to hold the involved personnel accountable and handle them.

[0137] Alarm Information Aggregation and Reporting: All audit results and alarm information will be aggregated and reported on the security monitoring platform to promptly detect and handle data leakage incidents. Meanwhile, this information can also serve as a basis for subsequent risk assessment and improvement, helping this embodiment continuously improve the data security protection system.

[0138] Through this series of construction plans, this embodiment will achieve the active discovery and timely warning of data leakage, ensure data security, and reduce data leakage incidents caused by human factors. Meanwhile, this will also help improve the security awareness and prevention capabilities of users, and promote the attention and maintenance of data security.

[0139] This method is designed based on the construction foundation of the integrated platform in terms of terminal security, network security, and data security, as Figure 1 shown, mainly including the following three parts:

[0140] 1. Watermark Integration: Integrate watermarks in the web applications and terminals that need to be protected to achieve invisible watermark protection for the applications and desktops. This solution will help quickly trace back to the specific terminal in case of a data leakage incident, improving the accuracy and efficiency of data leakage protection.

[0141] 2. Real-time Reporting of Outgoing Pictures: Enable the real-time reporting function of outgoing pictures in the terminal DLP and network DLP systems, and conduct real-time monitoring of all outgoing picture data according to the intelligent classification and grading rules of data, retain the outgoing pictures, and synchronize them to the leakage discovery and warning system in real time for leakage risk detection.

[0142] 3. Terminal Leakage Discovery and Alarm: The system will conduct in-depth content recognition of the outgoing picture data. For pictures carrying sensitive information, it will be synchronized to the leakage discovery and warning system in real time for watermark detection and watermark information tracing. Subsequently, through the capabilities of the digital watermark tracing platform, detailed tracing results will be obtained. According to the detection results, quickly discover the source of the terminal screen capture or photo in the terminal's outgoing data, and immediately report it to the security monitoring platform for alarm when abnormal outgoing situations are found.

[0143] As Figure 4 shown, through the above methods, build the ability to actively discover and detect the leakage behaviors of terminal screens, web application interface screenshots, or photos, and achieve the security effects of multiple protections, security linkage, and active defense from endpoints to gateways.

[0144] The first method: Integrate the PC screen watermark SDK through the terminal DLP software to actively detect and alarm leakage behaviors such as PC terminal screen capture and photographing. Through the leakage discovery and early warning system, it can detect the content of screenshots or photographs of the terminal screen and web application interfaces in real time and give abnormal alarms. It is linked with the digital watermark traceability platform to quickly track and trace the source, accurately locate the source of data leakage, and report the alarm information to the security monitoring platform for security managers to analyze and process.

[0145] 1. The client of the terminal DLP software integrates the digital watermark PC screen watermark SDK and enables the PC screen watermark function of the digital watermark traceability platform.

[0146] 2. The terminal DLP software provides a function for reporting externally sent pictures. The terminal DLP monitors all externally sent picture data in real time according to the intelligent classification and grading rules of the data, retains the externally sent pictures, and synchronizes them to the leakage discovery and early warning system in real time for leakage risk detection.

[0147] 3. For the picture content recognition function of the data leakage discovery and early warning system, the system will call a multi-modal large model, make full use of the semantic representation ability of the large-scale pre-trained model, and capture potential sensitive signals at the cross-modal level. Then the information is input into the feature cross-fusion module to more accurately capture possible sensitive information clues. Finally, the fused features are input into the decision scoring module to output the final judgment result of whether the picture carries sensitive information, enhancing the accuracy and generalization ability of image sensitive information detection.

[0148] 4. For the pictures that detect sensitive information, the watermark detection and traceability forensics module performs picture watermark recognition and extraction. Record the detection results and report the alarm information. Pictures for which the leaked account information can be successfully traced will have the traceability results supplemented in the alarm content; for pictures for which tracing is not successful, the specific reasons will be supplemented in the alarm content.

[0149] The second method: The DLP system of the network is deployed at the network exit and monitors the network exit traffic in real time. This system retrieves and analyzes the externally sent picture data through the picture sensitive information recognition module based on MLLM and transfers the pictures to the picture content recognition of the leakage discovery and early warning system in real time. Once an anomaly is found, the system will report the alarm information to the security monitoring platform through the leakage discovery engine for security managers to analyze and process.

[0150] 1. The network DLP software provides a function for real-time reporting of pictures and synchronizes the externally sent pictures with the leakage discovery and early warning system in real time.

[0151] 2. The picture recognition module of the leakage discovery and early warning system performs content recognition on the externally sent pictures through methods such as multi-modal large models and feature cross-fusion. Pictures without leakage risks are not processed, and pictures that may carry sensitive information are retained.

[0152] 3. The watermark detection and traceability evidence collection module performs image watermark recognition and extraction. Records the detection results and reports the alarm information. It can successfully trace the image that leaks account information, and the alarm content supplements the traceability result information; for images that cannot be successfully traced, the alarm content supplements the specific reasons.

[0153] The test results of the above solutions are as follows:

[0154] 1. PC terminal watermark embedding test

[0155] Test function: PC terminal watermark embedding function.

[0156] Evaluation criterion: On a PC terminal installed with the watermark client, the watermark can be normally generated to protect the terminal desktop.

[0157] Table 1 PC terminal watermark embedding test

[0158]

[0159] 2. Web page watermark embedding test

[0160] Test function: PC terminal web page watermark embedding function.

[0161] Evaluation criterion: In the web page window of the PC terminal, the watermark can be normally generated to protect the web application.

[0162] Table 2 Web page watermark embedding test

[0163]

[0164] 3. Terminal DLP outgoing picture interception test

[0165] Test function: Terminal DLP outgoing picture interception.

[0166] Evaluation criterion: Pictures sent from a terminal installed with DLP can be normally intercepted and sent to the picture leakage discovery warning system.

[0167] Table 3 Terminal DLP outgoing picture interception test

[0168]

[0169] 4. Network DLP outgoing picture interception test

[0170] Test function: Network DLP outgoing picture interception.

[0171] Evaluation criterion: When sending pictures via email, the network DLP can normally intercept and send them to the picture leakage discovery warning system.

[0172] Table 4 Network DLP Outbound Image Interception Test

[0173]

[0174] 5. Interception Keyword Function

[0175] Test Function: Sensitive Keyword Interception Setting

[0176] Evaluation Criteria: Support adding and deleting intercepted sensitive keywords

[0177] Table 5 Interception Keyword Function

[0178]

[0179] 6. Outbound Image Content Recognition

[0180] Test Function: Outbound Image Content Recognition Test

[0181] Evaluation Criteria: Support intercepting sensitive keywords according to settings

[0182] Table 6 Outbound Image Content Recognition

[0183]

[0184] 7. Outbound Image Watermark Traceability Test

[0185] Test Function: Outbound Image Watermark Traceability Test

[0186] Evaluation Criteria: Support watermark traceability for intercepted images

[0187] Table 7 Outbound Image Watermark Traceability Test

[0188]

[0189] 8. Abnormal Report Record Test

[0190] Test Function: Abnormal Report Record Function Test

[0191] Evaluation Criteria: Generate abnormal report records for images identified by sensitive keywords and send them to the security monitoring platform

[0192] Table 8 Abnormal Report Record Test

[0193]

[0194] 9. Test Conclusion

[0195] The system meets the design requirements for detecting sensitive information in images, with complete functions and qualified performance, and is ready for online deployment. It is recommended to continuously monitor after deployment in the formal environment and regularly optimize the algorithm and expand resources

[0196] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for detecting sensitive information in images, characterized in that, Including: S1. Embed watermarks on the terminal desktop and web application interfaces; S2. Detect the behavior of the user side to send out pictures and intercept the sent-out pictures; S3. Detect sensitive information in the sent-out pictures, specifically: Extract global visual features and local semantic features from the sent-out pictures to generate two pieces of semantic information with strong correlation of picture features; Identify and extract the text in the sent-out pictures. Specifically: Pass the sent-out pictures through the BLIP-2 model to obtain the overall semantics and visual information of the images, output the global visual representation vectors, and identify the important target objects in the images through the DETR model to output the corresponding local feature vectors; Recognize the text in the images through OCR, output the corresponding text content and its location, and combine the global visual representation vectors, local feature vectors, and the text content and its location information to generate a preliminary text description of the images; Generate two pieces of positive / negative hypothesis texts and corresponding answers by guiding the two pieces of semantic information with strong correlation of picture features and the text in the sent-out pictures through a large language model; Generate an enhanced synthetic evidence vector through contradiction verification and evidence synthesis of the two pieces of positive / negative hypothesis texts and corresponding answers; Evaluate whether the sent-out pictures carry sensitive information based on the synthetic evidence vector; S4. Conduct watermark detection and traceability on the sent-out pictures carrying sensitive information.

2. The method for detecting picture sensitive information according to claim 1, wherein The method for generating the two pieces of positive / negative hypothesis texts and corresponding answers in S3 is: Input the preliminary description of the sent-out pictures and the prompt words into the DeepSeek-R1 large language model; After receiving the preliminary description of the sent-out pictures and the prompt words, the DeepSeek-R1 large language model makes a round of reasoning according to the image content and prompt requirements and gives the thinking results; Among them, the two pieces of texts output by the DeepSeek-R1 large language model are: The first piece of text: Answer the "non-hypothetical" question - "Does this picture carry sensitive information?"; The second piece of text: Based on the result of the first piece, put forward the opposite hypothesis and let the DeepSeek-R1 large language model make further reasoning; If the first piece judges that "this picture does not contain sensitive information", then the second piece will assume that "the picture contains sensitive information" and require the model to give the corresponding argument or reason; If the first piece judges that "this picture contains sensitive information", then the second piece will assume that "the picture does not contain sensitive information" and then let the DeepSeek-R1 large language model reason about the possible reasons or context.

3. The method for detecting picture sensitive information according to claim 1 or 2, characterized in that The contradiction verification and evidence synthesis in S3 are specifically: Construct a multi-modal joint evidence by quantitatively analyzing the contradiction intensity between the generated positive and negative hypothesis texts; First, extract semantic embeddings based on the text encoder, and combine the logical contradiction probability output by the natural language inference model to calculate the comprehensive contradiction coefficient Γ through a weighted fusion formula to dynamically balance semantic divergence and logical paradox; Subsequently, splice the global and local visual features and input them into the cross-modal fusion network together with the Γ value, and use the gating mechanism to filter out effective information to generate an enhanced synthetic evidence vector; The contradiction intensity calculation formula is as follows: Where: T1 and T2 are positive and negative texts, λ is the balance factor of semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model; The evidence synthesis formula is as follows: Wherein: is the visual feature fusion weight matrix, where k×2d indicates that this is a matrix with k rows and 2d columns; the association mapping between the global feature and the local feature is established by learning the weight matrix W1; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, causing a significant shift in the high-contradiction region in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the concatenation of the global feature and the local feature.

4. The picture sensitive information detection method according to claim 3, characterized in that The specific method for evaluating whether the sent-out picture carries sensitive information based on the evidence vector in S3 is as follows: First, calculate the Mahalanobis distance of the synthesized evidence vector, and generate a feature legality score ρ through an exponential decay function. The score ρ directly participates in the sensitive probability calculation to form an anti-interference confidence level; Subsequently, introduce an entropy-aware dynamic threshold mechanism. When the score ρ value is lower than the set value, automatically relax the decision boundary and mark the sample as "pending review"; At the final decision-making, the system performs a strict threshold determination on samples with high ρ values, while initiating a defensive decision-making process for samples with medium and low score ρ values.

5. Image sensitive information detection system, characterized in that, Including: Watermark embedding module: Embed watermarks in the terminal desktop and web application interfaces; Detection module: Detect the sent-out pictures of the user side, where the sent-out pictures contain watermarks; Sensitive information detection module: Extract global visual features and local semantic features from the sent-out pictures to generate two segments of semantic information with strong association of picture features; Identify and extract the text in the sent-out pictures; Specifically: The sent-out pictures obtain the overall semantic and visual information of the image through the BLIP-2 model and output the global visual representation vector; Identify important target objects in the image through the DETR model and output the corresponding local feature vectors; Recognize the text in the image through OCR and output the corresponding text content and its location; Combine the global visual representation vector, local feature vector, and the text content and its location information to generate a preliminary text description of the image; Generate two segments of positive / negative hypothesis texts and corresponding answers from the two segments of semantic information with strong association of picture features and the text in the sent-out pictures through the guidance of a large language model; Generate an enhanced synthesized evidence vector from the two segments of positive / negative hypothesis texts and corresponding answers through contradiction verification and evidence synthesis; Evaluate whether the sent-out pictures carry sensitive information based on the synthesized evidence vector; Traceability module: Perform watermark detection and traceability on the sent-out pictures carrying sensitive information.

6. The picture sensitive information detection system according to claim 5, wherein The method for generating two segments of positive / negative hypothesis texts and corresponding answers in the sensitive information detection module is as follows: Input the preliminary description of the sent-out pictures and the prompt words into the DeepSeek-R1 large language model; After receiving the preliminary description of the pictures and the prompt words, the DeepSeek-R1 large language model makes a round of reasoning according to the image content and the prompt requirements and gives the thinking result; Among them, the two segments of texts output by the DeepSeek-R1 large language model are: The first segment of text: Answer the "non-hypothetical" question - "Does this picture carry sensitive information?"; Second paragraph of text: Based on the results of the first paragraph, propose the opposite hypothesis and let the DeepSeek-R1 large language model further reason. If the first paragraph determines that "the picture does not contain sensitive information", then the second paragraph will assume that "the picture contains sensitive information" and require the model to give corresponding arguments or reasons. If the first paragraph determines that "the picture contains sensitive information", then the second paragraph will assume that "the picture does not contain sensitive information" and then let the DeepSeek-R1 large language model reason about possible reasons or contexts.

7. The picture sensitive information detection system according to claim 5 or 6, characterized in that The specific methods of contradiction verification and evidence synthesis in the sensitive information detection module are as follows: By quantifying the contradiction intensity between the generated positive and negative hypothesis texts, a multi-modal joint evidence is constructed. First, based on the text encoder, semantic embeddings are extracted, and combined with the logical contradiction probability output by the natural language inference model. Through a weighted fusion formula, the comprehensive contradiction coefficient Γ is calculated to dynamically balance semantic divergence and logical paradox. Subsequently, after splicing the global and local visual features, they are jointly input into the cross-modal fusion network with the Γ value. The gating mechanism is used to filter out effective information and generate an enhanced synthetic evidence vector. The contradiction intensity calculation formula is as follows: Where: T1, T2 are the positive and negative texts, λ is the balance factor of semantic and logical contradictions, E is the text encoder, d is the text embedding dimension, and P(·) is the contradiction / implication probability output by the natural language inference model. The evidence synthesis formula is as follows: Wherein: is the visual feature fusion weight matrix, where k×2d indicates that this is a matrix with k rows and 2d columns; the correlation mapping between the global feature and the local feature is established by learning the weight matrix W1; is the contradiction intensity projection vector, which expands the scalar Γ value into a k-dimensional weight vector, causing a significant shift in the high-contradiction region in the feature space; GELU is the Gaussian error linear unit activation function, and [fglobal; flocal] is the operation of concatenating the global feature and the local feature.

8. The picture sensitive information detection system according to claim 7, characterized in that, The specific method for the sensitive information detection module to output whether the picture carries sensitive information based on the evidence vector is as follows: First, calculate the Mahalanobis distance of the synthetic evidence vector, and generate a feature legality score ρ through an exponential decay function. The score ρ directly participates in the sensitive probability calculation to form an anti-interference confidence level. Subsequently, an entropy-aware dynamic threshold mechanism is introduced. When the ρ value is lower than the set value, the decision boundary is automatically relaxed and the sample is marked as "pending review". At the final decision, the system performs a strict threshold determination on samples with high ρ values, while initiating a defensive decision-making process for samples with medium and low ρ values.

Citation Information

Patent Citations

  • Fine-grained multi-modal false news detection method

    CN113934882A

  • Multi-dimensional data authority management and privacy protection method for electric power information network

    CN119538276A

  • Software operation detection method and device, equipment, medium and product

    CN119759730A