Document tracing method and device based on dark watermark, computer equipment and medium

By embedding dark watermark encrypted strings in the document, the problem of difficult to trace after document leakage in the prior art is solved, and effective traceability and security protection for document leakage is achieved.

CN120012055AInactive Publication Date: 2025-05-16SHENZHEN TIANYUAN DIC INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510485086.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing business system has technical problems such as high risk of sensitive data leakage and difficult to trace after leakage in document download or data export functions.

Method used

The document traceability method based on dark watermark is adopted. By determining the watermark embedding location of the target document, a dark watermark encrypted string containing traceability information is generated, and embedded in the target document, thereby realizing the traceability of the leaked document.

Benefits of technology

Improve the security of documents, and make it difficult for attackers to detect and remove watermarks through the concealment and invisibility of dark watermarks, effectively prevent sensitive data leakage and accurately trace back to the source of the leakage after leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012055A_ABST
    Figure CN120012055A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information security, and provides a document tracing method and device based on a dark watermark, computer equipment and a medium. The method comprises the steps of determining a watermark embedding position of a target document; generating a dark watermark encrypted character string containing traceability information; embedding the dark watermark encrypted character string into the target document based on the watermark embedding position to obtain a watermark document; and carrying out reverse analysis on the watermark document so as to realize leakage traceability of the target document. According to the method and the device, the target document can be effectively identified and tracked after being leaked by embedding the dark watermark encrypted character string, so that the security of the document is improved. And due to the concealment and invisibility of the dark watermark, an attacker is difficult to perceive and remove the watermark, and the protection strength of the document is further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to a document tracing method, device, computer equipment and medium based on dark watermark. Background Art

[0002] In various business systems, the functions of downloading or exporting documents (such as WORD, PPT, EXCEL, and PDF) are very common, and these functions provide great convenience for users. However, if these functions are not protected, the risk of sensitive data leakage will increase significantly as documents are widely shared and forwarded, and once leaked, it is often difficult to trace its source.

[0003] Existing business systems provide security protection by adding visible watermarks to downloaded or exported documents. Although this approach has a deterrent effect to a certain extent, reminding users not to share or leak documents at will, the protection effect of visible watermarks is limited because visible watermarks are easy to tamper with or remove. Once a document is leaked, it is difficult to effectively trace its source and transmission path based on visible watermarks alone. Summary of the invention

[0004] Based on this, a document tracing method, device, computer equipment and medium based on dark watermark are proposed, aiming to solve the technical problems of high risk of sensitive data leakage and difficulty in tracing after leakage in document download or data export functions in existing business systems.

[0005] The first aspect of the present application provides a document tracing method based on dark watermark, the method comprising: Determine the watermark embedding position of the target document; Generate a dark watermark encrypted string containing traceability information; Embedding the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; The watermark document is reversely parsed to trace the leakage of the target document.

[0006] Optionally, generating a dark watermark encrypted string containing traceability information includes: Obtaining the operation time, operation object information and client information of the target document; Generate a combined string in a target format based on the operation time, operation object information and client information; The combined character string in the target format is encrypted to obtain a dark watermark encrypted character string.

[0007] Optionally, determining the watermark embedding position of the target document includes: Obtaining the media type of the target document; Identifying a document type of the target document according to the media type; The watermark embedding position of the target document is determined according to the document type.

[0008] Optionally, determining the watermark embedding position of the target document according to the document type includes: When the document type of the target document is WORD, EXCEL or PPT, the root node of the target document is obtained, and a first custom node is embedded under the root node as a watermark embedding position; When the document type of the target document is PDF type, the object node of the target document is obtained, and a second custom node is embedded before the object node as a watermark embedding position.

[0009] Optionally, the reverse parsing of the watermark document to achieve leakage tracing of the target document includes: Determining a watermark embedding position in the watermark document; Extracting the dark watermark encrypted string in the watermark embedding position; Decrypting the dark watermark encrypted string to obtain a decrypted string in a target format; The leakage of the target document is traced based on the decrypted character string in the target format.

[0010] Optionally, determining the watermark embedding position in the watermark document includes: Obtaining the media type of the watermark document; Identifying the document type of the watermark document according to the media type; The watermark embedding position in the watermark document is determined according to the document type.

[0011] Optionally, determining the watermark embedding position in the watermark document according to the document type includes: When the document type of the watermark document is WORD, EXCEL or PPT, the root node of the target document is obtained, and the first custom node under the root node is obtained to obtain the watermark embedding position; When the document type of the watermark document is PDF type, the object node of the watermark document is obtained, and the second custom node before the object node is obtained to obtain the watermark embedding position.

[0012] The second aspect of the present application provides a document tracing device based on dark watermark, the device comprising: A determination module, used for determining the watermark embedding position of the target document; A generation module, used to generate a dark watermark encrypted string containing traceability information; An embedding module, used for embedding the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; The tracing module is used to reversely parse the watermark document to achieve leakage tracing of the target document.

[0013] The third aspect of the present application provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the dark watermark-based document tracing method when executing the computer program.

[0014] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the document tracing method based on dark watermark are implemented.

[0015] The document tracing method, device, computer equipment and medium based on dark watermark provided in this application determine the watermark embedding position of the target document; generate a dark watermark encrypted string containing tracing information; embed the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; reverse parse the watermark document to achieve leakage tracing of the target document. By embedding a dark watermark encrypted string, the target document can be effectively identified and tracked after leakage, thereby improving the security of the document. The concealment and invisibility of the dark watermark make it difficult for attackers to detect and remove the watermark, further enhancing the protection of the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a flowchart of a document tracing method based on dark watermark provided in an embodiment of the present application.

[0018] Figure 2 It is a flowchart of another dark watermark-based document tracing method provided in an embodiment of the present application.

[0019] Figure 3 It is a functional module diagram of a dark watermark-based document tracing device provided in an embodiment of the present application.

[0020] Figure 4It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0022] In various business systems, the functions of downloading or exporting documents (such as WORD, PPT, EXCEL, and PDF) are extremely common, and these functions provide great convenience for users. However, if these functions are not protected, the risk of sensitive data leakage will increase significantly as documents are widely shared and forwarded, and once leaked, it is often difficult to trace its source. Existing business systems provide security protection by adding visible watermarks to downloaded or exported documents. Although this practice has a deterrent effect to a certain extent, reminding users not to share or leak documents at will, the protective effect of visible watermarks is limited. This is because visible watermarks are easily tampered with or removed. Once a document is leaked, it is difficult to effectively trace its source and transmission path based on visible watermarks alone.

[0023] In order to solve the above problems, the present application proposes a document tracing method, device, computer equipment and medium based on dark watermark.

[0024] Figure 1 A flowchart of a document tracing method based on dark watermark is provided in an embodiment of the present application. The document tracing method based on dark watermark includes the following steps.

[0025] S11, determining the watermark embedding position of the target document.

[0026] The target document refers to the original document that needs to be embedded with a watermark, which can be any type of electronic document, such as WORD, EXCEL, PPT, PDF, etc. The target document is the carrier for embedding the watermark and is also the basis for subsequent traceability.

[0027] Determining the watermark embedding position of the target document means determining where in the target document to embed the dark watermark. The structure of the target document can be analyzed to determine which positions are suitable for embedding watermarks, which positions may affect the readability or aesthetics of the document after embedding watermarks, and the concealment of watermarks embedded in different positions can also be evaluated. Determining the watermark embedding position provides a basis for subsequent watermark embedding and extraction, while ensuring the concealment and security of the watermark.

[0028] In an optional implementation, determining the watermark embedding position of the target document includes: Obtaining the media type of the target document; Identifying a document type of the target document according to the media type; The watermark embedding position of the target document is determined according to the document type.

[0029] The media type of the target document can be obtained through file extension, file content analysis or programming interface. The media type (MIME type) is a standard used to indicate the nature and format of a document, file or byte stream.

[0030] Based on the acquired media type, the specific document type of the target document is further identified. The document type refers to the specific format or type of the target document, such as WORD, EXCEL, PPT or PDF.

[0031] The document type of the target document can be identified based on a predefined mapping table or rule set and media type. The mapping table or rule set is used to map a specific media type to the corresponding document type. Different document types usually have specific media types. For example, a WORD document usually has a media type of application / vnd.openxmlformats-officedocument.wordprocessingml.document (for .docx format) or application / msword (for .doc format); a PDF document has a media type of application / pdf. Through a mapping table or rule set, the document type can be quickly and accurately identified based on the media type.

[0032] After determining the document type of the target document, select a suitable watermark embedding position according to the document type to ensure that the watermark can be effectively embedded without affecting the normal use and readability of the document.

[0033] In an optional implementation, determining the watermark embedding position of the target document according to the document type includes: When the document type of the target document is WORD, EXCEL or PPT, the root node of the target document is obtained, and a first custom node is embedded under the root node as a watermark embedding position; When the document type of the target document is PDF type, the object node of the target document is obtained, and a second custom node is embedded before the object node as a watermark embedding position.

[0034] The root node of a WORD document usually represents the entire document or the main part of a document, such as the document body (DocumentBody). The root node of a WORD document can be obtained by accessing the Document Object Model (DOM). The root node of an EXCEL document can be regarded as the entire workbook (Workbook) or a specific worksheet (Worksheet). The root node of an EXCEL document can be accessed through a Workbook object or a Worksheet object. The root node of a PPT document can be the entire presentation (Presentation) or a specific slide (Slide). The root node of a PPT document can be accessed through a Presentation object or a Slide object. For documents of the WORD, EXCEL or PPT type, the root node is the starting node of the document structure. After obtaining the root node of a WORD, EXCEL or PPT document, embedding a custom node (the first custom node) under the root node as the watermark embedding position can ensure that the watermark information can be effectively embedded in the document. The first custom node can be an invisible shape, text box or picture placeholder, etc.

[0035] A PDF document consists of multiple pages, each of which contains multiple objects (such as text objects, image objects, etc.). The structure and content of a PDF document can be accessed by traversing pages and objects. An object node usually represents a specific element in a PDF document, such as a piece of text or an image. For PDF documents, the object node is an important node in the document structure. After obtaining the object node, a custom node (the second custom node) is embedded in front of the object node as the watermark embedding position, which can ensure that the watermark information is recognized and extracted during the document parsing process. The second custom node can be an invisible PDF object, such as a transparent text box or image.

[0036] As a preferred implementation manner, the first custom node and the second custom node may include specific identifiers to facilitate identification and processing in a subsequent parsing process.

[0037] Compared with the prior art, the above optional implementation can flexibly adjust the watermark embedding position according to different document types, improve the accuracy and effectiveness of watermark embedding, and solve the problem in the prior art that the watermark embedding position is fixed and difficult to adapt to different document types.

[0038] S12, generating a dark watermark encrypted string containing traceability information.

[0039] The dark watermark encrypted string may be generated before or after the watermark embedded position of the target document is determined, and the dark watermark encrypted string includes the traceability information.

[0040] In an optional implementation, generating a dark watermark encrypted string containing traceability information includes: Obtaining the operation time, operation object information and client information of the target document; Generate a combined string in a target format based on the operation time, operation object information and client information; The combined character string in the target format is encrypted to obtain a dark watermark encrypted character string.

[0041] The operation time, operation object information and client information of the target document can be obtained through the document management system or operation log. The operation time is used to record the time when the target document is created, modified or accessed. The operation object information is used to record the identifier of the target document or document part being operated, such as the file name, paragraph ID, etc. The client information is used to record the information of the client device that performs the operation, such as the IP address, device ID, user agent, etc.

[0042] Based on a predefined format template, the obtained operation time, operation object information, and client information are concatenated in a certain order and format to obtain a string in a target format (for example, JSON format).

[0043] In order to ensure the security of traceability information, a symmetric encryption algorithm (such as AES) or an asymmetric encryption algorithm (such as RSA) can be used to encrypt the combined string in the target format to obtain a dark watermark encrypted string.

[0044] The dark watermark encrypted string records the operation details and source information of the target document, thus providing a reliable basis for subsequent leakage tracing. Compared with the visible watermark in the prior art, the dark watermark encrypted string of this application has higher concealment and security, and is not easy to be tampered with or removed.

[0045] S13, embedding the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document.

[0046] The generated dark watermark encrypted string is embedded into the watermark embedding position of the target document to obtain the watermark document.

[0047] For WORD, EXCEL or PPT documents, the generated dark watermark encrypted string is embedded in the first custom node; for PDF documents, the generated dark watermark encrypted string is embedded in the second custom node.

[0048] S14, reversely parsing the watermark document to trace the leakage of the target document.

[0049] Automatically embed the encrypted dark watermark string containing traceability information to ensure the confidentiality and security of the traceability information. Once the target document is leaked, in order to trace the leak of the target document, the watermark document can be reverse parsed to extract and decrypt the encrypted dark watermark string, thereby determining the source of the leak.

[0050] For example, in a practical application scenario, a company needs to trace the source of its important internal documents. Through this application, the company can automatically embed a dark watermark encrypted string containing the operation time, operation object information, and client information when the document is generated. Once the document is leaked, the company can obtain the dark watermark encrypted string through reverse analysis to determine the source of the leak.

[0051] In an optional implementation, the reverse parsing of the watermark document to achieve leakage tracing of the target document includes: Determining a watermark embedding position in the watermark document; Extracting the dark watermark encrypted string in the watermark embedding position; Decrypting the dark watermark encrypted string to obtain a decrypted string in a target format; The leakage of the target document is traced based on the decrypted character string in the target format.

[0052] According to the media type and document type of the document, use appropriate libraries or tools to parse the document structure and locate the embedded position of the watermark. For example, for PDF documents, you can use PyPDF2 or fitz library to traverse pages and objects; for Office documents (WORD, EXCEL, PPT), you can use python-docx, openpyxl, python-pptx and other libraries to access the tree structure of the document. Determining the embedded position of the watermark in the watermark document is the starting point of reverse parsing.

[0053] Dark watermarks are usually embedded in documents in an imperceptible way, which may be hidden in the pixel data of the image, the format attributes of the text, or the metadata field of the document. Extract the dark watermark encrypted string according to the embedding method of the watermark. For example, if the watermark is hidden in the pixel data of the image, you can use an image processing library (such as PIL / Pillow) to read the pixel value of the image; if the watermark is hidden in the format attributes of the text, you can parse the format information of the text (such as font, color, size, etc.) to extract the watermark.

[0054] The extracted dark watermark encrypted string usually needs to be decrypted to restore its original information. According to the algorithm and key used when encrypting the watermark, the dark watermark encrypted string is converted into a plaintext string. The decrypted string contains information for tracing, such as the creator of the document, creation time, modification record, etc. This information can be used to track the source of the document leak.

[0055] Compared with the traditional visible watermark method, the document tracing method based on dark watermark in this application has higher security and reliability. Since the dark watermark is difficult to be tampered with or removed, it can effectively prevent the leakage of sensitive data, and after the leakage occurs, it can accurately trace the source of the leakage, providing strong evidence support.

[0056] In an optional implementation, determining the watermark embedding position in the watermark document includes: Obtaining the media type of the watermark document; Identifying the document type of the watermark document according to the media type; The watermark embedding position in the watermark document is determined according to the document type.

[0057] The media type of the watermark document can be obtained by checking the file extension, reading the file header information, or using a specialized library (such as the mimetypes library in Python). The media type (also known as the MIME type) is a standardized string that represents the file format and can help programs identify the type and content of the file.

[0058] After obtaining the media type, the media type is mapped to the document type according to the pre-created mapping table, and then the corresponding document type is searched according to the media type. For example, application / pdf represents a PDF document, application / vnd.openxmlformats-officedocument.wordprocessingml.document represents a WORD document (.docx format), application / vnd.ms-excel represents an EXCEL document (old version .xls format), application / vnd.openxmlformats-officedocument.spreadsheetml.sheet represents an EXCEL document (new version .xlsx format), application / vnd.openxmlformats-officedocument.presentationml.presentation represents a PPT document (.pptx format), etc.

[0059] After the document type is identified, the embedding position of the watermark can be determined according to the document structure of that type.

[0060] In an optional implementation, determining the watermark embedding position in the watermark document according to the document type includes: When the document type of the watermark document is WORD, EXCEL or PPT, the root node of the target document is obtained, and the first custom node under the root node is obtained to obtain the watermark embedding position; When the document type of the watermark document is PDF type, the object node of the watermark document is obtained, and the second custom node before the object node is obtained to obtain the watermark embedding position.

[0061] When the document type of the watermark document is WORD, EXCEL or PPT, it indicates that the watermark document is in the document format of the Microsoft Office series. In an Office document, the structure of the document is represented by a tree structure, in which a root node represents the entire document. Obtaining the root node is the first step in parsing the document structure, and it serves as the starting point for accessing the document content. For a WORD document, the python-docx library can be used to access the root node of the document (i.e., the Document object); for an EXCEL document, the openpyxl library can be used to access the root node of the workbook (i.e., the Workbook object); for a PPT document, the python-pptx library can be used to access the root node of the presentation (i.e., the Presentation object). In the tree structure of an Office document, there are multiple child nodes under the root node, which represent the document's paragraphs, tables, charts and other elements. The first custom node generally refers to a specific node defined directly under the root node for storing watermark information. Once the first custom node is found, the embedding position of the watermark can be determined.

[0062] When the document type of the watermark document is PDF type, it indicates that the watermark document is in PDF document format. In a PDF document, an object node refers to a basic element in a document, such as a page, font, image, etc. Each PDF document consists of multiple objects, which are organized together according to a certain structure. You can use libraries such as PyPDF2, fitz (PyMuPDF) to parse PDF documents and access their object nodes. In the parsing process of a PDF document, the second custom node refers to a node defined before a specific object node for storing watermark information. Once the second custom node is found, the embedding position of the watermark can be determined.

[0063] This application achieves effective tracing of document leaks by embedding and parsing dark watermark encrypted strings. Compared with traditional visible watermarks, this application has higher security and reliability, and can effectively prevent watermarks from being tampered with or removed, thereby improving the effect of document security protection. In addition, this application can more accurately determine the watermark embedding position in different types of documents, ensure the consistency and reliability of watermark embedding, and thus improve the accuracy and effectiveness of document tracing. In this way, the problem of tracing failure caused by inaccurate watermark embedding position in the prior art can be effectively solved.

[0064] Figure 2 A flowchart of another dark watermark-based document tracing method provided in an embodiment of the present application is provided. The dark watermark-based document tracing method includes the following steps.

[0065] S21, obtaining the media type of the target document.

[0066] The target document refers to the original document that needs to be embedded with a watermark, which can be any type of electronic document, such as WORD, EXCEL, PPT, PDF, etc. The target document is the carrier for embedding the watermark and is also the basis for subsequent traceability.

[0067] The media type of the target document can be obtained through file extension, file content analysis or programming interface. The media type (MIME type) is a standard used to represent the nature and format of a document, file or byte stream. For example, .docx represents a WORD document, .xlsx represents an EXCEL document, .pptx represents a PPT document, and .pdf represents a PDF document.

[0068] S22: Identify the document type of the target document according to the media type.

[0069] Based on the acquired media type, the specific document type of the target document is further identified. The document type refers to the specific format or type of the target document, such as WORD, EXCEL, PPT or PDF.

[0070] The document type of the target document can be identified based on a predefined mapping table or rule set and media type. The mapping table or rule set is used to map a specific media type to the corresponding document type. Different document types usually have specific media types. For example, a WORD document usually has a media type of application / vnd.openxmlformats-officedocument.wordprocessingml.document (for .docx format) or application / msword (for .doc format); a PDF document has a media type of application / pdf. Through a mapping table or rule set, the document type can be quickly and accurately identified based on the media type.

[0071] WORD, EXCEL and PPT documents are actually compressed packages of the ZIP protocol, which contain key information such as data files, pictures and document definition files. Therefore, when processing WORD, EXCEL and PPT documents, you need to unzip the compressed package first before you can operate the internal document definition files.

[0072] When the target document is a WORD document, execute S231; when the target document is an EXCEL document, execute S232; when the target document is a PPT document, execute S233; when the target document is a PDF document, execute S234.

[0073] Among them, S231-S234 correspond to the steps of performing structural analysis on the target document. Different structural analysis operations need to be performed for different types of documents to obtain the document definition files or object nodes inside them.

[0074] S231, performing structural analysis on the WORD document to obtain a WORD document definition file.

[0075] After unzipping the compressed file, the WORD document definition file is: word / document.xml.

[0076] S232, performing structural analysis on the EXCEL document to obtain an EXCEL document definition file.

[0077] After unzipping the compressed package, the EXCEL document definition file is: xl / document.xml.

[0078] S233, performing structural analysis on the PPT document to obtain a PPT document definition file.

[0079] After unzipping the compressed package, the PPT document definition file is: ppt / presentation.xml.

[0080] S234, performing structural analysis on the PDF document to obtain object nodes.

[0081] The PDF document structure is similar to the XML structure, including the root node element (%PDF-version number), Pages object node (1 0obj), Page page object (2 0obj), Page Content object node (stream), etc.

[0082] S24, generating a dark watermark encrypted string.

[0083] The operation time, operation object information and client information of the target document can be obtained through the document management system or operation log. The operation time is used to record the time when the target document is created, modified or accessed. The operation object information is used to record the identifier of the target document or document part being operated, such as the file name, paragraph ID, etc. The client information is used to record the information of the client device that performs the operation, such as the IP address, device ID, user agent, etc.

[0084] Based on a predefined format template, the obtained operation time, operation object information, and client information are concatenated in a certain order and format to obtain a combined string in a target format (for example, JSON format).

[0085] In order to ensure the security of traceability information, a symmetric encryption algorithm (such as AES) or an asymmetric encryption algorithm (such as RSA) can be used to encrypt the combined string in the target format to obtain a dark watermark encrypted string.

[0086] S251-S252 correspond to the steps of embedding the dark watermark encrypted string into the target document.

[0087] For different types of documents, different methods are needed to implant the dark watermark encrypted string into the target document.

[0088] S251, implanting the dark watermark encrypted string into the target document based on the document definition file to obtain a watermark document.

[0089] For Word documents, in the root node of the word / document.xml file <w:document>Create a child node <dark:wartermark>Node, and then write the dark watermark encrypted string to <dark:wartermark>The node content can be.

[0090] For Excel documents, in the root node of the xl / document.xml file <workbook>Create a child node <dark:wartermark>Node, and then write the dark watermark encrypted string to <dark:wartermark>The node content can be.

[0091] For PPT documents, in the root node of the ppt / presentation.xml file <p:presentation>Create a sub-point <dark:wartermark>Node, and then write the dark watermark encrypted string to <dark:wartermark>The node content can be.

[0092] Since WORD, EXCEL and PPT documents are compressed packages of the ZIP protocol, you need to first unzip the compressed package, then modify the internal document definition file, and finally repackage it.

[0093] S252: implant the dark watermark encrypted string into the target document based on the object node to obtain a watermark document.

[0094] For a PDF document, create a 0 0 obj node before the 1 0 obj node, and then write the dark watermark encrypted string into the 0 0 obj node content.

[0095] Adding a custom node in the document definition file or in the object node and writing the dark watermark encrypted string into the custom node will not affect the reading and editing of the document, and users will not be able to detect the existence of the dark watermark during normal use.

[0096] S26, tracing the source of the leak based on the watermark document.

[0097] When a watermark document is leaked, the leak can be traced based on the watermark document. Tracing the leak based on the watermark document is to extract the dark watermark encrypted string from the watermark document and trace the source based on the dark watermark encrypted string. Extracting the dark watermark encrypted string is the reverse process of implanting the dark watermark encrypted string.

[0098] For WORD, EXCEL and PPT documents, obtain from the document definition file <dark:wartermark>The content under the node is then decrypted using the same encryption algorithm to obtain the dark watermark decryption string.

[0099] For a PDF document, the object node (such as 00obj) containing the dark watermark encryption string is found and the content in the object node (such as 00obj) is obtained, and then the obtained content is decrypted using the same encryption algorithm to obtain the dark watermark decryption string.

[0100] By parsing the dark watermark decryption string, we can obtain information such as the document download time, downloader, client IP, client information, and the type of sensitive data contained in the document. In this way, the leakage of the target document can be traced.

[0101] In summary, this application is divided into two stages: the first stage is document dark watermark implantation, corresponding to S21, S22, S231, S232, S233, S234, S24, S251, S252; the second stage is dark watermark extraction and tracing, corresponding to S26. Through these two stages, the security of the target document can be effectively protected and the source of the leak can be traced.

[0102] This application embeds a dark watermark encrypted string so that the target document can be effectively identified and tracked after being leaked, thereby improving the security of the document. The concealment and invisibility of the dark watermark make it difficult for attackers to detect and remove the watermark, further enhancing the protection of the document.

[0103] Figure 3 It is a functional module diagram of a dark watermark-based document tracing device provided in an embodiment of the present application.

[0104] In some embodiments, the dark watermark-based document tracing device 30 may include multiple functional modules composed of program code segments. The program code of each program segment in the dark watermark-based document tracing device 30 may be stored in the memory of a computer device and executed by at least one processor to execute (see Figure 1 and / or Figure 2 Description) The function of document tracing based on dark watermark.

[0105] In this embodiment, the dark watermark-based document tracing device 30 can be divided into multiple functional modules according to the functions it performs. The functional modules may include: a determination module 301, a generation module 302, an embedding module 303 and a tracing module 304. The module referred to in this application refers to a series of computer-readable instruction segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0106] The determination module 301 is used to determine the watermark embedding position of the target document; The generating module 302 is used to generate a dark watermark encrypted string containing traceability information; The embedding module 303 is used to embed the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; The tracing module 304 is used to reversely parse the watermark document to trace the leakage of the target document.

[0107] It should be understood that the various variations and specific embodiments of the dark watermark-based document tracing method provided in the above-mentioned embodiments are also applicable to the dark watermark-based document tracing device in the present embodiment. Through the detailed description of the aforementioned dark watermark-based document tracing method, those skilled in the art can clearly understand the implementation process of the dark watermark-based document tracing device in the present embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0108] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, all or part of the steps of the document tracing method based on dark watermark are implemented.

[0109] See also Figure 4 FIG. 4 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. In a preferred embodiment of the present application, the computer device 4 includes a memory 401 , at least one processor 402 , and at least one communication bus 403 .

[0110] Those skilled in the art should understand that Figure 4 The structure of the computer device shown does not constitute a limitation of the embodiments of the present application. The computer device 4 may also include more or less other hardware or software than shown in the figure, or a different arrangement of components.

[0111] In some embodiments, the computer device 4 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The computer device 4 may also include client devices, which include but are not limited to any electronic product that can interact with the client through a keyboard, mouse, remote control, touchpad, or voice control device, such as a personal computer, tablet computer, smart phone, digital camera, etc.

[0112] It should be noted that the computer device 4 is only an example, and other existing or future electronic products that are suitable for the present application should also be included in the protection scope of the present application and included here by reference.

[0113] In some embodiments, the memory 401 stores a computer program and an operating system, and when the computer program is executed by the at least one processor 402, all or part of the steps in the document tracing method based on dark watermark are implemented. The memory 401 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data. Further, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, and the like.

[0114] In some embodiments, the at least one processor 402 is the control core (Control Unit) of the computer device 4, and uses various interfaces and lines to connect various components of the entire computer device 4, and executes various functions and processes data of the computer device 4 by running or executing programs or modules stored in the memory 401, and calling data stored in the memory 401. For example, when the at least one processor 402 executes the computer program stored in the memory, it implements all or part of the steps of the document tracing method based on dark watermarks described in the embodiment of the present application; or implements all or part of the functions of the document tracing device based on dark watermarks. The at least one processor 402 can be composed of an integrated circuit, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips.

[0115] In some embodiments, the at least one communication bus 403 is configured to realize the connection and communication between the memory 401 and the at least one processor 402. Although not shown, the computer device 4 may also include a power supply (such as a battery) for supplying power to each component. Preferably, the power supply may be logically connected to the at least one processor 402 through a power management device, so as to realize the functions of managing charging, discharging, and power consumption management through the power management device. The power supply may also include any components such as one or more DC or AC power supplies, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The computer device 4 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, internal memory, network interfaces, input locations, and display screens, etc., which will not be repeated here.

[0116] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the method described in each embodiment of the present application.

[0117] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0118] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.< / dark:wartermark> < / dark:wartermark> < / dark:wartermark> < / p:presentation> < / dark:wartermark> < / dark:wartermark> < / workbook> < / dark:wartermark> < / dark:wartermark> < / w:document>

Claims

1. A document tracing method based on dark watermark, characterized in that: The method comprises: Determine the watermark embedding position of the target document; Generate a dark watermark encrypted string containing traceability information; Embedding the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; The watermark document is reversely parsed to trace the leakage of the target document.

2. The document tracing method based on dark watermark according to claim 1 is characterized in that: The generating of the dark watermark encrypted string containing the traceability information comprises: Obtaining the operation time, operation object information and client information of the target document; Generate a combined string in a target format based on the operation time, operation object information and client information; The combined character string in the target format is encrypted to obtain a dark watermark encrypted character string.

3. The document tracing method based on dark watermark according to claim 1 is characterized in that: Determining the watermark embedding position of the target document comprises: Obtaining the media type of the target document; Identifying a document type of the target document according to the media type; The watermark embedding position of the target document is determined according to the document type.

4. The document tracing method based on dark watermark according to claim 3 is characterized in that: Determining the watermark embedding position of the target document according to the document type comprises: When the document type of the target document is WORD, EXCEL or PPT, the root node of the target document is obtained, and a first custom node is embedded under the root node as a watermark embedding position; When the document type of the target document is PDF type, the object node of the target document is obtained, and a second custom node is embedded before the object node as a watermark embedding position.

5. The document tracing method based on dark watermark according to claim 1 is characterized in that: The reverse parsing of the watermark document to trace the leakage of the target document includes: Determining a watermark embedding position in the watermark document; Extracting the dark watermark encrypted string in the watermark embedding position; Decrypting the dark watermark encrypted string to obtain a decrypted string in a target format; The leakage of the target document is traced based on the decrypted character string in the target format.

6. The document tracing method based on dark watermark according to claim 5 is characterized in that: Determining the watermark embedding position in the watermark document comprises: Obtaining the media type of the watermark document; Identifying the document type of the watermark document according to the media type; The watermark embedding position in the watermark document is determined according to the document type.

7. The document tracing method based on dark watermark according to claim 6 is characterized in that: Determining the watermark embedding position in the watermark document according to the document type comprises: When the document type of the watermark document is WORD, EXCEL or PPT, the root node of the target document is obtained, and the first custom node under the root node is obtained to obtain the watermark embedding position; When the document type of the watermark document is PDF type, the object node of the watermark document is obtained, and the second custom node before the object node is obtained to obtain the watermark embedding position.

8. A document tracing device based on dark watermark, characterized in that: The device comprises: A determination module, used for determining a watermark embedding position of a target document; A generation module, used to generate a dark watermark encrypted string containing traceability information; An embedding module, used for embedding the dark watermark encrypted string into the target document based on the watermark embedding position to obtain a watermark document; The tracing module is used to reversely parse the watermark document to achieve leakage tracing of the target document.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the document tracing method based on dark watermark as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the document tracing method based on dark watermark as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data watermark generation method and device, storage medium and electronic equipment

    CN117436041A

  • Mail leakage prevention method, post-mail leakage tracing method and electronic equipment

    CN119254743A

  • Document processor and document processing method

    JP2006166091A

Cited By

  • PDF research report tracing method and device based on dark watermark

    CN120234788A

  • A PDF research report tracking and tracing method and device based on dark watermark

    CN120234788B