A Digital Rights Protection Method Based on the Extension of PDF Security Mechanism

By encrypting the cross-reference table and root node, combined with the hardware fingerprint of the reading device, high-strength encryption and permission control of PDF files are achieved, which solves the security risks of PDF documents and the inefficiency of traditional modes, and improves the security and efficiency of copyright protection.

CN115640548BActive Publication Date: 2025-07-08XIAN YOUKAN ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110814291.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-19
Publication Date
2025-07-08
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

The existing PDF document encryption methods have security risks, are easily cracked, cannot effectively control the dissemination rights, and the traditional model is inefficient and cannot meet the copyright protection needs of digital publications.

Method used

The digital copyright protection method based on the PDF security mechanism is extended. By encrypting the cross-reference table and root node, combined with the hardware fingerprint of the reading device, high-strength encryption and permission control of PDF files are realized, including custom security control options such as reading times and printing times.

Benefits of technology

It realizes efficient encryption of PDF files, prevents file spread, improves encryption and decryption performance, ensures the integrity and security of copyright protection, and supports a variety of security mechanisms to ensure data confidentiality and integrity.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention discloses a digital copyright protection method based on the extension of the PDF security mechanism. By encrypting the cross-reference table and the root node and blocking the physical and logical entrances, any PDF parser is prevented from entering the parsing process, thus achieving an encryption effect. In addition to the document control means stipulated by the standard, custom security control options can also be extended, including the number of words copied, the number of reading times, the number of printing times, etc. At the same time, since only a very small amount of data is encrypted, the encryption and decryption performance is greatly improved, and the larger the file, the more obvious the improvement. When the downloaded PDF file is opened on the current device, if it is copied to other computers, it is ensured that the PDF file cannot be opened by any PDF reading tool, thereby preventing the spread of PDF files and achieving the copyright protection of PDF files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital copyright protection, and in particular to a digital copyright protection method based on the extension of the PDF security mechanism. Background Art

[0002] Currently, the most common form of digital publications is stored in PDF format. The PDF document is an electronic document format born in the 1990s and widely used in various industries. However, its Owner Password encryption method has great security risks, and only relies on the "gentlemen's agreement" of the reader developers to ensure its security control behaviors (such as permissions for copying, printing, document splitting, etc.). Therefore, there are many cracking tools for PDF documents encrypted with Owner Password, which can easily break through these restrictions and even directly restore the PDF document to unprotected raw data. This poses a severe challenge to the security of PDF documents.

[0003] The development of domestic digital copyright protection technology has lagged far behind the development of digitalization. The industry urgently needs a simple, efficient, and powerful copyright protection system. Although there are such products in China, they are either complex to use or simply protected and easy to crack, and none of them fundamentally solve the copyright protection needs of publishing enterprises.

[0004] Currently, there are still many government agencies that use the most primitive human delivery mode for the transmission of confidential documents. The fundamental reason is that the dissemination of digitized confidential documents cannot be precisely controlled, so only the traditional mode with high cost and low efficiency can be adopted for distribution. For enterprises, generally, the problem of "who can view" has been solved, but the behaviors during and after viewing cannot be controlled, so enterprise leaks are not uncommon.

[0005] For the dissemination of content creators, creative creators, and personal private documents, there is a very large market demand. However, there is no dedicated copyright protection product on the market for this demand.

[0006] In summary, the existing PDF digital publications generally have problems such as the inability to precisely control dissemination permissions during the dissemination process, lax restriction control, cumbersome control processes, and poor platform compatibility. Summary of the Invention

[0007] To solve the above problems, the present invention provides a digital copyright protection method based on the extension of the PDF security mechanism, which delves into the underlying layer of electronic documents, uses the PDF format parsing engine as the basis of the entire copyright protection system, combines the internal data characteristics and dissemination medium characteristics of the PDF format for identification, and forms complete encryption and permission control for the PDF file by encrypting local key information of the PDF file with high strength.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A digital copyright protection method based on the extension of the PDF security mechanism, comprising the following steps:

[0010] Step 1: First, mark the end position of the PDF file; at the end of the PDF file, find the start position (startref) node of the cross-reference table and obtain the cross-reference table (xref) entry, record the cross-reference table (xref) entry position and mark it; then traverse the cross-reference table (xref) to find the trailer dictionary, and search for the object ID corresponding to the root object entry in the trailer dictionary; find the catalog object (Catalog) that serves as the logical entry of the PDF file from the object ID and record and mark the logical entry position; finally, traverse the catalog object (Catalog) dictionary to find the end position of the catalog object (Catalog) and mark it;

[0011] Step 2: Encrypt the content from the cross-reference table (xref) entry position to the end position of the PDF file; encrypt the content from the logical entry position to the end position of the catalog object (Catalog);

[0012] Step 3: Starting from the end position of the PDF file, append file access permission configuration information for copyright protection. The file access permission configuration information includes, but is not limited to, the number of readings, reading duration, number of prints, printing duration, bound hardware fingerprint, whether copying is allowed, and whether screenshotting is allowed;

[0013] Step 4: First, encrypt the file access permission configuration information starting from the end position of the PDF file and mark the end position of the file access permission configuration information; finally, write the value at the end position of the file access permission configuration information at the end of the file for PDF file parsing.

[0014] In the above digital copyright protection method based on the extension of the PDF security mechanism, symmetric encryption (3DES) is used for encryption.

[0015] In the above digital copyright protection method based on the extension of the PDF security mechanism, the length of the file access permission configuration information is 1024 bytes.

[0016] In the above digital copyright protection method based on the extension of the PDF security mechanism, the access permission of the reading device to the encrypted PDF file is determined according to the hardware fingerprint bound to the reading device; the reading device obtains the access permission after parsing the PDF file; after the PDF file is downloaded to the reading device, the PDF file is encrypted again based on the hardware fingerprint.

[0017] The beneficial effects produced by adopting the present invention are as follows:

[0018] (1) By encrypting the cross-reference table and the root node, and blocking the physical and logical entrances, the present invention makes it impossible for any PDF parser (reader and cracking tool) to enter the parsing process, thus achieving an encryption effect. In addition to the document control means stipulated by the standard, custom extended security control options can be defined, including the number of copied words, the number of reading times, the number of printing times, etc. At the same time, because only a very small amount of data is encrypted, the encryption and decryption performance is greatly improved, and the larger the file, the more obvious the improvement.

[0019] (2) The present invention realizes that when the downloaded PDF file is opened on the current device, if it is copied to other computers, it is ensured that the PDF file cannot be opened by any PDF reading tool, thereby preventing the spread of the PDF file and achieving the copyright protection of the PDF file. Detailed implementation manners

[0020] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0021] A PDF file mainly consists of four parts: a file header, file content, a cross-reference table (xref), and a file trailer.

[0022] File header: Generally, it indicates the PDF file version that the pdf file follows, such as 1.3, 1.7, etc.

[0023] File content: The part that actually stores the file content, which is stored according to obj objects. Each obj object has a number, but it is not necessarily stored in sequence.

[0024] Cross-reference table (xref): Records the position of each obj in the entire file and is located through the offset. A file can have multiple cross-tables.

[0025] File trailer: Relative to the file header, the file trailer stores the most critical file meta information: 1. The offset position of the cross-reference table (xref) in the entire file; 2. The number of the root object (root) of the document; 3. The number of objects in the document.

[0026] The body of a PDF file consists of objects representing the document content. Objects are the basic types of the document, representing the various components of the document, such as fonts, pages, and instance graphics. Since PDF 1.5, the main body can also contain object streams, each of which contains a series of indirect objects. The cross-reference table (xref) contains information that allows random access to the indirect objects in the file, so that any particular object can be located without having to read the entire file. Since PDF 1.5, some or all of the reference table information can also be contained in a reference stream.

[0027] The cross-reference table (xref) is the only part of a PDF file with a fixed format. Each cross-reference table (xref) starts with a line containing the keyword xref. Following this line are one or more reference subsections, which can appear in any order. The structure of the subsections is beneficial for incremental updates because it allows a new reference section to be appended to the PDF file, and the options included are only for objects that have been appended or deleted. For a file that has never been updated, the reference section contains only one subsection, and its object numbers start from 0.

[0028] The digital rights protection method based on the PDF security mechanism extension includes the following steps:

[0029] Step 1, first mark the end position of the PDF file as position A; at the end of the PDF file, find the startref node and obtain the cross-reference table (xref) entry, record the cross-reference table (xref) entry position and mark it as position B; the startref is the starting position of the cross-reference table (xref). The last line of the PDF file should only contain the file end marker %%EOF. The number below the keyword startref represents the byte offset from the beginning of the last cross-reference table (xref) keyword. The cross-reference table (xref) contains information about the indirect objects in the file to allow random access to these objects, so that any particular object can be located without having to read the entire file. The cross-reference table starts with xref, followed by two numbers separated by a space, and then each line is an object information.

[0030] Then traverse the cross-reference table (xref) to find the trailer dictionary, and look up the object ID corresponding to the root object entry in the trailer dictionary. Object types include direct objects and indirect objects. A direct object stores the actual content directly, while an indirect object is a pointer to another object. For example, the data stored in a page object can be simply understood as a dict, including actual text content, image media information, font styles, and various other resources, which can also be indirectly represented by specifying an object.

[0031] All the content in a PDF file is described by individual objects. These objects have their own numbers (object IDs) and start and end positions. The cross-reference table (xref) records the object numbers and start positions of all objects. Through the cross-reference table (xref), any object can be retrieved.

[0032] Find the Catalog object, which is the logical entry point of the PDF file, from the object ID, and mark the position where the logical entry point is located as position C while recording it. Finally, traverse the Catalog object dictionary to find the end position of the Catalog object and mark it as position D.

[0033] Between the trailer and the startref is the trailer dictionary. A dictionary is a type of PDF object, presented in the form of key-value pairs, where each key corresponds to a value. The trailer object appears in dictionary form and contains several key pieces of information in the PDF, such as the root node of the document and the document author. A PDF reader starts parsing the file from the end of the PDF, and through the trailer section, the cross-reference table and certain special objects can be quickly found.

[0034] Step 2: Encrypt the content from position B to position A using the symmetric encryption algorithm (3DES); encrypt the content from position C to position D using the symmetric encryption algorithm (3DES).

[0035] Step 3: Starting from position A, append the file access permission configuration information for copyright protection. The file access permission configuration information includes, but is not limited to, the number of reading times, reading duration, number of printing times, printing duration, bound hardware fingerprint, whether copying is allowed, whether screenshotting is allowed, and the length of the file access permission configuration information is 1024 bytes.

[0036] Step 4: First, start encrypting the file access permission configuration information from position A and mark the end position of the file access permission configuration information as position E; finally, write the value at position E at the end of the file for PDF file parsing.

[0037] Step 5: The server-side handler determines the access permission of the encrypted PDF file for the reading device based on the hardware fingerprint bound to the reading device; the server-side encryption service program automatically encrypts the software to ensure that the encrypted PDF file cannot be opened by any PDF tool, thus preventing the illegal access to server files.

[0038] The reading device obtains the access permission after parsing the PDF file; after the PDF file is downloaded to the reading device, it is encrypted again based on the hardware fingerprint. The execution entity for the re-encryption is the PDF client reading control. The PDF client reading control re-encrypts the PDF file based on the hardware fingerprint of the current device (PC, Android, and iOS phones) during download to ensure that the PDF file can only be opened on the current device and cannot be opened if copied to other computers.

[0039] In terms of system deployment, it is divided into the server-side PDF file encryption service program and the client reading engine. The server-side encryption program supports both the Linux and Windows operating systems, and the client reading control supports Windows, Android, and iOS, and supports mainstream server and client devices. At the same time, all details of encryption and permission control are encapsulated, and the user operation is very simple, no different from reading a normal file.

[0040] PDF file parsing process: First, parse the file trailer to obtain the cross-reference table and the root object number. Then, through the cross-reference table (xref) and the root object number, parse the document layer by layer to construct the document tree. Note that in the root directory, there is usually a / PAGES object, which represents the root node of the pages.

[0041] There are two key entry points for PDF parsing. One is the physical entry point, i.e., the starting position (startref), which is used to obtain the cross-reference table of the PDF file, that is, the offset positions of all objects. By obtaining the offset positions of the objects, the content of all objects can be further obtained. The other is the logical entry point, i.e., the Catalog object (root node). Through the root node, all objects can be strung together in units of pages, and finally complete parsing can be achieved. The logical entry point is the starting position of the root object (root) node. Only by obtaining the information of the root object (root) node can each page be further parsed, and then each object within the page can be parsed. The physical entry point is the starting position of the first object. Only by parsing out each object from here can the value of each object be obtained according to the object ID later.

[0042] The present invention encrypts the cross-reference table and the root node, blocks the physical entry point and the logical entry point, so that any PDF parser (reader and cracking tool) cannot enter the parsing process, thereby achieving an encryption effect. In addition to the document control means stipulated by the standard, custom extended security control options can also be defined, including the number of copied words, the number of reading times, the number of printing times, etc. At the same time, because only a very small amount of data is encrypted, the encryption and decryption performance is greatly improved, and the larger the file, the more obvious the improvement.

[0043] The present invention provides a variety of security mechanisms to ensure the confidentiality and integrity of data and guarantee the normal operation of the business. A variety of security mechanisms are provided through identity authentication, role assignment, user reading monitoring, log reporting, information security level setting, etc. to ensure the confidentiality and integrity of data. At the same time, as needed, it can be supplemented by the method of user identity lock, which is a security guarantee measure for secondary authentication when logging in to the system. It is a supplement to the "username + password" login method. Not only can users log in to the system securely at any time and place, but also the system can effectively identify the legitimate identity of users.

[0044] The above content is a further detailed description of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited to this. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious modifications are made, and the performance or use is the same, and all should be regarded as belonging to the patent protection scope determined by the claims submitted by the present invention.

Claims

1. A digital copyright protection method based on the extension of PDF security mechanism, characterized in that: It includes the following steps: Step 1: First, mark the end position of the PDF file; at the end of the PDF file, find the start position (startref) of the cross-reference table and obtain the cross-reference table (xref) entry, record the position of the cross-reference table (xref) entry and mark it; then traverse the cross-reference table (xref) to find the trailer dictionary, and look up the object ID corresponding to the root object entry in the trailer dictionary; find the catalog object (Catalog) that serves as the logical entry of the PDF file from the object ID and record and mark the logical entry position; Finally, traverse the catalog object (Catalog) dictionary to find the end position of the catalog object (Catalog) and mark it; Step 2: Encrypt the content from the position of the xref entry to the end of the PDF file; encrypt the content from the logical entry position to the end of the catalog object (Catalog); Step 3: Starting from the end position of the PDF file, append file access permission configuration information for copyright protection. The file access permission configuration information includes, but is not limited to, the number of reading times, reading duration, number of printing times, printing duration, bound hardware fingerprint, whether copying is allowed, and whether screenshotting is allowed; Step 4: First, start from the end position of the PDF file to encrypt the file access permission configuration information and mark the end position of the file access permission configuration information; finally, write the value at the end position of the file access permission configuration information at the end of the file for PDF file parsing.

2. The digital rights protection method based on the extension of the PDF security mechanism according to claim 1, characterized in that: Symmetric encryption (3DES) is used for encryption.

3. The digital copyright protection method based on the extension of the PDF security mechanism according to claim 1, characterized in that: The length of the file access permission configuration information is 1024 bytes.

4. The digital copyright protection method based on the extension of the PDF security mechanism according to claim 1, characterized in that: Determine the access permission of the encrypted PDF file for the reading device based on the hardware fingerprint bound to the reading device; the reading device obtains the access permission after parsing the PDF file; the PDF file is encrypted again based on the hardware fingerprint after being downloaded to the reading device.

Citation Information

Patent Citations

  • PDF (Portable Document Format) file information embedding and extracting method based on PDF cross reference table

    CN102622562A

  • Image-text related robust steganography method and system based on PDF file

    CN109784082A