Watermark embedding method and device for PDF (Portable Document Format) data document, electronic equipment and product

By extracting the structural features and text content of PDF documents and using classification models and optical character recognition to generate watermark content, the problem of low efficiency in batch watermark embedding of PDF documents in the automotive industry is solved, automatic and intelligent watermark embedding is achieved, and labor costs are reduced.

CN120744889APending Publication Date: 2025-10-03BEIJING BAICHEBAO TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511240910.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, the watermark embedding efficiency of PDF data documents in the automotive industry is low, and batch processing cannot be achieved, resulting in high labor costs.

Method used

By extracting the document structure features and text content of PDF documents, using the classification model to determine the document type, and generating watermark content based on the document type and vehicle information, optical character recognition and the fitz library are used to extract page image description information, and the watermark is automatically embedded through the pdfcpu library.

Benefits of technology

It realizes the automatic and intelligent watermark embedding of PDF data documents, reduces labor costs, and the embedded watermark is associated with the document content, with better compatibility and adaptive adjustment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744889A_ABST
    Figure CN120744889A_ABST
Patent Text Reader

Abstract

The invention discloses a PDF (Portable Document Format) data document watermark embedding method and device, electronic equipment and a product, and relates to the technical field of watermark embedding. Comprising the following steps: extracting document structure features and text contents of a target PDF data document; determining corresponding vehicle information based on the text content; inputting the document structure features into a classification model to obtain document types; if the document type of the target PDF data document is a circuit diagram, extracting description information of each page picture through optical character recognition, and generating watermark content of each page picture based on the description information and the vehicle information; if the document type is a user manual or a maintenance manual, extracting description information of each page picture through a fit library, and generating watermark content of each page picture based on the description information and the vehicle information; and based on the watermark content of each page picture, embedding a watermark into the region where each page picture is located through the pdfcpu library. According to the method, watermark embedding of batch PDF data documents can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of watermark embedding, and in particular relates to a method, device, electronic equipment and product for embedding watermarks in PDF documents. Background Art

[0002] With the deepening digital transformation of the automotive industry, PDF documents have become the standard for document storage (such as user manuals, repair manuals, and circuit diagrams) due to their unified format, stable content, and strong cross-platform compatibility. However, automotive companies process a large number of PDF documents annually, and their electronic distribution faces severe risks of content leakage, including unauthorized dissemination, theft of trade secrets by competitors, and misuse of user privacy data. To safeguard intellectual property and information security, watermarking technology, as a key means of copyright tracing and content provenance tracking, is widely used to protect automotive PDF documents.

[0003] In the existing technology, watermark embedding for PDF data documents is mostly done manually based on the watermark content set by humans. However, the automotive industry involves a large number of technical documents and manuals. The manual watermarking method is inefficient and cannot achieve batch processing of large-scale documents, which greatly increases labor costs.

[0004] Therefore, how to provide an effective solution to achieve watermark embedding in batches of PDF documents has become a difficult problem to be solved in the prior art. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, electronic equipment and product for embedding watermarks in PDF documents to solve the above-mentioned problems existing in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for embedding a watermark in a PDF document, which is used to embed a watermark in a PDF document of an automobile, comprising: Extract the document structure features and text content of the target PDF document to be embedded with the watermark; Determining vehicle information corresponding to the target PDF document based on the text content; Inputting the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, wherein the document type is a user manual, a maintenance manual, or a circuit diagram; If the document type of the target PDF document is a circuit diagram, extracting description information of each page image in the target PDF document through optical character recognition, and generating watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; If the document type of the target PDF document is a user manual or a maintenance manual, extracting description information of each page image in the target PDF document through the fitz library, and generating watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; Based on the watermark content of each page image in the target PDF document, a watermark is embedded into the area where each page image in the target PDF document is located through the pdfcpu library.

[0007] Based on the above-disclosed content, the present invention extracts the document structure features and text content of the target PDF document to be watermarked; and determines the vehicle information corresponding to the target PDF document based on the text content; then inputs the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, which is a user manual, a maintenance manual, or a circuit diagram; if the document type of the target PDF document is a circuit diagram, optical character recognition is used to extract the description information of each page image in the target PDF document, and watermark content of each page image in the target PDF document is generated based on the description information of each page image and the vehicle information; if the document type of the target PDF document is a user manual or a maintenance manual, the fitz library is used to extract the description information of each page image in the target PDF document, and watermark content of each page image in the target PDF document is generated based on the description information of each page image and the vehicle information; based on the watermark content of each page image in the target PDF document, a watermark is embedded in the area where each page image in the target PDF document is located using the pdfcpu library. In this way, PDF documents can be automatically classified, and the descriptive information of page images in PDF documents can be identified, the corresponding watermark content can be generated, and then the watermark can be automatically embedded, thereby realizing the automatic generation and embedding of watermarks. The embedded watermark content is associated with the content of the PDF document, which can realize the automatic and intelligent embedding of watermarks for batches of PDF documents, reduce labor costs, and can be widely used in the protection of large numbers of PDF documents in the automotive industry.

[0008] In one possible design, based on the watermark content of each page image in the target PDF document, embedding a watermark into the area where each page image in the target PDF document is located through the pdfcpu library includes: Determining a watermark embedding region corresponding to each page image in the target PDF document; if the target PDF document is a user manual or a maintenance manual, randomly generating a watermark embedding region corresponding to each page image in the target PDF document; and if the target PDF document is a circuit diagram, using the region where each page image in the target PDF document is located as the watermark embedding region corresponding to each page image in the target PDF document; Obtaining the background color of the watermark embedding area corresponding to each page image in the target PDF document through OpenCV, and determining the transparency of the watermark corresponding to each page image in the target PDF document based on the background color of the watermark embedding area corresponding to each page image in the target PDF document; Based on the transparency of the watermark corresponding to each page image in the target PDF document, a watermark is embedded into the watermark embedding area corresponding to each page image in the target PDF document through the pdfcpu library.

[0009] In one possible design, before determining the watermark embedding area corresponding to each page image in the target PDF document, the method further includes: Detecting whether the number of characters in the watermark content of each page image in the target PDF document exceeds a preset number of characters; If the number of characters in the watermark content of any page image in the target PDF document exceeds a preset number of characters, the watermark content of any page image is regenerated.

[0010] In one possible design, after determining the vehicle information corresponding to the target PDF document based on the text content, the method further includes: Storing the target PDF document in a file directory corresponding to the vehicle information corresponding to the target PDF document; The vehicle information includes vehicle brand, vehicle model and / or vehicle year.

[0011] In a possible design, the document structure features and text content of the target PDF document to be watermarked are extracted, including: The document structure features and text content of the target PDF document to be embedded with watermark are extracted through optical character recognition.

[0012] In one possible design, the document structure features include the number of pages, the number of pictures and / or the page ratio of text to pictures.

[0013] In one possible design, the classification model is a convolutional neural network model.

[0014] In a second aspect, the present invention provides a device for embedding watermarks into a PDF document, which is used to embed watermarks into a PDF document of an automobile, comprising: An extraction unit, used for extracting the document structure features and text content of the target PDF document to be embedded with a watermark; a determining unit, configured to determine vehicle information corresponding to the target PDF document based on the text content; a classification unit, configured to input the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, wherein the document type is a user manual, a maintenance manual, or a circuit diagram; a watermark generating unit configured to extract description information of each page image in the target PDF document by optical character recognition if the document type of the target PDF document is a circuit diagram, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark generating unit is further configured to extract description information of each page image in the target PDF document through a fitz library if the document type of the target PDF document is a user manual or a maintenance manual, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark embedding unit is used to embed a watermark into the area where each page image in the target PDF document is located through the pdfcpu library based on the watermark content of each page image in the target PDF document.

[0015] In a third aspect, the present invention provides an electronic device comprising a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the PDF data document watermark embedding method as described in the first aspect or any possible design of the first aspect.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on a computer, the method for embedding watermarks in a PDF document according to the first aspect or any possible design of the first aspect is executed.

[0017] In a fifth aspect, the present invention provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the method for embedding watermarks in a PDF document as described in the first aspect or any possible design of the first aspect.

[0018] Beneficial effects: The present invention can automatically classify PDF documents, identify descriptive information of page images in PDF documents, generate corresponding watermark content, and then automatically embed the watermark, thereby achieving automatic watermark generation and embedding. The embedded watermark content is associated with the content of the PDF document, enabling automated and intelligent watermark embedding for batches of PDF documents, reducing labor costs and being widely used in the automotive industry to protect large numbers of PDF documents. Furthermore, because the embedded watermark content is associated with the content of the PDF document, the content of the embedded watermark varies from page to page, making the embedded watermark more compatible and, unlike ordinary watermarks, not easily removed using watermark removal tools.

[0019] Furthermore, the watermark form, position and transparency can be dynamically adjusted according to the content of the PDF document to achieve adaptive adjustment of the watermark. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Flowchart of the method for embedding watermarks in a PDF document provided by an embodiment of the present application; Figure 2 A schematic block diagram of a device for embedding watermarks into a PDF document according to an embodiment of the present application; Figure 3 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.

[0022] It should be understood that although the terms "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the exemplary embodiments of the present invention.

[0023] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may indicate three situations: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" that may appear in this document describes another type of association object relationship, indicating that two relationships may exist. For example, A / and B may indicate two situations: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0024] In order to realize watermark embedding in batches of PDF data documents, the embodiments of the present application provide a PDF data document watermark embedding method, device, electronic equipment and product. The PDF data document watermark embedding method, device, electronic equipment and product can realize automatic and intelligent watermark embedding in batches of PDF data documents.

[0025] The PDF document watermark embedding method provided in the embodiment of the present application can be applied to a user terminal or server for embedding a watermark in a PDF document. It is understood that the execution subject does not constitute a limitation on the embodiment of the present application.

[0026] like Figure 1 , which is a flow chart of a method for embedding watermarks into a PDF document provided in the first aspect of this embodiment. The method for embedding watermarks into a PDF document may include, but is not limited to, the following steps S101 to S106.

[0027] Step S101: extracting the document structure features and text content of the target PDF document to be watermarked.

[0028] In one or more embodiments, a user may upload a target PDF document to be watermarked. After receiving the target PDF document, optical character recognition (OCR) may be used to directly extract the document structure features and text content of the target PDF document.

[0029] The document structure features of the target PDF document may include, but are limited to, content chapter distribution, page count (i.e., number of pages), image count, image location, and / or page ratio of text and images in the target PDF document.

[0030] Step S102: Determine the vehicle information corresponding to the target PDF document based on the text content.

[0031] In one or more embodiments, the vehicle information corresponding to the target PDF document can be identified from the text content using keyword recognition technology. The vehicle information may include, but is not limited to, vehicle brand, vehicle model, and / or vehicle year.

[0032] In one or more embodiments, PDF documents can be stored in separate directories based on different vehicle information. After the vehicle information corresponding to a target PDF document is determined based on its text content, the target PDF document can be stored in a file directory corresponding to the vehicle information corresponding to the target PDF document. For example, file storage directories corresponding to different vehicle information can be set up according to a directory configuration scheme based on "vehicle brand / vehicle model / vehicle year." This allows for categorized storage of PDF documents with different vehicle information.

[0033] Step S103: Input the document structure features into a pre-trained classification model to obtain the document type of the target PDF document.

[0034] Common types of PDF documents in the automotive industry include user manuals, maintenance manuals, and circuit diagrams. These documents have distinct characteristics. User manuals usually contain a large amount of text instructions, operating steps, diagrams, etc. Maintenance manuals usually contain some technical maintenance steps, tool instructions, troubleshooting information, etc. Circuit diagrams are mostly pictures and usually contain a large amount of technical symbols and component information.

[0035] Therefore, in one or more embodiments, to facilitate the classification of PDF documents, a classification model may be pre-established for classifying the document types of PDF documents, such as user manuals, maintenance manuals, or circuit diagrams. The classification model may be trained using the structural features of a large number of sample PDF documents as sample inputs and the document types corresponding to the sample PDF documents as sample outputs. The classification model may be, but is not limited to, a convolutional neural network (CNN) model or a support vector machine (SVM) model. It is understood that the classification of PDF documents may also be based on the text content, text content title, and / or document name of the PDF document.

[0036] It should be noted that the order of the aforementioned step S102 and step S103 is not limited.

[0037] Step S104. If the document type of the target PDF document is a circuit diagram, extract the description information of each page image in the target PDF document through optical character recognition, and generate watermark content for each page image in the target PDF document based on the description information of each page image and vehicle information.

[0038] Circuit diagrams are mostly images, typically containing a large amount of technical symbols and component information. Therefore, in one or more embodiments, if the document type of the target PDF document is a circuit diagram, descriptive information of each page image in the target PDF document can be extracted through optical character recognition. The descriptive information of the page image in the circuit diagram can include, but is not limited to, the circuit name, core component name, design number name, and / or technical terminology (for example, "Module-1: Power Control Module," "R1001: Resistor," etc.). Watermark content is then generated for each page image in the target PDF document based on the descriptive information of each page image and vehicle information (vehicle brand, vehicle model, and / or vehicle year, etc.). When generating the watermark content for each page image, the watermark content for each page image can be generated using a large language model based on the descriptive information and vehicle information of each page image.

[0039] Step S105. If the document type of the target PDF document is a user manual or a maintenance manual, the description information of each page image in the target PDF document is extracted through the Fitz library, and the watermark content of each page image in the target PDF document is generated based on the description information of each page image and the vehicle information.

[0040] The fitz library is a core module in the PyMuPDF library (a powerful Python library for processing PDF, XPS, EPUB, and other documents). It's based on the MuPDF engine, a lightweight, high-performance document rendering library. fitz provides sophisticated document manipulation capabilities, particularly PDFs, including text extraction, image processing, page editing, annotation, and form filling. It's a common tool for working with PDF documents.

[0041] Similarly, when generating a watermark for each page image in the target PDF document based on the description information of each page image and the vehicle information, the watermark can also be generated through a large language model.

[0042] In one or more embodiments, if the document type of the target PDF data document is a user manual or a maintenance manual, the description information of each page image in the target PDF data document can be extracted through the Fitz library. The description information of the page images in the user manual or maintenance manual can be, but is not limited to, some of its keywords or key sentences. Then, the watermark content of each page image in the target PDF data document can be generated based on the description information of each page image and the vehicle information.

[0043] Step S106: Based on the watermark content of each page image in the target PDF document, embed the watermark into the area where each page image in the target PDF document is located through the pdfcpu library.

[0044] In one or more embodiments, embedding a watermark into the area where the image of each page in the target PDF document is located by using the pdfcpu library may include, but is not limited to, the following steps S1061-S1063.

[0045] Step S1061: Determine the watermark embedding area corresponding to each page image in the target PDF document.

[0046] Among them, if the target PDF data document is a user manual or a maintenance manual, the watermark embedding area corresponding to each page image in the target PDF data document can be randomly generated; if the target PDF data document is a circuit diagram, the area where each page image in the target PDF data document is located can be used as the watermark embedding area corresponding to each page image in the target PDF data document.

[0047] In one or more embodiments, before determining the watermark embedding region corresponding to each page image in the target PDF document, a check may be performed to determine whether the number of characters in the watermark content of each page image in the target PDF document exceeds a preset number of characters. The preset number of characters may be set based on actual circumstances. If the number of characters in the watermark content of any page image in the target PDF document exceeds the preset number of characters, the watermark content of any page image in the target PDF document may be regenerated.

[0048] Step S1062: Obtain the background color of the watermark embedding area corresponding to each page image in the target PDF document through OpenCV, and determine the transparency of the watermark corresponding to each page image in the target PDF document based on the background color of the watermark embedding area corresponding to each page image in the target PDF document.

[0049] The darker the background color of the watermark embedding area, the lower the transparency of the watermark corresponding to the page image in the target PDF document (that is, the darker the watermark color); the lighter the background color of the watermark embedding area, the higher the transparency of the watermark corresponding to the page image in the target PDF document (that is, the lighter the watermark color).

[0050] Step S1063: Based on the transparency of the watermark corresponding to each page image in the target PDF document, embed the watermark into the watermark embedding area corresponding to each page image in the target PDF document through the pdfcpu library.

[0051] In summary, the present invention provides a method for embedding watermarks in PDF documents. The method extracts the document structural features and text content of a target PDF document to be watermarked, determines the vehicle information corresponding to the target PDF document based on the text content, inputs the document structural features into a pre-trained classification model, and obtains the document type of the target PDF document, which is a user manual, a maintenance manual, or a circuit diagram. If the document type of the target PDF document is a circuit diagram, optical character recognition is used to extract description information of each page image in the target PDF document, and watermark content of each page image in the target PDF document is generated based on the description information of each page image and the vehicle information. If the document type of the target PDF document is a user manual or a maintenance manual, the fitz library is used to extract description information of each page image in the target PDF document, and watermark content of each page image in the target PDF document is generated based on the description information of each page image and the vehicle information. Based on the watermark content of each page image in the target PDF document, the pdfcpu library is used to embed watermarks into the areas where each page image in the target PDF document is located. This system can automatically classify PDF documents, identify descriptive information for page images within PDF documents, generate corresponding watermark content, and then automatically embed the watermark, achieving automatic watermark generation and embedding. The embedded watermark content is linked to the PDF document content, enabling automated and intelligent watermark embedding for batches of PDF documents, reducing labor costs and widely applicable to protecting large numbers of PDF documents in the automotive industry. Furthermore, because the embedded watermark content is linked to the PDF document content, the embedded watermark content varies from page to page, making the embedded watermark more compatible and unable to be uniformly removed using watermark removal tools like standard watermarks. Furthermore, the watermark's form, position, and transparency can be dynamically adjusted based on the PDF document's content, achieving adaptive watermark adjustment.

[0052] See also Figure 2 In a second aspect, an embodiment of the present application provides a device for embedding watermarks into a PDF document, which is used to embed watermarks into a PDF document of an automobile, comprising: An extraction unit, used for extracting the document structure features and text content of the target PDF document to be embedded with a watermark; a determining unit, configured to determine vehicle information corresponding to the target PDF document based on the text content; a classification unit, configured to input the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, wherein the document type is a user manual, a maintenance manual, or a circuit diagram; a watermark generating unit configured to extract description information of each page image in the target PDF document by optical character recognition if the document type of the target PDF document is a circuit diagram, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark generating unit is further configured to extract description information of each page image in the target PDF document through a fitz library if the document type of the target PDF document is a user manual or a maintenance manual, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark embedding unit is used to embed a watermark into the area where each page image in the target PDF document is located through the pdfcpu library based on the watermark content of each page image in the target PDF document.

[0053] The working process, working details and technical effects of the PDF document watermark embedding device provided in the second aspect of this embodiment can be found in the first aspect of the embodiment and will not be described in detail here.

[0054] like Figure 3 As shown, the third aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the PDF data document watermark embedding method as described in the first aspect of the embodiment.

[0055] For example, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out memory (FIFO) and / or first-in-last-out memory (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series, an ARM (Advanced RISC Machines), an X86 or other architecture processor, or a processor with an integrated NPU (neural-network processing units); the transceiver may include, but is not limited to, a WiFi (Wireless Fidelity) wireless transceiver, a Bluetooth wireless transceiver, a General Packet Radio Service (GPRS) wireless transceiver, a ZigBee protocol (a low-power local area network protocol based on the IEEE802.15.4 standard, ZigBee) wireless transceiver, a 3G transceiver, a 4G transceiver and / or a 5G transceiver, etc.

[0056] A fourth aspect of this embodiment provides a computer-readable storage medium storing instructions for the method for embedding watermarks in PDF documents as described in the first aspect of this embodiment. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, execute the method for embedding watermarks in PDF documents as described in the first aspect. The computer-readable storage medium refers to a data storage medium and may include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a memory stick. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device.

[0057] A fifth aspect of this embodiment provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the PDF document watermark embedding method as described in the first aspect of the embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0058] It should be understood that certain details are provided in the following description to facilitate a thorough understanding of the example embodiments. However, one of ordinary skill in the art will appreciate that the example embodiments can be practiced without these specific details. For example, a system may be shown in block diagrams to avoid obscuring the example with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the example embodiments.

[0059] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A method for embedding watermarks into a PDF document, for embedding watermarks into a PDF document of an automobile, characterized in that: include: Extract the document structure features and text content of the target PDF document to be embedded with the watermark; Determining vehicle information corresponding to the target PDF document based on the text content; Inputting the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, wherein the document type is a user manual, a maintenance manual, or a circuit diagram; If the document type of the target PDF document is a circuit diagram, extracting description information of each page image in the target PDF document through optical character recognition, and generating watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; If the document type of the target PDF document is a user manual or a maintenance manual, extracting description information of each page image in the target PDF document through the fitz library, and generating watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; Based on the watermark content of each page image in the target PDF document, a watermark is embedded into the area where each page image in the target PDF document is located through the pdfcpu library.

2. The method for embedding watermarks in a PDF document according to claim 1, wherein: Based on the watermark content of each page image in the target PDF document, embedding the watermark into the area where each page image in the target PDF document is located through the pdfcpu library, including: Determining a watermark embedding region corresponding to each page image in the target PDF document; if the target PDF document is a user manual or a maintenance manual, randomly generating a watermark embedding region corresponding to each page image in the target PDF document; and if the target PDF document is a circuit diagram, using the region where each page image in the target PDF document is located as the watermark embedding region corresponding to each page image in the target PDF document; Obtaining the background color of the watermark embedding area corresponding to each page image in the target PDF document through OpenCV, and determining the transparency of the watermark corresponding to each page image in the target PDF document based on the background color of the watermark embedding area corresponding to each page image in the target PDF document; Based on the transparency of the watermark corresponding to each page image in the target PDF document, a watermark is embedded into the watermark embedding area corresponding to each page image in the target PDF document through the pdfcpu library.

3. The method for embedding watermarks into a PDF document according to claim 2, wherein: Before determining the watermark embedding area corresponding to each page image in the target PDF document, the method further includes: Detecting whether the number of characters in the watermark content of each page image in the target PDF document exceeds a preset number of characters; If the number of characters in the watermark content of any page image in the target PDF document exceeds a preset number of characters, the watermark content of any page image is regenerated.

4. The method for embedding watermarks into a PDF document according to claim 1, wherein: After determining the vehicle information corresponding to the target PDF document based on the text content, the method further includes: Storing the target PDF document in a file directory corresponding to the vehicle information corresponding to the target PDF document; The vehicle information includes vehicle brand, vehicle model and / or vehicle year.

5. The method for embedding watermarks into a PDF document according to claim 1, wherein: Extract the document structure features and text content of the target PDF document to be watermarked, including: The document structure features and text content of the target PDF document to be embedded with watermark are extracted through optical character recognition.

6. The method for embedding watermarks into a PDF document according to claim 1, wherein: The document structure features include the number of pages, the number of images and / or the page ratio of text to images.

7. The method for embedding watermarks into a PDF document according to claim 1, wherein: The classification model is a convolutional neural network model.

8. A PDF document watermark embedding device for embedding watermarks into automobile PDF documents, characterized in that: include: An extraction unit, used for extracting the document structure features and text content of the target PDF document to be embedded with a watermark; a determining unit, configured to determine vehicle information corresponding to the target PDF document based on the text content; a classification unit, configured to input the document structure features into a pre-trained classification model to obtain the document type of the target PDF document, wherein the document type is a user manual, a maintenance manual, or a circuit diagram; a watermark generating unit configured to extract description information of each page image in the target PDF document by optical character recognition if the document type of the target PDF document is a circuit diagram, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark generating unit is further configured to extract description information of each page image in the target PDF document through a fitz library if the document type of the target PDF document is a user manual or a maintenance manual, and generate watermark content for each page image in the target PDF document based on the description information of each page image and the vehicle information; The watermark embedding unit is used to embed a watermark into the area where each page image in the target PDF document is located through the pdfcpu library based on the watermark content of each page image in the target PDF document.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a transceiver which are sequentially communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the PDF document watermark embedding method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or the instruction is executed by a computer, the method for embedding watermarks in a PDF document as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and device for page watermark generation, method and device for page watermark identification, equipment and storage medium

    CN111191414A

  • Document batch watermark adding method and system and storage medium

    CN114969683A

  • Automatic identification method and device for academic paper directory page, and electronic equipment

    CN118411730A

  • Watermarking digital documents

    US8189861B1