Watermark file processing method, tracing method and device

By generating a hidden watermark in the file that combines the recipient's identity and regional characteristics, the problem of existing watermarking technologies being easily detected and removed is solved, achieving high file security and resistance to attacks, and supporting rapid source tracing.

CN121580366APending Publication Date: 2026-02-27THE PEOPLES BANK OF CHINA DIGITAL CURRENCY INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510840461.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing watermarking technologies, explicit watermarks are easy to detect and remove, while traditional implicit watermarks have low robustness and resistance to attacks when faced with advanced technical means, resulting in reduced file security and difficulty in tracing the source of leakage.

Method used

By combining the identity information of the file recipient and the file feature information of different areas of the file to be processed, a hidden watermark corresponding to different areas is generated and embedded in the file to avoid single, duplicate watermark information that cannot be identified by the naked eye and is difficult to detect or remove in batches.

Benefits of technology

It improves the security and anti-attack capabilities of files, ensuring that watermarked files do not affect normal use during dissemination, and can quickly and accurately identify the original source of files, providing support for tracing the source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580366A_ABST
    Figure CN121580366A_ABST
Patent Text Reader

Abstract

The invention discloses a watermark file processing method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the following steps: acquiring identity information of a to-be-processed file and a file receiver; determining file feature information corresponding to different areas of the to-be-processed file; the file feature information comprises keywords and / or abstract values of the corresponding areas; generating hidden watermarks corresponding to different areas according to the identity information and the file feature information; and embedding the hidden watermark into a corresponding area of the file to be processed, and obtaining and outputting a watermark file. According to the embodiment, the security and the anti-attack capability of the file are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, tracing method, and apparatus for processing watermarked files. Background Technology

[0002] To improve file security, watermarking technology is generally used to trace the source of files that are subsequently distributed.

[0003] However, while existing technologies that use watermarks to trace the source of documents can achieve some degree of tracking of protected information, the obvious visual features of explicit watermarks and the large amount of repetitive watermark text in ordinary implicit watermarks make them easy to detect and remove. This makes it difficult to trace the source of the leak and reduces the security of the documents. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method, method, and apparatus for processing watermarked files. By combining the identity information of the file recipient with the file feature information corresponding to different regions of the file to be processed, implicit watermarks corresponding to different regions are generated. These implicit watermarks are then embedded into the corresponding regions of the file to be processed to obtain a watermarked file. This avoids repetitive watermark information, and the implicit watermarks cannot be detected by the naked eye. Therefore, implicit watermarks in watermarked files are difficult to detect or remove in batches, while not affecting the normal use of the file. This improves file security and resistance to attacks, and also enhances the user experience.

[0005] To achieve the above objectives, according to one aspect of the present invention, a method for processing watermarked documents is provided, the method comprising:

[0006] Obtain the file to be processed and the identity information of the file recipient;

[0007] Determine the file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions;

[0008] Based on identity information and document feature information, generate hidden watermarks corresponding to different regions;

[0009] The hidden watermark is embedded in the corresponding area of ​​the file to be processed, and the watermarked file is obtained and output.

[0010] Optionally, the file to be processed is a text file, and the file feature information corresponding to different regions of the file to be processed is determined, including:

[0011] This paper extracts multiple keywords from a text file using a text summarization algorithm and determines the regions of each keyword within the text file. The keywords are then used as file feature information for the corresponding regions.

[0012] Optionally, the regions of multiple keywords in the text file are determined separately, including: determining the semantic paragraphs corresponding to multiple keywords separately, and using the regions of the semantic paragraphs in the text file as the regions of the keywords in the text file;

[0013] Optionally, keywords can be used as file feature information for the corresponding region, including: if multiple keywords correspond to the same region of a text file, the keyword with the highest weight is selected from the multiple keywords corresponding to the same region as the file feature information.

[0014] Optionally, before extracting multiple keywords from a text file based on a text summarization algorithm, the method further includes: determining the number of keywords to be extracted based on a preset sensitivity level of the text file, wherein the number of keywords extracted is positively correlated with the sensitivity level.

[0015] Optionally, the implicit watermark can be embedded in the corresponding area of ​​the document to be processed, including: embedding the implicit watermark in the beginning and end areas of the corresponding semantic paragraph.

[0016] Optionally, the file to be processed includes watermark information corresponding to the file sender; the file feature information determined based on the text digest algorithm includes the watermark information; the implicit watermark is embedded in the beginning and end regions of the corresponding semantic paragraph, including:

[0017] Embed the hidden watermark into the watermark information.

[0018] Optionally, the file to be processed includes watermark information corresponding to the file sender; multiple keywords are extracted from the text file based on a text digest algorithm, including:

[0019] Detect and remove watermark information from the file to be processed to obtain text information;

[0020] Extracting multiple keywords from text information based on text summarization algorithms.

[0021] Optionally, the implicit watermark is embedded in the corresponding area of ​​the document to be processed, including:

[0022] Embed the hidden watermark in the corresponding area of ​​the text information.

[0023] Optionally, determining the file feature information corresponding to different regions of the file to be processed includes: splitting the file to be processed into multiple file blocks, with different file blocks corresponding to different regions of the file to be processed; calculating multiple summary values ​​corresponding to multiple file blocks, and using the summary values ​​as file feature information.

[0024] Optionally, the file to be processed can be split into multiple file blocks, including: splitting the file to be processed into multiple file blocks according to a preset file block size; or, splitting the file to be processed into multiple file blocks according to a preset number of file blocks.

[0025] Optionally, based on identity information and file feature information, a hidden watermark corresponding to different regions is generated, including: concatenating the identity information and file feature information to obtain a watermark code; encrypting the watermark code and processing it with zero-width characters to obtain a hidden watermark.

[0026] Optionally, the watermark code is encrypted, including: encrypting the watermark code using the national cryptographic algorithm.

[0027] According to a second aspect of the present invention, a method for tracing the source of a watermarked file is provided. The method further includes: obtaining a watermarked file to be traced, and extracting one or more implicit watermarks from the watermarked file to be traced; wherein the watermarked file to be traced is obtained according to any of the watermarked file processing methods of the first aspect described above; matching the extracted one or more implicit watermarks with the pre-stored multiple watermarks according to pre-stored identity information and multiple watermarks to determine whether target identity information exists, wherein the matching rate between the watermark corresponding to the target identity information and the extracted one or more implicit watermarks is higher than a preset threshold; and determining tracing information according to the target identity information if target identity information exists.

[0028] Optionally, if the document feature information includes keywords, the source tracing information is determined based on the target identity information, including:

[0029] Keywords are extracted from the watermarked file to be traced using a text summarization algorithm; the watermarked file to be traced is a text file; it is determined whether the extracted keywords are the same as the keywords corresponding to the target identity information. If they are the same, the target identity information is used as the traceability information.

[0030] Optionally, when the file feature information includes a digest value, the source information is determined based on the target identity information, including:

[0031] According to the splitting method corresponding to the target watermark, the watermark file to be traced is split into multiple file blocks; multiple digest values ​​of the split file blocks are calculated; it is determined whether the calculated multiple digest values ​​are the same as the digest value corresponding to the target identity information. If they are the same, the target identity information is used as the traceability information.

[0032] Optionally, if it is determined that there are multiple target identity information, the source information is determined based on the target identity information, including: determining the file transmission path based on the timestamp of the watermark corresponding to each of the multiple target identity information; and using the file transmission path as the source information.

[0033] Optionally, one or more hidden watermarks can be extracted from the watermarked file to be traced, including:

[0034] Repeat the following operations until a hidden watermark of the preset length is extracted;

[0035] The starting position of the first watermark is detected in the watermark file to be traced, and it is determined whether there is a starting position of the second watermark within a preset length. The starting positions of the first and second watermarks correspond to different hidden watermarks.

[0036] If so, the starting position of the second watermark will be used as the new first starting position;

[0037] If not, extract the hidden watermark based on the first end position corresponding to the first watermark start position.

[0038] To achieve the above objectives, according to a third aspect of the present invention, a watermark file processing apparatus is provided, comprising: an acquisition module, a feature information determination module, a watermark generation module, and a watermark embedding module; wherein,

[0039] The module is configured to retrieve the file to be processed and the identity information of the file recipient.

[0040] The feature information determination module is configured to: determine the file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions;

[0041] The watermark generation module is configured to generate hidden watermarks for different regions based on identity information and file feature information.

[0042] The watermark embedding module is configured to embed a hidden watermark into the corresponding area of ​​the file to be processed, and then output the watermarked file.

[0043] According to a fourth aspect of the present invention, a source tracing device for watermarked documents is provided, comprising: a watermark extraction module, a matching module, and a source tracing information determination module; wherein,

[0044] The watermark extraction module is configured to: acquire the watermark file to be traced, and extract one or more hidden watermarks from the watermark file to be traced; the watermark file to be traced is obtained according to the watermark file processing device provided in the third aspect above;

[0045] The matching module is configured as follows: based on the pre-stored identity information and multiple watermarks, the extracted one or more implicit watermarks are matched with the pre-stored multiple watermarks to determine whether target identity information exists. The matching rate between the watermark corresponding to the target identity information and the extracted one or more implicit watermarks is higher than a preset threshold. If target identity information exists, the source information determination module is triggered.

[0046] The traceability information determination module is configured to determine traceability information based on the target's identity information.

[0047] To achieve the above objectives, according to a fifth aspect of the present invention, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing processor-executable instructions, wherein the processor is configured to execute the instructions to implement a watermark file processing method or a source tracing method according to an embodiment of the present invention.

[0048] To achieve the above objectives, according to a sixth aspect of the present invention, a computer-readable storage medium is provided, which, when the instructions in the computer-readable storage medium are executed by a processor of a file processing server, enables the file processing server to perform a watermark file processing method or a source tracing method according to an embodiment of the present invention.

[0049] To achieve the above objectives, according to another aspect of the present invention, a computer program product is provided, including a computer program that, when executed by a processor, implements a watermark file processing method or a source tracing method according to the present invention.

[0050] One embodiment of the above invention has the following advantages or beneficial effects: by combining the identity information of the file recipient and the file feature information corresponding to different areas of the file to be processed, a hidden watermark corresponding to each area is generated. Then, the hidden watermark is embedded into the corresponding area of ​​the file to be processed to obtain a watermarked file. This avoids single, repetitive watermark information, and the hidden watermark cannot be detected by the naked eye. Therefore, the hidden watermark in the watermarked file is difficult to detect or remove in batches, while not affecting the normal use of the file, thereby improving file security and anti-attack capabilities, and also improving user experience.

[0051] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description

[0052] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:

[0053] Figure 1 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;

[0054] Figure 2 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention;

[0055] Figure 3 This is a flowchart illustrating a watermark file processing method provided in an embodiment of the present invention;

[0056] Figure 4 This is a flowchart illustrating a method for determining file feature information based on keywords, provided in an embodiment of the present invention.

[0057] Figure 5 This is a flowchart illustrating another watermark file processing method provided in an embodiment of the present invention;

[0058] Figure 6 This is a flowchart illustrating another watermark file processing method provided in an embodiment of the present invention;

[0059] Figure 7 This is a flowchart illustrating a method for tracing the source of watermarked files provided in an embodiment of the present invention;

[0060] Figure 8 This is a schematic diagram illustrating the embedding of a hidden watermark into watermark information according to an embodiment of the present invention;

[0061] Figure 9 This is a schematic diagram of the main modules of a watermarked document processing device provided in an embodiment of the present invention;

[0062] Figure 10 This is a schematic diagram of the main modules of a watermarked document tracing device provided in an embodiment of the present invention. Detailed Implementation

[0063] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0064] To enable those skilled in the art to better understand this disclosure, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific implementation methods. It should be noted that, unless otherwise specified, the embodiments of the present invention and the technical features thereof can be combined with each other.

[0065] It should be noted that the technical solutions disclosed in the embodiments of the present invention, regarding the collection, updating, analysis, processing, use, transmission, and storage of user personal information, all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0066] As a crucial factor of production, the secure sharing of data has become an increasingly important issue. However, in practice, once a document is provided to other users or processors, the data owner may completely lose control and oversight of the data. To improve the security of document use, existing technologies employ watermarking for document traceability. However, current technologies typically use explicit watermarks or ordinary implicit watermarks for document traceability. While explicit watermarking technology achieves a certain degree of embedded and traceable protection, its obvious visual characteristics make it easy to detect and remove. For example, after obtaining internal documents with explicit watermarks (such as contracts, financial reports, or other confidential documents), company employees can remove the explicit watermark layer to obtain a watermark-free document, and then leak or publicly distribute the watermark-free document. On the other hand, traditional implicit watermarks, due to their large amount of repetitive text, are also easily detected and removed, especially when facing advanced technologies such as AI. The protective effect of traditional implicit watermarks on documents is greatly reduced, resulting in lower robustness and resistance to attacks for documents containing traditional implicit watermarks. It is evident that explicit watermarks and traditional implicit watermarks can not only lead to the leakage of internal documents and reduce document security, but also make it difficult to trace the source of the leakage.

[0067] To address the aforementioned issues, this invention provides a method for processing watermarked files. By combining the recipient's identity information with file feature information corresponding to different regions of the file to be processed, a hidden watermark is generated for each region. This hidden watermark is then embedded into the corresponding regions of the file to obtain a watermarked file, thereby improving file security and resistance to attacks. Furthermore, since the hidden watermark includes the recipient's identity information, extracting and verifying the watermark information allows for quick and accurate identification of the file's original source, providing support for file traceability.

[0068] The watermark file processing method provided in this invention can be applied to electronic devices such as mobile terminals or servers with displays. A mobile terminal can also be referred to as a terminal device, user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device. A mobile terminal can be a cellular phone, cordless phone, personal digital assistant (PDA) device, handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, computer, laptop computer, handheld communication device, handheld computing device, satellite wireless device, customer premises equipment (CPE), and / or other devices used for communication over wireless systems, as well as next-generation communication systems, such as mobile terminals in 5G networks or mobile terminals in future evolved Public Land Mobile Network (PLMN) networks. A server can be a file server and / or a database server, etc. This application does not specifically limit the form of the mobile terminal and server described above.

[0069] Figure 1 This is an exemplary architecture 100 that can be applied to a watermark file processing method or apparatus provided in the embodiments of the present invention. For example... Figure 1As shown, the system architecture may include a first terminal device 101, a network 102, and a second terminal device 103. The network 102 serves as a medium for providing a communication link between the first terminal device 101 and the second terminal device 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables. Wireless communication links include wireless LAN, Bluetooth, global navigation satellite system, frequency modulation, short-range wireless communication technology, infrared technology, and other wireless communication methods. A file recipient (such as user A) can use the first terminal device 101 to obtain the file to be processed. After the first terminal device 101 processes the file to be processed using the watermark file processing method provided in this embodiment of the invention to obtain a watermarked file, it sends the watermarked file to the second terminal device 103 through the network 102. The watermarked file obtained by the second terminal device 103 is a file with an embedded implicit watermark, and the implicit watermark includes user A's identity information. Since the hidden watermark cannot be detected by the naked eye, and the hidden watermark is generated based on the file feature information corresponding to different areas of the file to be processed, it avoids the single and repeated watermark information. Therefore, the hidden watermark in the watermark file will not affect the normal reading and use of the file by the user of the second terminal device 103 (such as user B), and it is also difficult to be detected and removed in batches. Therefore, it is beneficial to improve the security and anti-attack capability of the file, and also beneficial to improve the user experience.

[0070] Both the first terminal device 101 and the second terminal device 103 can install the watermark file processing device and various client applications provided in this embodiment of the invention, such as instant messaging tools and social platform software.

[0071] It is understandable that when user B reads and uses the watermarked file through the second terminal device 103, if the second terminal device 103 is equipped with the watermarked file processing device provided in this embodiment of the invention, then the second terminal device 103 can also use the watermarked file processing device provided in this embodiment of the invention to generate a watermarked file. Thus, the watermarked file generated by the second terminal device 103 includes a hidden watermark corresponding to user B's identity information; that is, the hidden watermark is generated based on user B's identity information and file feature information. Therefore, when user B continues to transmit the watermarked file to other users through the second terminal 103, the hidden watermark in the watermarked file is difficult to detect and remove in batches, and it will not affect other users' normal reading and use of the file, which is beneficial to improving the security and anti-attack capability of the file. Furthermore, since the watermarked file is sent to user B by user A through the first terminal device 101, the watermarked file generates a hidden watermark corresponding to user A in the first terminal device 101. In the second terminal device 103, a hidden watermark corresponding to user B is generated. Therefore, the watermarked file sent via the second terminal device 103 can include the hidden watermarks corresponding to users A and B respectively. If user A or user B causes file leakage during file use, since the hidden watermark includes the identity information of user A and user B, the user who leaked the file can be identified by extracting the watermark information from the leaked watermarked file (the watermarked file to be traced) and verifying the identity information therein, thus providing support for file tracing. Alternatively, the second terminal device 103 can detect and remove user A's watermark, and then generate a hidden watermark corresponding to user B, so that the watermarked file only includes the watermark information corresponding to the current user, thereby preventing the watermarked file from gradually increasing in size during transmission. During file tracing, the user who leaked the file can also be identified by extracting the watermark information of the last file recipient from the watermarked file to be traced, providing support for file tracing.

[0072] Figure 2 A schematic diagram of the structure of a computer system 200 suitable for implementing an embodiment of the present invention is shown. Figure 2 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0073] like Figure 2 As shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 202 or programs loaded from storage section 208 into random access memory (RAM) 203. The RAM 203 also stores various programs and data required for the operation of the system 200. The CPU 201, ROM 202, and RAM 203 are interconnected via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0074] The following components are connected to I / O interface 205: an input section 206 including a keyboard, mouse, etc.; an output section 207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, modem, etc. The communication section 209 performs communication processing via a network such as the Internet. Drive 210 is also connected to I / O interface 205 as needed. Removable media 211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 210 as needed so that computer programs read from them can be installed into storage section 208 as needed.

[0075] In particular, according to embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 209, and / or installed from removable medium 211. When the computer program is executed by central processing unit (CPU) 201, it performs the functions defined above in the system of this invention.

[0076] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0078] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described system can also be located in a processor, for example, it can be described as: a processor including an acquisition module, a feature information determination module, a watermark generation module, and a watermark embedding module. The names of these modules do not necessarily limit the system itself; for example, the acquisition module can also be described as "a module for acquiring the file to be processed and the identity information of the file recipient".

[0079] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to: acquire a file to be processed and the identity information of the file recipient; determine file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions; generate implicit watermarks corresponding to different regions based on the identity information and the file feature information; embed the implicit watermarks into the corresponding regions of the file to be processed, and obtain and output a watermarked file.

[0080] Specifically, such as Figure 3 As shown, this embodiment of the invention provides a method for processing watermarked files, which mainly includes the following steps S301-S304:

[0081] Step S301: Obtain the file to be processed and the identity information of the file recipient.

[0082] The files to be processed are those that require processing using the watermark file processing method provided in this embodiment of the invention. For example, in the process of hierarchical management of files within an enterprise, higher-level files require traceability control. When a user obtains these higher-level files through a mobile terminal, the mobile terminal will execute the watermark file processing method provided in this embodiment of the invention. These files that need watermark processing are the files to be processed. If the files to be processed are sent to the electronic device of the file recipient via a mobile terminal equipped with the watermark file processing device provided in this embodiment of the invention, then the files to be processed include watermark information corresponding to the file sender. This watermark information is generated based on the file sender's identity information and file feature information. Therefore, when generating the corresponding implicit watermark based on the file recipient's identity information, the watermark information corresponding to the file sender can be retained, or the watermark information can be removed before embedding the implicit watermark corresponding to the file recipient. That is, the generated watermark file can include watermark information corresponding to each file sender and file recipient in the file transmission path, or it can only include the implicit watermark corresponding to the current file recipient. The identity information of the file recipient or file sender may include: user ID, timestamp, IP address, and MAC address, etc.

[0083] Step S302: Determine the file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions.

[0084] In this embodiment of the invention, file feature information can be generated in various ways. The following are exemplary descriptions of various ways to generate file feature information:

[0085] Example 1

[0086] This first embodiment targets text files: keywords are extracted from the text file and used as file feature information. Specifically, multiple keywords can be extracted from the text file based on a text summarization algorithm, and the regions of each keyword in the text file are determined. These keywords are then used as file feature information for the corresponding regions. In this embodiment, the text file is a file containing text content, such as a txt, doc, docx, or xlsx file containing plain text, or a pdf file with editable text.

[0087] In this embodiment, the text summarization algorithm can be an algorithm for calculating word frequency, the TextRank algorithm, the LDA topic model, the BERT deep language model, etc. Among them, the algorithm for calculating word frequency determines keywords by calculating the number of times a word appears in a file. That is, the importance of a word increases in direct proportion to the number of times it appears in the file. When determining keywords, they are selected in descending order of word frequency. The TextRank algorithm is a graph-based ranking algorithm that regards words (sentences) in a file as a network, and the semantic relationships between words (sentences) as links in the network. By iterative calculation, the weight values of words (sentences) are obtained, and the keywords of the file are determined based on the weight values. The LDA (Latent Dirichlet Allocation) topic model is used to discover hidden topic information from a document collection. It regards a document as a probability mixture of multiple topics, and each topic is composed of a probability distribution of vocabulary. The LDA topic model estimates parameters by establishing a document-topic distribution and a topic-word distribution and using probabilistic inference methods, thereby automatically identifying the most representative topic set in the document. Each topic is represented by a group of high-frequency words, and these high-frequency words can be used as keywords in the embodiments of the present invention.

[0088] It can be understood that before using the text summarization algorithm to extract keywords from a text file, preprocessing such as word segmentation and filtering can be performed on the content of the text file to remove meaningless words (such as "de", "shi", "zai", etc.) in the text file to avoid their negative impact on the keyword selection result, thereby improving the accuracy of keyword screening.

[0089] Exemplarily, using the text summarization algorithm, multiple keywords K1, K2,..., Kn can be extracted from a text file. For example, based on the word frequency algorithm, the n keywords that appear most frequently in the text file can be extracted. Using the TextRank algorithm, the n keywords with the highest weight values in the text file can be extracted. Using the LDA topic model, the n high-frequency words corresponding to the most representative topic set in the text file can be extracted. Among them, n can be determined according to the sensitivity level of the text file or set according to actual needs. Examples of determining n according to a preset sensitivity level will be further introduced in subsequent embodiments for hierarchical management of files.

[0090] It can be understood that in the process of extracting keywords, the above several text summarization algorithms can be used alone or in combination, and the embodiments of the present invention do not limit this.

[0091] When determining keywords, or after determining keywords, the keywords can be used as file feature information for the corresponding regions in the text file. For example, when determining the regions of keywords, multiple semantic paragraphs corresponding to keywords can be determined separately, and the regions of the semantic paragraphs in the text file can be used as the regions of the keywords in the text file.

[0092] For LDA topic modeling, we can first determine the semantic paragraphs corresponding to the most representative set of topics. The region that represents the high-frequency words of the most representative set of topics is the region of the determined semantic paragraph in the text file. For example, if the most representative set of topics is located in the 5th and 10th paragraphs of the text file, then the region of the corresponding high-frequency words is also the 5th and 10th paragraphs of the text file.

[0093] For the TextRank algorithm, we can combine the context window to mark the semantic paragraphs P1, P2, ..., Pn where n keywords (K1, K2, ..., Kn) are located. Then, the regions of P1, P2, ..., Pn in the text file are taken as the regions of the keywords in the text file. That is, the region of P1 in the text file is taken as the region of the keyword K1 in the text file. For example, if P1 is in the first sentence of the third paragraph of the text file, then the region of K1 is the first sentence of the third paragraph. And so on, the regions of keywords K2-Kn in the text file can be determined respectively.

[0094] Specifically, the sentence containing each keyword can be extracted, and x sentences before and after it can be used as a context window to form a semantic paragraph. The region of this semantic paragraph in the text file is then used as the region of the corresponding keyword in the text file. For example, if the same keyword appears multiple times in the text file, the context near the regions where the keyword appears multiple times can be merged into a single semantic paragraph. For instance, if the same keyword appears multiple times in adjacent sentences, these adjacent sentences can be merged into a larger context window; if the regions where the keyword appears multiple times are far apart, such as y paragraphs apart, the regions where the keyword appears multiple times can be processed separately, that is, multiple semantic paragraphs can be determined based on the regions where the keyword appears multiple times. Furthermore, if the context window restricts the original paragraph boundaries when determining semantic paragraphs, it may affect semantic coherence. To avoid such problems, in this embodiment of the invention, the context window can be restricted to the same paragraph, or the window size can be adjusted according to the paragraph boundaries to maintain semantic consistency.

[0095] The following is for reference. Figure 4 The specific method for extracting keywords and determining the region of keywords in a text file using the TextRank algorithm in this embodiment of the invention is described, such as... Figure 4 As shown, this implementation includes the following steps:

[0096] Step S401: Perform word segmentation on the text file, preprocess the segmented words to remove meaningless words; record the regional information of each word in the text file.

[0097] In this step, natural language processing tools can be used to perform word segmentation first, and then remove meaningless words such as "de", "shi", "zai", etc. to avoid their negative impact on the selection of keywords.

[0098] The regional information of a word in the text file refers to the sentences and / or semantic paragraphs in the text file where each word is located. For example, by establishing an index structure, the regions of the sentences and / or semantic paragraphs where each word is located in the text file can be recorded.

[0099] Step S402: Use the TextRank algorithm to extract keywords.

[0100] Step S403: For each extracted keyword, determine the sentence in which it is located in the text file, and determine the semantic paragraph corresponding to the keyword according to the set context window.

[0101] In this step, according to the pre-established index structure, the sentence where the keyword is located can be obtained, and then the adjacent sentences before and after can be obtained according to the set context window size, and these adjacent sentences can be merged to form a context paragraph to form the semantic paragraph corresponding to the keyword.

[0102] Step S404: Determine the region of the corresponding keyword in the text file according to the region of the semantic paragraph corresponding to the keyword in the text file.

[0103] Step S405: Use the keyword as the file feature information of its corresponding region.

[0104] It can be further understood that two or more keywords extracted by the text summarization algorithm may correspond to the same semantic paragraph. That is to say, if multiple keywords correspond to the same region of the text file, in order to reduce the complexity of the text feature information, in this case, the keyword with the highest weight can be selected from the multiple keywords corresponding to the same region as the file feature information. This makes only one keyword selected for the same region, and the selected keyword is the vocabulary that best represents the semantic paragraph, thereby not only reducing the complexity of the text feature information but also improving the accuracy of the text feature information.

[0105] Furthermore, to facilitate hierarchical management of files and further enhance file security, in this embodiment of the invention, before extracting multiple keywords from a text file based on a text summarization algorithm, the number of keywords to be extracted is determined according to the preset sensitivity level of the text file. This number of extractions is positively correlated with the sensitivity level. That is, the higher the sensitivity level of the text file, the more keywords are extracted. For example, if files are divided into levels 1, 2, 3...z according to their sensitivity levels from low to high, then the number of keywords corresponding to each level gradually increases, i.e., n1 < n2 < n3 <...nz, where n is the number of keywords. This allows for dynamic adjustment of the watermark quantity based on the file's sensitivity level, in accordance with industry-specific data classification requirements. This facilitates more refined watermarking of text files according to their sensitivity levels. Files with higher sensitivity levels have more watermark embedding points, which enhances the watermark effect and thus improves file security.

[0106] Example 2

[0107] This second embodiment primarily uses a digest value as file feature information. Specifically, the file to be processed can be divided into multiple file blocks, with different file blocks corresponding to different regions of the file. Then, multiple digest values ​​corresponding to the multiple file blocks are calculated, and these digest values ​​are used as file feature information. The file to be processed can be a text file, or it can be an image, compressed file, or other formats.

[0108] In this embodiment, the file to be processed can be split in different ways. For example, it can be split according to paragraphs in the file, with different paragraphs corresponding to different regions of the file, thus obtaining file blocks corresponding to different regions. Alternatively, the file to be processed can be split into multiple file blocks according to a preset file block size or a preset number of file blocks, with different file blocks also corresponding to different regions of the file. For example, if the preset file block size is 100KB, then a 1MB file can be split into 10 100KB file blocks and 1 24KB file block. As another example, if the preset number of file blocks is 10, then a 1MB file can be split into 10 file blocks, each 102.4KB in size; or it can be split into 9 100KB file blocks and 1 124KB file block; or the file can be split into 10 file blocks by combining text semantics or the original text paragraphs.

[0109] To facilitate hierarchical file management, the preset file block size and number of preset file blocks can be determined based on the preset sensitivity level of the files to be processed. Specifically, the file block size is negatively correlated with the preset sensitivity level, while the preset number of file blocks is positively correlated. That is, for files with higher sensitivity levels, each file block is smaller when split according to the file block size; conversely, when processing files with higher sensitivity levels according to the preset number of file blocks, more file blocks are used. This allows for dynamic adjustment of the watermark quantity based on the file's sensitivity level, in accordance with industry-specific data classification requirements. This enables more precise watermarking based on sensitivity level, with higher-sensitivity files having more watermark embedding points, thus enhancing the watermark effect and improving file security.

[0110] After obtaining multiple file blocks, a digest value H1, H2...Hn for each file block can be calculated using algorithms such as hashing or SM3. These digest values ​​H1, H2...Hn can then be used as file feature information for the corresponding region.

[0111] Example 3

[0112] This implementation combines the methods of Embodiment 1 and Embodiment 2 described above to generate text file feature information, that is, the file feature information includes summary values ​​and keywords. For example, keywords are first extracted using any one or more text summarization algorithms provided in Embodiment 1. Then, the text file is split into multiple file blocks according to the semantic paragraphs corresponding to the keywords, and the summary value of each file block is calculated. Thus, the summary value of each file block and the keywords included in that file block are used together as the file feature information of that file block.

[0113] Step S303: Generate hidden watermarks corresponding to different regions based on identity information and document feature information.

[0114] In this embodiment of the invention, implicit watermarks can be generated in various ways. For example, after concatenating identity information and text feature information to obtain a watermark code, the font color of the watermark code is set according to the background color of the corresponding area in the file to be processed, so that the font color is the same as the background color, thereby forming an implicit watermark that matches the area.

[0115] For example, a hidden watermark can be generated through zero-width character processing. Specifically, identity information can be concatenated with file feature information to obtain the watermark code. Taking keywords as file feature information as an example, user identity information is concatenated with keywords K1, K2, ..., Kn from different regions to generate n watermark codes W′1, W′2, ..., W′n. Then, the watermark codes are encrypted and processed with zero-width characters to obtain the hidden watermark.

[0116] In this embodiment of the invention, the watermark code can be encrypted using a national cryptographic algorithm. For example, the SM4 encryption algorithm is used to encrypt the watermark codes W'1, W'2…W'n respectively, obtaining the encrypted strings corresponding to W'1, W'2…W'n. Then, zero-width character processing is applied to each encrypted string to generate the implicit watermarks W1, W2…Wn corresponding to the encrypted strings. Specifically, the encrypted strings corresponding to W'1, W'2…W'n can be converted into binary strings respectively. Then, zero-width characters are selected to represent binary 0 and 1. For example, a zero-width non-hyphen (\u200c) is selected to represent binary 0, and a zero-width hyphen (\u200d) is selected to represent binary 1. Then, the selected zero-width characters are used to convert the binary code into a zero-width character sequence, thereby converting the encrypted watermark code into a zero-width character sequence to generate the implicit watermarks W1, W2…Wn.

[0117] Step S304: Embed the hidden watermark into the corresponding area of ​​the file to be processed, and obtain and output the watermark file.

[0118] For example, when embedding a hidden watermark into a file to be processed, the hidden watermark can be embedded in the beginning and end regions of the semantic paragraphs corresponding to each keyword, or it can be embedded in random regions of the file blocks corresponding to the summary values ​​and / or the beginning and end of the file, thereby obtaining a watermarked file. Since different regions of the watermarked file have hidden watermarks corresponding to their keywords embedded, the robustness of the watermark can be improved, and its anti-tampering ability can be enhanced.

[0119] The following section further explains the watermark file processing method provided in this embodiment of the invention, based on the generation of different file feature information and the watermark embedding method. For example... Figure 5 As shown, in the implementation method for generating file feature information based on keywords, the watermark file processing method provided in this embodiment of the invention mainly includes the following steps:

[0120] Step S501: Obtain the text file and the identity information of the file recipient.

[0121] Step S502: Determine the number of keywords to be extracted based on the preset sensitivity level of the text file, and extract multiple keywords from the text file based on the text summarization algorithm according to the number of keywords extracted.

[0122] Step S503: Determine the semantic paragraphs corresponding to multiple keywords respectively, and determine the region of the keywords in the text file based on the region of the semantic paragraphs in the text file.

[0123] Step S504: For each keyword: concatenate the keyword with the identity information to obtain the watermark code, encrypt the watermark code using the national cryptographic algorithm and process it with zero-width characters to obtain the hidden watermark in the area corresponding to the keyword.

[0124] Step S505: Based on the region corresponding to the hidden watermark, embed the hidden watermark into the corresponding region of the text file to obtain and output the watermark file.

[0125] Additionally, in implementations that generate file feature information based on digest values, such as Figure 6 As shown, the watermark file processing method provided in this embodiment of the invention mainly includes the following steps:

[0126] Step S601: Obtain the file to be processed and the identity information of the file recipient.

[0127] Step S602: Determine the preset file block size or number of file blocks according to the preset sensitivity level of the file to be processed, and split the file to be processed into multiple file blocks according to the preset file block size or number of file blocks.

[0128] Step S603: Calculate multiple digest values ​​corresponding to multiple file blocks.

[0129] Step S604: For each digest value: concatenate the digest value with the identity information to obtain the watermark code, encrypt the watermark code using the national cryptographic algorithm and process it with zero-width characters to obtain the hidden watermark corresponding to the digest value.

[0130] Step S605: Embed the hidden watermark into the corresponding file block to obtain and output the watermark file.

[0131] According to the above embodiments, by combining the identity information of the file recipient and the file feature information corresponding to different areas of the file to be processed, implicit watermarks corresponding to different areas are generated. These implicit watermarks are then embedded into the corresponding areas of the file to be processed to obtain a watermarked file. This avoids single, repetitive watermark information, and the implicit watermarks cannot be detected by the naked eye. Therefore, the implicit watermarks in the watermarked file are difficult to detect or remove in batches, while not affecting the normal use of the file, thereby improving file security and anti-attack capabilities, and also improving user experience. Furthermore, because the implicit watermarks are embedded in the file to be processed, even if the watermarked file undergoes a certain degree of compression, format conversion, or copying and pasting, the watermark information can still be effectively preserved, thereby improving the robustness of the implicit watermark and enhancing its anti-attack capabilities.

[0132] In addition, to trace the source of suspected leaked watermarked files, this embodiment of the invention also provides a method for tracing the source of watermarked files, such as... Figure 7 As shown, this tracing method includes the following steps:

[0133] Step S701: Obtain the watermark file to be traced, and extract one or more hidden watermarks from the watermark file to be traced.

[0134] Step S702: Based on the pre-stored identity information and multiple watermarks, match one or more extracted implicit watermarks with the pre-stored multiple watermarks to determine whether target identity information exists. The matching rate between the watermark corresponding to the target identity information and the one or more extracted implicit watermarks is higher than a preset threshold. If target identity information exists, proceed to step S703.

[0135] Step S703: Determine the source information based on the target identity information.

[0136] In this embodiment of the invention, after generating the encrypted string or implicit watermark corresponding to the watermark code, the encrypted string or implicit watermark can be stored in correspondence with its corresponding identity information and / or the identifier of the watermark file. It is understood that since multiple areas in the watermark file embed implicit watermarks generated based on the identity information, one identity information or watermark file identifier corresponds to multiple implicit watermarks or multiple encrypted strings. The stored encrypted string can be in plaintext form before being converted to binary, or it can be in binary form; the stored implicit watermark is a zero-width character sequence.

[0137] After obtaining the watermarked file to be traced, such as through a web-shared document or a document upload monitoring path on a terminal, the implicit watermark is extracted by detecting zero-width characters. The extracted implicit watermark can then be directly matched with a stored zero-width character sequence, or the implicit watermark can be converted back to binary and matched with a stored binary encrypted string, or the binary implicit watermark can be restored to plaintext and matched with a plaintext encrypted string. This embodiment of the invention does not limit the format of the stored watermark or its matching format with the extracted implicit watermark.

[0138] In this embodiment, the hidden watermark extracted from the watermark file to be traced may have been intercepted or tampered with. Therefore, during the matching process, this embodiment of the invention uses fuzzy matching to compare the extracted watermark with the pre-stored watermark. That is, if the extracted hidden watermark matches a portion of the watermark corresponding to a certain pre-stored identity information, and the proportion of the matching portion (matching rate) is higher than a preset threshold, then the identity information corresponding to the extracted hidden watermark is considered to match the pre-stored target identity information. The user corresponding to this target identity information may be the leaker of the watermark file to be traced, and therefore the tracing information is determined based on this target identity information. The preset threshold for the matching rate can be adjusted according to actual needs, such as setting it to 85%. In this case, for one or more extracted hidden watermarks, they are matched with the pre-stored watermark information in sequence (such as the order in the watermark file to be traced). If the matching rate between a certain extracted hidden watermark and a certain pre-stored watermark information exceeds 85%, it can be considered that the hidden watermark has a corresponding target identity information, and this target identity information is the identity information corresponding to the matched watermark. Alternatively, embodiments of the present invention may also employ precise matching to compare the extracted watermark with the pre-stored watermark. In this case, the preset threshold for the matching rate can be set to 100%.

[0139] To further improve the accuracy of source tracing information, in this embodiment of the invention, when the file feature information includes keywords, the source tracing information can be determined by combining implicit watermarks and keywords. Specifically, a text summarization algorithm is used to extract keywords from the watermarked file to be traced, where the watermarked file is a text file; it is then determined whether the extracted keywords are the same as the keywords corresponding to the target identity information. If they are the same, the target identity information is used as the source tracing information.

[0140] In this embodiment, after generating the implicit watermark, the implicit watermark, its corresponding text digest algorithm, identity information, and keywords can be stored. Once the target identity information is determined, the keywords corresponding to that target identity information can be determined based on the stored information. On the other hand, keywords can be extracted from the watermark file (text file) to be traced using a text digest algorithm. This text digest algorithm is the same as the one used in generating the watermark file, such as using the TextRank algorithm to extract keywords. If the keywords extracted from the watermark file to be traced are the same as the keywords corresponding to the target identity information, then the watermark file to be traced is considered to match the watermark file corresponding to the target identity information, and thus the identity information corresponding to the target watermark is used as the tracing information. Since the identity information corresponding to the target watermark includes user ID, timestamp, IP address, and MAC address, the user and time of the leaked watermark file can be determined based on the identity information, thereby facilitating accurate location of the leaker.

[0141] In another embodiment of the present invention, when the file feature information includes a digest value, the source information can also be determined by combining the implicit watermark and the digest value. Specifically, the watermarked file to be traced can be split into multiple file blocks according to the splitting method corresponding to the target watermark, multiple digest values ​​of the split file blocks can be calculated, and it can be determined whether the calculated multiple digest values ​​are the same as the digest value corresponding to the target identity information. If they are the same, the target identity information is used as the source information.

[0142] In this embodiment, after generating the implicit watermark, the implicit watermark and its corresponding file splitting method, identity information, and digest value can be stored accordingly. After determining the target identity information, the keywords corresponding to the target identity information can be determined based on the stored information. On the other hand, the file splitting method used when generating the target watermark corresponding to the target identity information can be used to split the watermark file to be traced into multiple file blocks, such as splitting them according to a set file block size. Then, the digest value of each file block is calculated. If the calculated digest value is the same as the digest value corresponding to the target identity information, the watermark file to be traced is considered to match the watermark file corresponding to the target identity information, and thus the target identity information is used as the tracing information. Since the identity information corresponding to the target watermark includes user ID, timestamp, IP address, and MAC address, the user and time of the leaked watermark file can be determined based on the identity information, thereby facilitating the accurate location of the leaker.

[0143] The following examples illustrate the situation where a file is traced back to its source after being transmitted through multiple terminal devices. Each terminal device is exemplarily equipped with the watermarked file processing device provided in this embodiment of the invention.

[0144] Example 4

[0145] This embodiment uses the example of user A processing a file to be processed through terminal device A and then sending the file to user B's corresponding terminal device B to illustrate the watermark file processing method and watermark file tracing method provided in this embodiment of the invention. User A, as the file sender, processes the file to be processed through terminal device A; therefore, the file to be processed includes watermark information corresponding to the file sender (user A). After the file receiver (user B) receives the file to be processed through terminal device B, the watermark file processing device configured on terminal device B, during the process of determining the file feature information corresponding to the file to be processed based on a text digest algorithm, may extract the watermark information as a keyword and use it as the file feature information of the corresponding area because the watermark information appears multiple times in different areas of the file to be processed. In other words, the file feature information determined based on the text digest algorithm includes the watermark information in the file to be processed. After generating a hidden watermark based on the file feature information and the identity information of the file receiver, when embedding the hidden watermark into the beginning and end areas of the corresponding semantic paragraph, since the watermark information has already been embedded in the corresponding semantic paragraph, the newly generated hidden watermark may be embedded in the watermark information, such as... Figure 8 As shown. In Figure 8 In the diagram, a1-a2 represent the watermark information corresponding to user A, which is embedded as a hidden watermark in the file to be processed. b1-b2 represent the hidden watermark corresponding to user B, which may be inserted into a1-a2 when inserting the hidden watermark b1-b2. Here, a1 and a2 are the start and end positions of the watermark information corresponding to user A, respectively, and b1 and b2 are the start and end positions of the hidden watermark corresponding to user B, respectively. a1-a2 and b1-b2 have the same number of bits, which can be pre-configured in the watermark file processing device.

[0146] Therefore, during the tracing process of watermarked files, when extracting hidden watermarks from the watermarked file to be traced, the hidden watermarks can be extracted by detecting the start and end positions of the hidden watermarks. Specifically, the following operations can be repeated until a hidden watermark of a preset length is extracted (i.e., a complete hidden watermark from the start to the end position is extracted): detect the start position of the first watermark in the watermarked file to be traced, and determine whether there is a second watermark start position within the preset length. The first watermark start position and the second watermark start position correspond to different hidden watermarks. If they exist, ignore the first watermark start position and use the new second watermark start position as the new first watermark start position to continue the loop. If they do not exist, extract the hidden watermark based on the first watermark end position corresponding to the first watermark start position.

[0147] Taking the example of user A sending a file to be processed to user B's corresponding terminal device B through terminal device A, if user B leaks the watermarked file through terminal device B, then the watermarked file to be traced obtained by the watermarked file tracing device includes, for example, the following: Figure 8 The example shows a hidden watermark of a preset length, which may be embedded in the watermark file to be traced as a zero-width character sequence. When extracting the hidden watermark, the watermark file tracing device detects the first watermark start position (e.g., a1) in the watermark file to be traced based on the structural information of the hidden watermark (e.g., fixed header information fields and preset length), and determines whether there is a second watermark start position (e.g., b1) within the preset length (e.g., 128 bits). In this example, since b1-b2 is embedded in a1-a2, there is a second watermark start position within the preset length where a1 is the first watermark start position. Therefore, the watermark file tracing device ignores the first watermark start position, that is, it no longer uses a1 as the first watermark start position to continue extracting the hidden watermark, but instead uses the second watermark start position (b1) as the new first watermark start position to continue extracting the hidden watermark. During the extraction of the hidden watermark using b1 as the starting position of the first watermark, if no other starting position of the second watermark exists within a preset length, then the hidden watermark b1-b2 is extracted based on the ending position b2 of the first watermark corresponding to b1. In other words, the watermark file tracing device can extract the hidden watermark b1-b2 from the watermark file to be traced, thereby determining the watermark file's tracing information. In this example, user B can be identified as the file leaker. This not only improves the robustness of the hidden watermark and enhances its resistance to attacks, but also facilitates the accurate extraction of the hidden watermark from the watermark file to be traced, thus precisely locating the leaker.

[0148] Example 5

[0149] In this embodiment, taking the example of user C processing a file to be processed through terminal device C and then sending the file to user D's corresponding terminal device D, the watermark file processing method and watermark file tracing method provided by this embodiment of the invention will be exemplarily described. Wherein, user C, as the file sender, processes the file to be processed through terminal device C, and the file to be processed includes watermark information corresponding to the file sender (user C). After the file receiver (user D) receives the file to be processed through terminal device D, in order to avoid the watermark information corresponding to user C affecting the determination of file feature information, terminal device D detects and removes the watermark information from the file to obtain text information, and then extracts multiple keywords from the text information based on a text summarization algorithm, using the keywords as file feature information for the corresponding regions. This ensures that the file feature information accurately reflects the characteristics of the text file itself, improving the accuracy of the text feature information.

[0150] When embedding the implicit watermark corresponding to user D, terminal device D can still embed the implicit watermark into the file to be processed (including the watermark information corresponding to user C). This is similar to the embedding situation in Embodiment 4—since the watermark information corresponding to user C has already been embedded in the corresponding semantic paragraph of the file to be processed, the implicit watermark corresponding to user D generated this time may be embedded in the watermark information. Therefore, during the tracing process of watermarked files, when extracting the implicit watermark from the watermarked file to be traced, the same extraction method as in Embodiment 4 can be adopted, and will not be repeated here.

[0151] Example 6

[0152] In this embodiment, taking the example of user C processing a file through terminal device C and then sending the file to user D's terminal device D, the watermark file processing method and watermark file tracing method provided by this embodiment of the invention will be described exemplarily. Furthermore, similar to Embodiment 5, to avoid the watermark information corresponding to user C affecting the determination of file feature information, terminal device D detects and removes the watermark information from the file to be processed, obtaining text information. Then, based on a text summarization algorithm, it extracts multiple keywords from the text information to use the keywords as file feature information for the corresponding regions.

[0153] When embedding the implicit watermark corresponding to user D, terminal device D, to prevent the file size from continuously increasing during transmission, can embed the implicit watermark into the corresponding area of ​​the text information. That is, when embedding the implicit watermark into the file to be processed, terminal device D first removes the watermark information corresponding to user C from the file, and then embeds the implicit watermark corresponding to user D into the corresponding area of ​​the file, so that the generated watermark file only includes the implicit watermark corresponding to user D. Therefore, during file transmission, the watermark file only includes the implicit watermark corresponding to the user who last received the file, thus avoiding the continuous increase in file size due to the accumulation of watermarks during transmission.

[0154] During the document tracing process, the watermark document tracing device can extract the hidden watermark by detecting zero-width characters, thereby determining the target identity information and the corresponding tracing information.

[0155] Example 7

[0156] In this embodiment, the watermarked file to be traced may be transmitted through multiple users' electronic devices before being acquired by the watermarked file tracing device. If all these users' electronic devices are equipped with the watermarked file processing device provided in this embodiment, then during the file transmission process, each electronic device can generate a hidden watermark corresponding to the user information. During the embedding of the hidden watermark, there is no situation where a later hidden watermark is embedded in an earlier one, i.e., there is no situation as shown in Embodiments 4 or 5 where a later hidden watermark is embedded in an earlier watermark. Therefore, the watermarked file to be traced contains watermark information corresponding to multiple users respectively. In other words, multiple hidden watermarks can be extracted from the watermarked file to be traced. Thus, during the file tracing process, when the pre-stored identity information and multiple watermarks are matched with the extracted hidden watermarks, multiple target identity information may be identified. In this case, the file transmission path can be determined based on the timestamps of the watermarks corresponding to the multiple target identity information, and the file transmission path can be used as tracing information.

[0157] The file transmission path includes the identity information of multiple users transmitting the watermarked file to be traced, as well as the chronological order in which these users transmitted the watermarked file. This allows the transmission process of the watermarked file to be traced to be determined through the file transmission path, thereby improving the accuracy of the traceability information. For example, the watermarked file to be traced is transmitted by user E through a first terminal device to a second terminal device of user F. Then, user F transmits the watermarked file to a public network. Both the first and second terminal devices are equipped with the watermarked file processing device provided in this embodiment of the invention. Therefore, the watermarked file transmitted by user F to the public network includes implicit watermarks corresponding to the identity information of users E and F. After obtaining the watermarked file to be traced through the terminal document upload monitoring path, the watermarked file traceability device provided in this embodiment of the invention can extract the implicit watermarks corresponding to the identity information of users E and F from the watermarked file. After matching the extracted implicit watermarks with the pre-stored identity information and multiple watermarks, the target identity information of users E and F can be determined. Furthermore, since user identity information includes data such as user ID, timestamp, IP address, and MAC address, the file transmission path of the watermark file to be traced can be determined based on the timestamp included in the target identity information of user E and user F. In this example, it is user E (first terminal device) → user F (second terminal device). Because the first terminal device cannot generate the implicit watermark corresponding to user E, the file transmission path can determine both the file transmission process and the user who transmitted the watermark file to the public network (user F in this example), thereby improving the accuracy of the traceability information.

[0158] Based on the inventive concept of the above-described watermarked document processing method, embodiments of the present invention also provide a watermarked document processing apparatus, such as... Figure 9 As shown, the watermarked document processing device 900 includes an acquisition module 901, a feature information determination module 902, a watermark generation module 903, and a watermark embedding module 904; wherein,

[0159] The acquisition module 901 is configured to acquire the file to be processed and the identity information of the file recipient;

[0160] The feature information determination module 902 is configured to: determine the file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions;

[0161] The watermark generation module 903 is configured to generate hidden watermarks corresponding to different areas based on identity information and file feature information.

[0162] The watermark embedding module 904 is configured to embed a hidden watermark into the corresponding area of ​​the file to be processed, and then obtain and output the watermarked file.

[0163] In one embodiment of the present invention, the file to be processed is a text file, and the feature information determination module 902 is configured to: extract multiple keywords from the text file based on a text summarization algorithm, and determine the regions of the multiple keywords in the text file respectively; and use the keywords as the file feature information of the corresponding regions.

[0164] In one embodiment of the present invention, the feature information determination module 902 is configured to: determine the semantic paragraphs corresponding to multiple keywords respectively, and use the region of the semantic paragraph in the text file as the region of the keyword in the text file.

[0165] In one embodiment of the present invention, the feature information determination module 902 is configured to: if the same area of ​​a text file corresponds to multiple keywords, then select the keyword with the highest weight from the multiple keywords corresponding to the same area as the file feature information.

[0166] In one embodiment of the present invention, the feature information determination module 902 is configured to: before extracting multiple keywords from a text file based on a text summarization algorithm, determine the number of keywords to be extracted according to a preset sensitivity level of the text file, wherein the number of keywords extracted is positively correlated with the sensitivity level.

[0167] In one embodiment of the present invention, the watermark embedding module 904 is configured to embed a hidden watermark into the beginning and end regions of the corresponding semantic paragraph.

[0168] In one embodiment of the present invention, the watermark embedding module 904 is configured to embed a hidden watermark into the watermark information.

[0169] In one embodiment of the present invention, the file to be processed includes watermark information corresponding to the file sender; the feature information determination module 902 is configured to: detect and remove the watermark information from the file to be processed to obtain text information; and extract multiple keywords from the text information based on a text summarization algorithm.

[0170] In one embodiment of the present invention, the watermark embedding module 904 is configured to embed a hidden watermark into the corresponding area of ​​the text information.

[0171] In one embodiment of the present invention, the feature information determination module 902 is configured to: split the file to be processed into multiple file blocks, with different file blocks corresponding to different regions of the file to be processed; calculate multiple summary values ​​corresponding to the multiple file blocks, and use the summary values ​​as file feature information.

[0172] In one embodiment of the present invention, the feature information determination module 902 is configured to: split the file to be processed into multiple file blocks according to a preset file block size; or, split the file to be processed into multiple file blocks according to a preset number of file blocks.

[0173] In one embodiment of the present invention, the watermark generation module 903 is configured to: concatenate identity information and file feature information to obtain a watermark code; encrypt the watermark code and perform zero-width character processing to obtain a hidden watermark.

[0174] In one embodiment of the present invention, the watermark generation module 903 is configured to encrypt the watermark code using the national cryptographic algorithm.

[0175] In addition, embodiments of the present invention also provide a device for tracing the source of watermarked documents, such as... Figure 10 As shown, the watermarked document tracing device 1000 includes: a watermark extraction module 1001, a matching module 1002, and a tracing information determination module 1003; wherein,

[0176] The watermark extraction module 1001 is configured to: acquire a watermark file to be traced, and extract one or more hidden watermarks from the watermark file to be traced; wherein the watermark file to be traced is obtained by the watermark file processing device provided in any of the above embodiments;

[0177] The matching module 1002 is configured to: match one or more extracted implicit watermarks with the pre-stored watermarks based on the pre-stored identity information and multiple watermarks to determine whether target identity information exists, and the matching rate between the watermark corresponding to the target identity information and the one or more extracted implicit watermarks is higher than a preset threshold; if target identity information exists, the source information determination module 1003 is triggered.

[0178] The traceability information determination module 1003 is configured to determine traceability information based on the target identity information.

[0179] In one embodiment of the present invention, when the file feature information includes keywords, the source information determination module 1003 is configured to: extract keywords from the watermarked file to be traced using a text digest algorithm, wherein the watermarked file to be traced is a text file; determine whether the extracted keywords are the same as the keywords corresponding to the target identity information, and if so, use the target identity information as the source information.

[0180] In one embodiment of the present invention, when the file feature information includes a digest value, the source information determination module 1003 is configured to: split the watermarked file to be traced into multiple file blocks according to the splitting method corresponding to the target identity information; calculate multiple digest values ​​of the split multiple file blocks; determine whether the calculated multiple digest values ​​are the same as the digest value corresponding to the target identity information; if so, use the target identity information as source information.

[0181] In one embodiment of the present invention, when the matching module 1002 determines that the target identity information corresponds to multiple watermarks, the source information determination module 1003 is configured to: determine the file transmission path according to the timestamps corresponding to the multiple watermarks; and use the file transmission path as source information.

[0182] In one embodiment of the present invention, the watermark extraction module 1001 is configured to: perform the following operations cyclically until a hidden watermark of a preset length is extracted; detect the starting position of the first watermark from the watermark file to be traced, and determine whether there is a second watermark starting position within the preset length, wherein the first watermark starting position and the second watermark starting position correspond to different hidden watermarks; if yes, take the second watermark starting position as the new first watermark starting position; if no, extract the hidden watermark according to the first watermark ending position corresponding to the first watermark starting position.

[0183] According to the above embodiments, by combining the identity information of the file recipient and the file feature information corresponding to different areas of the file to be processed, implicit watermarks corresponding to different areas are generated. These implicit watermarks are then embedded into the corresponding areas of the file to be processed to obtain a watermarked file. This avoids single, repetitive watermark information, and the implicit watermarks cannot be detected by the naked eye. Therefore, the implicit watermarks in the watermarked file are difficult to detect or remove in batches, while not affecting the normal use of the file, thereby improving file security and anti-attack capabilities, and also improving user experience. Furthermore, because the implicit watermarks are embedded in the file to be processed, even if the watermarked file undergoes a certain degree of compression, format conversion, or copying and pasting, the watermark information can still be effectively preserved, thereby improving the robustness of the implicit watermark and enhancing its anti-attack capabilities.

[0184] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for processing watermarked documents, characterized in that, include: Obtain the file to be processed and the identity information of the file recipient; Determine the file feature information corresponding to different regions of the file to be processed; The file feature information includes: keywords and / or summary values ​​for the corresponding region; Based on the identity information and file feature information, generate hidden watermarks corresponding to different regions; The hidden watermark is embedded in the corresponding area of ​​the file to be processed, and the watermarked file is obtained and output.

2. The processing method according to claim 1, characterized in that, The file to be processed is a text file, and determining the file feature information corresponding to different regions of the file to be processed includes: Multiple keywords are extracted from the text file based on a text summarization algorithm, and the regions of the multiple keywords in the text file are determined respectively; The keywords are used as file feature information for the corresponding regions.

3. The processing method according to claim 2, characterized in that, The step of determining the regions of the plurality of keywords in the text file includes: determining the semantic paragraphs corresponding to the plurality of keywords, and using the regions of the semantic paragraphs in the text file as the regions of the keywords in the text file; And / or, The step of using the keyword as the file feature information of the corresponding region includes: if the same region of the text file corresponds to multiple keywords, then the keyword with the highest weight is selected from the multiple keywords corresponding to the same region as the file feature information.

4. The processing method according to claim 2, characterized in that, Before extracting multiple keywords from the text file based on the text summarization algorithm, the method further includes: The number of keywords to be extracted is determined based on the preset sensitivity level of the text file, and the number of keywords to be extracted is positively correlated with the sensitivity level.

5. The processing method according to claim 3, characterized in that, The step of embedding the hidden watermark into the corresponding area of ​​the file to be processed includes: The hidden watermark is embedded in the beginning and end areas of the corresponding semantic paragraph.

6. The processing method according to claim 5, characterized in that, The file to be processed includes watermark information corresponding to the file sender; the file feature information determined based on the text digest algorithm includes the watermark information; embedding the implicit watermark into the beginning and end regions of the corresponding semantic paragraph includes: The hidden watermark is embedded in the watermark information.

7. The processing method according to claim 2 or 3, characterized in that, The file to be processed includes watermark information corresponding to the file sender; the text digest algorithm extracts multiple keywords from the text file, including: The watermark information is detected and removed from the file to be processed to obtain the text information; Multiple keywords are extracted from the text information based on the text summarization algorithm.

8. The processing method according to claim 7, characterized in that, The step of embedding the hidden watermark into the corresponding area of ​​the file to be processed includes: The hidden watermark is embedded in the corresponding area of ​​the text information.

9. The processing method according to claim 1, characterized in that, The step of determining the file feature information corresponding to different regions of the file to be processed includes: The file to be processed is split into multiple file blocks, with different file blocks corresponding to different regions of the file to be processed. Calculate multiple digest values ​​corresponding to the multiple file blocks, and use the digest values ​​as the file feature information.

10. The processing method according to claim 9, characterized in that, The step of splitting the file to be processed into multiple file blocks includes: The file to be processed is split into multiple file blocks according to a preset file block size; or, The file to be processed is split into multiple file blocks according to a preset number of file blocks.

11. The processing method according to claim 1, characterized in that, The step of generating hidden watermarks corresponding to different regions based on the identity information and file feature information includes: The identity information is concatenated with the file feature information to obtain the watermark code; The watermark encoding is encrypted and processed with zero-width characters to obtain the hidden watermark.

12. The processing method according to claim 10, characterized in that, The encryption of the watermark code includes: encrypting the watermark code using a national cryptographic algorithm.

13. A method for tracing the origin of watermarked documents, characterized in that, include: Obtain the watermarked file to be traced, and extract one or more hidden watermarks from the watermarked file to be traced; The watermarked file to be traced is obtained by the watermarked file processing method according to any one of claims 1-9; Based on the pre-stored identity information and multiple watermarks, the extracted one or more hidden watermarks are matched with the pre-stored multiple watermarks to determine whether there is target identity information. The matching rate between the watermark corresponding to the target identity information and the extracted one or more hidden watermarks is higher than a preset threshold. If target identity information exists, traceability information is determined based on the target identity information.

14. The tracing method according to claim 13, characterized in that, When the file feature information includes keywords, determining the source information based on the target identity information includes: Keywords are extracted from the watermarked file to be traced using a text summarization algorithm; the watermarked file to be traced is a text file. Determine whether the extracted keywords are the same as the keywords corresponding to the target identity information. If they are the same, use the target identity information as the source information.

15. The tracing method according to claim 13, characterized in that, When the file feature information includes a digest value, determining the tracing information based on the target identity information includes: According to the splitting method corresponding to the target identity information, the watermarked file to be traced is split into multiple file blocks; Calculate multiple digest values ​​for the split file blocks; Determine whether the calculated multiple digest values ​​are the same as the digest value corresponding to the target identity information. If they are the same, use the target identity information as the tracing information.

16. The tracing method according to claim 13, characterized in that, When it is determined that multiple target identity information exists, the step of determining the tracing information based on the target identity information includes: Based on the timestamps of the watermarks corresponding to the multiple target identity information, the file transmission path is determined; the file transmission path is used as the source tracing information.

17. The tracing method according to claim 13, characterized in that, The step of extracting one or more hidden watermarks from the watermarked file to be traced includes: Repeat the following operations until a hidden watermark of the preset length is extracted; The starting position of the first watermark is detected from the watermark file to be traced, and it is determined whether there is a starting position of the second watermark within a preset length. The starting positions of the first watermark and the second watermark correspond to different hidden watermarks. If so, the starting position of the second watermark will be used as the new starting position of the first watermark; If not, extract the hidden watermark based on the end position of the first watermark corresponding to the start position of the first watermark.

18. A processing apparatus for watermarked documents, characterized in that, include: The module comprises an acquisition module, a feature information determination module, a watermark generation module, and a watermark embedding module; among which, The acquisition module is configured to acquire the file to be processed and the identity information of the file recipient; The feature information determination module is configured to: determine the file feature information corresponding to different regions of the file to be processed; the file feature information includes: keywords and / or summary values ​​of the corresponding regions; The watermark generation module is configured to generate hidden watermarks corresponding to different regions based on the identity information and file feature information. The watermark embedding module is configured to embed the implicit watermark into the corresponding area of ​​the file to be processed, thereby obtaining and outputting the watermarked file.

19. A device for tracing the origin of watermarked documents, characterized in that, include: The module includes a watermark extraction module, a matching module, and a source information determination module; among which, The watermark extraction module is configured to: acquire a watermark file to be traced, and extract one or more hidden watermarks from the watermark file to be traced; the watermark file to be traced is obtained by the watermark file processing device according to claim 18. The matching module is configured to: match the extracted one or more implicit watermarks with the pre-stored watermarks based on the pre-stored identity information and multiple watermarks to determine whether target identity information exists, wherein the matching rate between the watermark corresponding to the target identity information and the extracted one or more implicit watermarks is higher than a preset threshold; and trigger the source information determination module if target identity information exists. The traceability information determination module is configured to determine traceability information based on the target identity information.

20. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method as described in any one of claims 1-17.

21. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of a server processing files, enable the server processing files to perform the method as described in any one of claims 1-17.