Watermark embedding method and device, tracing method and device, equipment and storage medium

By combining watermark algorithms and database tools, embedding multiple watermark data, the security and traceability efficiency problems caused by a single algorithm in the existing watermark technology are solved, and data protection and traceability with high security and high secrets are achieved.

CN120234786APending Publication Date: 2025-07-01GUANGDONG ESHORE TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311850583.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing watermark technology has the problems of single watermark algorithm, low traceability efficiency and insufficient secrets, which leads to poor data security and easily leads to data leakage.

Method used

A combination of pseudo-line watermarking algorithm, pseudo-column watermarking algorithm, and desensitized watermarking algorithm is used to generate multiple types of watermark data, and embed them in the original file, combining a relational database and Elasticsearch tool for traceability recording and analysis.

Benefits of technology

It improves the security and secrets of files, enhances the success rate and efficiency of traceability, and ensures that data can be quickly located and disposed of when leaked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234786A_ABST
    Figure CN120234786A_ABST
Patent Text Reader

Abstract

The invention provides a watermark embedding method, a tracing method, devices, equipment and a storage medium, and the watermark embedding method comprises the steps: obtaining an original file, generating first watermark data according to the original file and an invisible watermark algorithm, determining a watermark algorithm combination according to the file type of the original file and a preset rule, and storing the first watermark data and the watermark algorithm combination. According to the method, the original file and the watermark algorithm combination are combined, combined watermark data are generated according to the watermark algorithm combination and the original file, the watermark algorithm combination comprises at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm and a desensitization watermark algorithm, multiple types of watermark algorithms are provided, and the confidentiality is improved; and writing the first watermark data and the combined watermark data into the original file to obtain the target file, thereby effectively improving the security of the target file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of watermarks, and in particular, to a watermark embedding method, a traceability method, a device, a device, and a storage medium. Background Art

[0002] The purpose of using watermark technology is to protect the security of data without affecting the original data and to quickly locate and respond to data leakage. At present, the existing publicly disclosed watermark technology models generally have the problems of a single watermark algorithm, poor traceability efficiency, and low traceability success rate. The single watermark algorithm results in insufficient concealment and poor security performance, making it easy to cause data leakage. Therefore, an effective method needs to be designed to solve the existing problems. Summary of the Invention

[0003] Embodiments of this application provide a watermark embedding method, a traceability method, a device, a device, and a storage medium to solve at least one problem existing in the related art. The technical solutions are as follows:

[0004] In a first aspect, embodiments of this application provide a watermark embedding method, including:

[0005] Obtain the original file;

[0006] Generate first watermark data according to the original file and the invisible watermark algorithm;

[0007] Determine a watermark algorithm combination according to the file type of the original file and a preset rule, and generate combined watermark data according to the watermark algorithm combination and the original file. The watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm;

[0008] Write the first watermark data and the combined watermark data into the original file to obtain a target file.

[0009] In an implementation manner, the generating first watermark data according to the original file and the invisible watermark algorithm includes:

[0010] Read the digest information of the original file;

[0011] Generate a unique distribution serial number, and encrypt the distribution serial number through an encryption algorithm to obtain an encrypted distribution serial number;

[0012] Convert the encrypted distribution serial number into a zero-width character to obtain first watermark data;

[0013] Wherein, the first watermark data is used to be written into the digest information.

[0014] In one embodiment, when the watermark algorithm combination includes the pseudo-row watermark algorithm, generating combined watermark data according to the watermark algorithm combination and the original file includes:

[0015] Extracting several rows of data from the original file as row samples;

[0016] Extracting the data of each column from the row samples to obtain column data of several columns, and determining the attributes of the column data;

[0017] Forging pseudo-row data of several rows according to the attributes of the column data to obtain second watermark data;

[0018] Among them, each row of pseudo-row data is respectively used for insertion in the original file at intervals of a preset number of rows.

[0019] In one embodiment, when the watermark algorithm combination includes the pseudo-column watermark algorithm, generating combined watermark data according to the watermark algorithm combination and the original file includes:

[0020] Determining the usage characteristics of the data in the original file, and determining a target attribute from a preset digital dictionary attribute library as the pseudo-column attribute name according to the usage characteristics;

[0021] Generating pseudo-column data of several columns according to the pseudo-column attribute name and the generation algorithm to obtain third watermark data;

[0022] Among them, the pseudo-column data is used to be inserted into the corresponding position of the original file according to a preset insertion position strategy.

[0023] In one embodiment, when the watermark algorithm combination includes the desensitization watermark algorithm, generating combined watermark data according to the watermark algorithm combination and the original file includes:

[0024] Determining the attributes of the column data in the original file;

[0025] Determining the target column to be desensitized according to the attributes of the column data;

[0026] According to the sensitive algorithm for matching corresponding sensitive words from the sensitive word library according to the attributes, performing desensitization processing on the target column through the sensitive algorithm to obtain fourth watermark data;

[0027] Among them, the fourth watermark data is used to replace the target column in the original file.

[0028] In one embodiment, the method further includes:

[0029] Record the first watermark data and the combined watermark data line by line, recording the unique distribution serial number, the positions and times at which the first watermark data and the combined watermark data are written into the original file, to obtain a log file, where the distribution serial number is generated during the generation process of the first watermark data;

[0030] Store the log file in a specified directory, and save the distribution information of the distribution serial number to a relational database.

[0031] In a second aspect, an embodiment of the present application provides a tracing method, including:

[0032] Obtain the leaked content, where the leaked content is the target file or a part of the target file, and the target file is obtained by the watermark embedding method;

[0033] Determine the distribution serial number of the leaked content through a preset tracing algorithm;

[0034] Perform tracing according to the distribution serial number.

[0035] In a third aspect, an embodiment of the present application provides a watermark embedding device, including:

[0036] An acquisition module, configured to acquire an original file;

[0037] A first generation module, configured to generate first watermark data according to the original file and a stealth watermark algorithm;

[0038] A second generation module, configured to determine a watermark algorithm combination according to the file type of the original file and a preset rule, and generate combined watermark data according to the watermark algorithm combination and the original file, where the watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm;

[0039] A writing module, configured to write the first watermark data and the combined watermark data into the original file to obtain a target file.

[0040] In an implementation manner, the watermark embedding device further includes a recording module, and the recording module is configured to:

[0041] Record the first watermark data and the combined watermark data line by line, recording the unique distribution serial number, the positions and times at which the first watermark data and the combined watermark data are written into the original file, to obtain a log file, where the distribution serial number is generated during the generation process of the first watermark data;

[0042] Store the log file in a specified directory, and save the distribution information of the distribution serial number to a relational database.

[0043] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory. Instructions are stored in the memory and are loaded and executed by the processor to implement the method in any one of the above aspects.

[0044] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed implements the method in any one of the above aspects.

[0045] The beneficial effects in the above technical solutions at least include:

[0046] By obtaining the original file, generating first watermark data according to the original file and the invisible watermark algorithm, determining a watermark algorithm combination according to the file type of the original file and preset rules, and generating combined watermark data according to the watermark algorithm combination and the original file, and the watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm, providing multiple types of watermark algorithms is beneficial to improving the secrecy; writing the first watermark data and the combined watermark data into the original file to obtain the target file, thereby effectively improving the security of the target file.

[0047] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.

[0049] Figure 1 It is a schematic flow chart of the steps of a watermark embedding method according to an embodiment of the present application;

[0050] FIG. 2(a) is a schematic diagram before processing the pseudo-row watermark algorithm according to an embodiment of the present application, and FIG. 2(b) is a schematic diagram after processing the pseudo-row watermark algorithm according to an embodiment of the present application;

[0051] Figure 3 It is a schematic diagram after processing the desensitization watermark algorithm according to an embodiment of the present application;

[0052] Figure 4 It is a schematic diagram of a target file according to an embodiment of the present application;

[0053] Figure 5 It is a schematic flowchart of the steps of a traceability method according to an embodiment of the present application;

[0054] Figure 6 It is a structural block diagram of a watermark embedding device according to an embodiment of the present application;

[0055] Figure 7 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0056] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature rather than restrictive.

[0057] Glossary:

[0058] Document watermark: Document watermark is a branch of digital watermark, and its carrier is a text file. It usually refers to implanting some specific data codes or information that are not easily recognized and cracked into the text through a certain method. Without affecting the text content, it provides a verification and protection mechanism for digital content. Document watermark can also provide a tracking and monitoring mechanism during the transmission of digital content to protect the integrity, security and reliability of digital content, and can effectively prevent behaviors such as text content tampering, forgery, and leakage.

[0059] Pseudo-row watermark: It refers to a watermark technology that forges one or several rows of data in the row dimension by a certain method and embeds it into the original file.

[0060] Pseudo-column watermark: It refers to a watermark technology that forges a new attribute column in the column dimension by a certain method and embeds it into the original file. The data of pseudo-rows and pseudo-columns should be kept in the same type format as the original data in the column as much as possible, or related to the data content characteristics of the file, so as to be not easily noticed.

[0061] Invisible watermark: Embeds invisible data that cannot be perceived visually in the text, and the content and embedding position of the invisible data can be customized according to requirements.

[0062] Desensitized watermark: It refers to desensitizing or transforming some data in the original data according to desensitization rules to protect sensitive privacy data.

[0063] Watermark traceability: It refers to obtaining key information such as data distribution objects, distribution times, and distribution methods by reading, identifying and analyzing watermarks and code information in the file, so as to trace the whole process of data leakage.

[0064] Elasticsearch: An open-source, highly scalable, distributed full-text search engine that can store and retrieve data almost in real time. It has good scalability itself and can be scaled to hundreds of servers to handle PB-level data. ES has the following characteristics: a distributed real-time document storage engine where each field can be indexed and searched; a distributed real-time analysis search engine that supports various query and aggregation operations; capable of scaling to hundreds of service nodes and supporting PB-level structured or unstructured data.

[0065] Filebeat: A lightweight shipper used to forward and centralize log data. It is used to monitor specified log files or directories, collect log information, and forward them to storage media such as Elasticsearch for indexing.

[0066] Refer to Figure 1 , which shows a flowchart of a watermark embedding method according to an embodiment of the present application. The watermark embedding method may at least include steps S100 - S400:

[0067] S100. Obtain the original file.

[0068] S200. Generate first watermark data according to the original file and the invisible watermark algorithm.

[0069] S300. Determine the watermark algorithm combination according to the file type of the original file and the preset rules, and generate combined watermark data according to the watermark algorithm combination and the original file.

[0070] Optionally, the watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm.

[0071] S400. Write the first watermark data and the combined watermark data into the original file to obtain the target file.

[0072] The watermark embedding method of the embodiment of the present application can be executed by an electronic control unit, a controller, a processor, etc. of terminals such as a computer, a mobile phone, a tablet, and a vehicle-mounted terminal, or can also be executed by a cloud server, such as the system of a cloud server.

[0073] The technical solution of the embodiment of the present application generates first watermark data by obtaining the original file and according to the original file and the invisible watermark algorithm, determines the watermark algorithm combination according to the file type of the original file and the preset rules, and generates combined watermark data according to the watermark algorithm combination and the original file. The watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm. Providing various types of watermark algorithms is beneficial to improving the secrecy; writing the first watermark data and the combined watermark data into the original file to obtain the target file, thereby effectively improving the security of the target file.

[0074] In one implementation, the user can upload the original file to the system. The system reads the original file in the form of an IO stream and can obtain the file type. The original file includes, but is not limited to, structured or semi-structured data text files of file types such as EXCEL, CVS, TXT, and WORD.

[0075] In one implementation, step S200 includes steps S210 - S230:

[0076] S210. Read the summary information of the original file.

[0077] Optionally, when reading the summary information of the original file, the content, title, author, body, identifier, and other attributes of the summary and the corresponding attribute values of each attribute can be read out.

[0078] S220. Generate a unique distribution serial number, and encrypt the distribution serial number through an encryption algorithm to obtain the encrypted distribution serial number.

[0079] Optionally, generate a unique distribution serial number distributeID, use the RAS key generator to generate public and private keys, and encrypt the distribution serial number distributeID through the RAS asymmetric encryption algorithm to obtain the encrypted distribution serial number distributeID. It should be noted that in other embodiments, other encryption algorithms can be used, and no specific limitation is made.

[0080] S230. Convert the encrypted distribution serial number into a zero-width character to obtain the first watermark data.

[0081] Optionally, convert the encrypted distribution serial number distributeID into a binary string, and then convert the binary string into a zero-width character to obtain the first watermark data. Among them, the first watermark data is used to be written into the summary information; the zero-width character is a real Unicode character that is hidden and not displayed or printable in the file, so it has high secrecy.

[0082] In one embodiment, preset principles can be set in advance, and one or more watermark algorithms can be configured for each file type, and then the watermark algorithm combination can be determined according to the file type of the original file. For example, the supported file types include but are not limited to EXCEL, CVS, TXT, WORD, etc. For example, for an EXCEL format file, one or more of the pseudo-row watermark algorithm, pseudo-column watermark algorithm, and desensitization watermark algorithm can be selected, and for TXT, the desensitization watermark algorithm can be selected. In some embodiments, the preset rule can also select the null algorithm (i.e., equivalent to not selecting any watermark algorithm). For example, the null algorithm can be set for WORD. It should be noted that in the embodiments of the present application, for the convenience of introducing each watermark method, it is assumed that the watermark algorithm combination includes the pseudo-row watermark algorithm, pseudo-column watermark algorithm, and desensitization watermark algorithm at the same time, so as to maximize the robustness and security of the watermark.

[0083] In one embodiment, when the watermark algorithm combination includes the pseudo-row watermark algorithm, generating the combined watermark data according to the watermark algorithm combination and the original file in step S300 includes steps S301 - S303:

[0084] S301. Extract several rows of data from the original file as row samples.

[0085] Optionally, read the original file, and extract several rows of data from the original file as row samples. For example, it can be extracted by random sampling.

[0086] S302. Extract the data of each column from the row samples to obtain the column data of several columns, and determine the attributes of the column data.

[0087] Optionally, extract by column from the row samples, extract the data of each column to obtain the column data of several columns, obtain the column sample data set C0...Cn, where C0 represents the column data of the first column, and determine the attributes of the column data. It should be noted that the original file can be read and analyzed to know the attributes of each column of data, including but not limited to column names, column subscripts, column categories (such as mobile phone numbers, administrative regions, ages), data types (such as character type, numerical type), data formats (such as the number of decimal places, character length), etc.

[0088] S303. Forge several rows of pseudo-row data according to the attributes of the column data to obtain the second watermark data.

[0089] Optionally, forge a pseudo-row of data according to the attributes of the columns. The values of each column in this row are generated according to the constraints using C0...Cn as samples, that is, having the same column category, data type, data format, and reasonable value range (the difference is less than the value threshold) as the original column data of the column where it is located. Then repeat the above steps to obtain several rows of pseudo-row data, and obtain the second watermark data. Among them, each row of pseudo-row data is used to be inserted every preset number of rows K in the original file, and the preset number of rows K is adjusted according to the actual situation.

[0090] For example, as shown in Figures 2(a) and 2(b), 1. First, read the original file "Unit Donation Summary Table" to obtain the data attributes of the columns, including column names (which are serial number, donation unit, donation amount, contact person, contact phone number respectively. If there is no column name, obtain the serial number of the column), column category attributes (which are serial number, company name, amount, name, mobile phone number respectively), data type and format (which are integer, character type, two decimal places, character type, character type respectively).

[0091] 2. Randomly select several rows as samples. The specific number of rows selected should meet the requirements of the sampling sample, not too many or too few. For example, if the original file has 60 rows, then 10% of the number of rows can be selected. Then extract the sample data set with columns as the dimension. For example, the column sample data set extracted for the donation amount is:

[0092] C{8900.00,7950.00,3205.00,5600.00,3150.00,5000.00}.

[0093] 3. Forge a row of data. At this time, it is necessary to forge the serial number, donation unit, donation amount, contact person, and contact phone number respectively. Based on the analysis in step 1, the serial number can obtain the row number of the previous row plus 1 when inserting the pseudo-row, and then the subsequent row numbers are shifted backward. The donation unit can be randomly obtained by selecting the "company name" generation program in the "built-in common digital dictionary attribute library". The donation amount can be randomly generated as an integer amount with 2 decimal places within the minimum and maximum values of the above sample set C. Similarly, the contact person and contact phone number are also randomly generated by the built-in "name" and "mobile phone number" generation programs respectively. According to this method, several rows of pseudo-row data can be forged and form a pseudo-row data set R.

[0094] 4. Define an insertion strategy, such as taking a row of pseudo-row data from the data set R and inserting it every K rows. This K value should be appropriately selected. For example, if the original file has 60 rows, then K can take 10% of the total number of rows, that is, insert a row of pseudo-row data every 6 rows. If the K value is too small, it will affect the robustness, and if it is too large, it will lead to a large total embedding amount and thus affect the accuracy of the original data. For example, the serial numbers 7 and 13 are the inserted pseudo-row data.

[0095] In the embodiments of the present application, the principle of the pseudo-column watermark algorithm is that the forged data should consider various constraints, and should have the same type, format, reasonable value range as the original data in the column, and be as relevant as possible to other attributes of the file data, etc., so that the watermark is more concealed and authentic.

[0096] In one implementation, when the watermark algorithm combination includes the pseudo-column watermark algorithm, generating combined watermark data according to the watermark algorithm combination and the original file in step S300 includes steps S311 - S312:

[0097] S311. Determine the usage characteristics of the data in the original file, and determine the target attribute as the pseudo-column attribute name from the preset digital dictionary attribute library according to the usage characteristics.

[0098] Optionally, the usage characteristics include, but are not limited to, information such as the industry to which it belongs and the usage. After determining the usage characteristics of the data in the original file, determine the target attribute as the pseudo-column attribute name from the preset digital dictionary attribute library. It should be noted that the number of target attributes can be one or more, and the pseudo-column attribute name should be as relevant as possible to the other column attribute names in the file data relationship table, that is, relevant to the other column attribute names.

[0099] S312. Generate pseudo-column data for several columns according to the pseudo-column attribute name and the generation algorithm to obtain the third watermark data.

[0100] Optionally, set in advance a generation algorithm that provides built-in column values. The generation algorithm supports input of parameters, including the number of generated column values, value range, character length, numerical precision, and upper and lower limits of variance, etc. Thus, according to the pseudo-column attribute name and the parameters, generate pseudo-column data for several columns to obtain the third watermark data L0...Ln, where L0 is the pseudo-column data of the first column. Among them, the pseudo-column data is used to be inserted into the corresponding position of the original file according to the preset insertion position strategy.

[0101] In one implementation, when the watermark algorithm combination includes the desensitization watermark algorithm, generating combined watermark data according to the watermark algorithm combination and the original file in step S300 includes steps S321 - S323:

[0102] S321. Determine the attributes of the column data in the original file.

[0103] Similarly, the attributes of the column data include, but are not limited to, column name, column subscript, column category to which it belongs (such as mobile phone number, administrative region, age), data type (such as character type, numerical type), data format (such as the number of decimal points, character length), etc.

[0104] S322. Determine the target column to be desensitized according to the attributes of the column data.

[0105] Optionally, based on a preset matching rule, the attributes of column data can be used to calculate the similarity with sensitive content in the database, and columns with certain content greater than or equal to the similarity threshold are used as target columns to be desensitized; alternatively, the user can input a selection instruction by viewing the attributes of the column data, and in response to the selection instruction input by the user, directly determine the target columns to be desensitized.

[0106] S323. A sensitive algorithm that matches corresponding sensitive words from a sensitive word library according to attributes, and desensitizes the target columns through the sensitive algorithm to obtain the fourth watermark data.

[0107] Optionally, match the corresponding sensitive words from the sensitive word library according to the attributes of the column data, then query the built-in sensitive algorithm corresponding to the sensitive words, and desensitize the target columns through the sensitive algorithm to obtain the fourth watermark data. It should be noted that the desensitization process includes but is not limited to masking, replacement, scrambling, rounding, and mapping. Among them, the fourth watermark data is used to replace the target columns in the original file.

[0108] For example, an example of the desensitized watermark algorithm:

[0109] 1. First, read the original file "Unit Donation Summary Table" in Figure 2(a) to obtain the attributes of the column data, including column names (which are serial number, donation unit, donation amount, contact person, contact phone number respectively. If there is no column name, obtain the serial number of the column), column category attributes (which are serial number, company name, amount, name, mobile phone number respectively), data types and formats (which are integer, character type, two decimal places, character type, character type respectively).

[0110] 2. Compare and match the attributes of the column data obtained from the above analysis with the built-in sensitive word library to identify the columns that need to be desensitized. In the case where there is no column category attribute, the columns that need to be desensitized can be directly selected manually and associated with the sensitive words in the sensitive word library, so as to confirm the columns that need to be desensitized.

[0111] 3. Multiple desensitization algorithms / sensitive algorithms are built in for each sensitive word, and algorithms such as masking, replacement, scrambling, rounding, and mapping can be selected as needed. For example, the first digit of the donation amount can be masked, and for the mobile phone number, an algorithm for replacing the middle four digits can be selected. For example, four digits in each phone number are replaced with other digits.

[0112] 4. Each desensitized column desensitizes the original value according to the selected desensitization strategy and replaces the original value. As Figure 3 shown is the effect after executing the desensitized watermark.

[0113] In one embodiment, step S400 writes the first watermark data and the combined watermark data into the original file to obtain the target file. Optionally, the first watermark data, the second watermark data, the third watermark data, and the fourth watermark data are sequentially written into the original file. Each watermark data is embedded in the writing process by directly replacing the content of the original file or creating a temporary file and then writing, so as to obtain the target file.

[0114] For example, 1) insert the first watermark data, i.e., zero-width characters, after the attribute values of the five attributes of the abstract information content, title, author, subject, and identifier; 2) insert a line of pseudo-line data every preset number of lines K in the original file. This line of pseudo-line data is randomly selected from several lines of pseudo-line data. In other embodiments, other insertion strategies can be set, not limited to inserting a line of pseudo-line data every preset number of lines K; 3) insert pseudo-column data into the corresponding positions of the original file according to the preset insertion position strategy. For example, insert one or more columns of pseudo-column data after the last column or the column in the preset order; 4) replace the content of the target column in the original file with the content of the fourth watermark data after desensitization processing, i.e., the content in the desensitized target column. Through the above filling, the final target file is obtained.

[0115] It should be noted that after obtaining the final target file, the original file is deleted to ensure that the original file does not land or remain, maximizing data security. Thus, the watermark data embedding process is completed, and the watermark file for external distribution, i.e., the target file, is output.

[0116] In one embodiment, the watermark embedding method of the embodiments of the present application may further include steps S510 - S520:

[0117] S510: Record the first watermark data and the combined watermark data line by line, recording the unique distribution serial number, the positions and times where the first watermark data and the combined watermark data are written into the original file, to obtain a log file.

[0118] Optionally, record the first watermark data and the combined watermark data line by line, recording the unique distribution serial number distributeID, the positions (line numbers or column numbers), times, watermark types, etc. where the first watermark data and the combined watermark data are written into the original file, and write them line by line into the log file to obtain the final log file.

[0119] S520: Store the log file in a specified directory, and save the distribution information of the distribution serial number to a relational database.

[0120] Then, store the log file in the specified directory, and save the distribution information of the distribution serial number distributeID, such as the corresponding distributor, distribution object, distribution time, distribution method strategy, etc., in a relational database.

[0121] It should be noted that when Filebeat monitors that there are incremental log files or content in this specified directory, it collects and forwards them to Elasticsearch for classification indexing and distributed storage, that is, stores the watermark data of each watermark algorithm, that is, the content of the first watermark data, the second watermark data, the third watermark data, and the fourth watermark data.

[0122] For example, the process data of generating each watermark data is recorded throughout the process and written into the log file line by line. Each line of log data includes the distribution serial number distributeID, the JSON-formatted watermark data, the insertion position of the watermark data (line number or column number), the watermark type, and the log generation time. After performing four hybrid watermark strategies of the invisible watermark algorithm, the pseudo-row watermark algorithm, the pseudo-column watermark algorithm, and the desensitization watermark algorithm on the original file "Unit Donation Summary Table" in Figure 2(a), the output watermark file, that is, the target file, is as Figure 4 shown, then four types of log data such as the red part in the file (the circled part is part of the red part) will be generated, and they all have the same distributeID, which are respectively:

[0123] ① Invisible watermark log:

[0124] T000000000000001||{}||{}||Invisible watermark||20231021203422

[0125] ② Pseudo-row watermark log:

[0126] T000000000000001||{7,"Zhanjiang Branch",4180.00,"Lao Qian","17822256212"}||{7}||Pseudo-row watermark||20231021203422 ......

[0128] T000000000000001||{13,"Zhongshan Branch",4930.00,"Lao Sun","17822267382"}||{13}||Pseudo-row watermark||20231021203428

[0129] ③ Pseudo-column watermark log:

[0130] T000000000000001||{"Email", "5341343@qq.com", "lihsi@163.com", "vk2023@126.com", "......", "vsnnd@wy.co m", "......", "ewsse@163.com", "......", "lhe32@126.com", "bisev@qq.com", "rdh21@189.com"}||{6}||Pseudo-column watermark||20231021203421

[0131] ④ Desensitized watermark log:

[0132] T000000000000001||{"Donation amount (yuan)", *900.00, *950.00, *205.00,......, *600.00, *150.00, *000.00}||{3}||Desensitized watermark||20231021203427

[0133] T000000000000001||{"Contact phone number", "18922562231", "18024372234", "13654520098",......, "18814293313", "13671135543", "18919884409"}||{5}||Desensitized watermark||20231021203425.

[0134] Refer to Figure 5 , the embodiment of the present application also provides a traceability method, including steps S600 - S800:

[0135] S600. Obtain the leaked content.

[0136] Among them, the leaked content is the target file or part of the content in the target file, and the target file is obtained through the watermark embedding method.

[0137] S700. Determine the distribution serial number of the leaked content through a preset traceability algorithm.

[0138] Optionally, when there is leaked content, determine the distribution serial number of the leaked content through a preset traceability algorithm. The preset traceability algorithm may include a watermark file traceability method and a keyword index traceability method. First, determine whether the target file itself has been obtained, that is, whether the leaked content is the target file. If not, that is, the leaked content is not the target file but part of the target file, then determine the distribution serial number of the leaked content through the keyword index traceability method; if the target file itself has been obtained, that is, the leaked content is the target file, then combine the watermark file traceability method and the keyword index traceability method to determine the distribution serial number of the leaked content.

[0139] For example, the watermark file traceability method: 1. Read the watermark file (i.e., the target file) and obtain and parse the zero-width characters embedded in the five attributes of the file digest; 2. Through the RSA decryption algorithm, decrypt to obtain the respective distribution serial numbers distributeID (denoted as the first sub-distribution serial number) in the above five attribute watermark information, and multiple first sub-distribution serial numbers form the first distributeID set.

[0140] For example, the keyword index traceability method: 1. Input a specified number of keywords in a set of leaked content. To balance the accuracy and efficiency of traceability, the specified number should be appropriate, and the specific number should ensure that it can cover the watermark data; 2. Use the above keywords as input parameters and asynchronously call the high-performance full-text index interface provided by Elasticsearch to quickly search for the watermark data log records of all content containing the keywords stored in Elasticsearch, that is, match the content of the stored first watermark data, second watermark data, third watermark data, and fourth watermark data with the keywords to determine the matching content, and then parse the matching content to parse out the distributeID corresponding to the log record (denoted as the second sub-distribution serial number), and obtain the second distributeID set through repeated matching of multiple keywords.

[0141] S800. Trace back according to the distribution serial number.

[0142] Among them, when the two methods are combined, take the intersection of the first distributeID set and the second distributeID set obtained according to the above two traceability methods to obtain the final target distributeID, and then query the relational database according to this target distributeID for the key information such as the data distributor, distribution object, distribution time, and distribution method corresponding to the target distributeID, and thus complete the watermark traceability process.

[0143] Otherwise, if only the keyword index tracing method is used, the second distributeID set is used as the target distributeID, and the key information such as the data distributor, distribution object, distribution time, and distribution method corresponding to the target distributeID in the relational database is queried to complete the watermark tracing process.

[0144] For example, suppose Figure 4 The red part in (the circled part is part of the red part) is the leaked content. At this time: 1. Watermark file tracing method: First read the target file, obtain the zero-width characters embedded in the five attribute values ​​of the file summary information, title, author, subject, and identifier, and parse out 5 distributed IDs through conversion and RSA decryption algorithm. At this time, deduplication is performed to obtain the possible first distributed ID set. 2. Keyword index tracing method: Regardless of whether the above target file is obtained, or only part of the leaked information is obtained, it can be traced through a set of keywords. Because the data in the red part is watermark data, this data has been collected by Filebeat and distributedly stored in Elasticsearch, so the keywords in the red part can be used to index the log records in Elasticsearch, thereby parsing the distributed ID of each log. Among them, in order to ensure the success rate of tracing, the number of keywords should be appropriate. If the number is too small, the data containing watermarks may not be obtained, and if it is too large, the efficiency will be affected. Therefore, the "return" point method can be adopted, and the values ​​of the four points of the upper, lower, left and right points can be obtained in a loop nested manner to ensure that the watermark data in the red part can be obtained. At this time, deduplication is performed based on the multiple distributeIDs parsed above to obtain a possible second distributeID set. Then, the first distributeID set and the second distributeID set are intersected to obtain the final target distributeID, and then the relational database is queried through the target distributeID to obtain key information such as the data distributor, distribution object, distribution time, and distribution method corresponding to the target distributeID, and finally the traceability is successfully completed.

[0145] By the method of the embodiment of the present application:

[0146] First of all, it can keep the data authentic and difficult to detect, and it is highly secure and not easily destroyed.

[0147] Secondly, four hybrid watermark algorithms, namely pseudo-row watermark, pseudo-column watermark, invisible watermark, and desensitized watermark, are adopted. At the same time, a built-in common digital dictionary attribute library (such as administrative regions, industry types, etc.), a sensitive word library, and multiple sensitive word desensitization algorithms are built to ensure the security, fidelity, and robustness of the watermark as much as possible, solving the problem of single watermark algorithm. Moreover, the invisible watermark is realized by implanting invisible zero-width characters in five attributes of the file summary information, with high concealment. The forged data of the pseudo-row, pseudo-column, and desensitized watermarks are realized through matching with the built-in digital dictionary attribute library, sensitive word library, and sensitive word desensitization algorithms, providing rich options for watermark strategy execution and ensuring the fidelity and secrecy of the watermark.

[0148] Then, in conjunction with the use of relational databases, Filebeat, and Elasticsearch tools, the watermark data logs generated during the watermark embedding process are recorded, collected, and classified and stored throughout the process. The watermark information can be analyzed and indexed by reading the watermark file or inputting the keyword of the leaked content, ensuring the success rate and high efficiency of watermark traceability as much as possible, and solving the problems of low traceability success rate and poor efficiency.

[0149] Finally, the watermark embedding, watermark traceability technologies and ideas are practical and feasible, with strong practicability.

[0150] Refer to Figure 6 , which shows the structural block diagram of the watermark embedding device according to an embodiment of the present application. The device may include:

[0151] An acquisition module, configured to acquire an original file;

[0152] A first generation module, configured to generate first watermark data according to the original file and the invisible watermark algorithm;

[0153] A second generation module, configured to determine a watermark algorithm combination according to the file type of the original file and a preset rule, and generate combined watermark data according to the watermark algorithm combination and the original file. The watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitized watermark algorithm;

[0154] A writing module, configured to write the first watermark data and the combined watermark data into the original file to obtain a target file.

[0155] In an implementation manner, the watermark embedding device further includes a recording module, and the recording module is configured to:

[0156] Record the first watermark data and the combined watermark data line by line, record the unique distribution serial number, the positions and times when the first watermark data and the combined watermark data are written into the original file, to obtain a log file. The distribution serial number is generated during the generation process of the first watermark data;

[0157] Store the log file in the specified directory and save the distribution information of the distribution serial number in a relational database.

[0158] For the functions of each module in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, which will not be elaborated herein.

[0159] In one implementation, an embodiment of the present application further provides a traceability device, including:

[0160] An acquisition module, configured to acquire leaked content, where the leaked content is a target file or a part of the target file, and the target file is obtained by a watermark embedding method;

[0161] A determination module, configured to determine the distribution serial number of the leaked content through a preset traceability algorithm;

[0162] A traceability module, configured to perform traceability according to the distribution serial number.

[0163] For the functions of each module in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, which will not be elaborated herein.

[0164] Refer to Figure 7 , which shows the structural block diagram of an electronic device according to an embodiment of the present application. The electronic device includes: a memory 310 and a processor 320. Instructions that can run on the processor 320 are stored in the memory 310, and the processor 320 loads and executes the instructions to implement the watermark embedding method or the traceability method in the above embodiments. Among them, the number of the memory 310 and the processor 320 can be one or more.

[0165] In one implementation, the electronic device further includes a communication interface 330, configured to communicate with external devices and perform data interaction and transmission. If the memory 310, the processor 320, and the communication interface 330 are implemented independently, the memory 310, the processor 320, and the communication interface 330 can be interconnected through a bus and complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.

[0166] Optionally, in a specific implementation, if the memory 310, the processor 320, and the communication interface 330 are integrated on a single chip, the memory 310, the processor 320, and the communication interface 330 can communicate with each other through an internal interface.

[0167] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the watermark embedding method or the traceability method provided in the above embodiments.

[0168] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.

[0169] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute code in the memory, and when the code is executed, the processor is configured to execute the method provided in the embodiment of the application.

[0170] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced RISC machines (ARM) architecture.

[0171] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0173] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0174] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.

[0175] Any process or method description represented in a flowchart or described in other ways herein can be understood as representing a module, segment or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed.

[0176] The logic and / or steps represented in a flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in combination with these instruction execution systems, apparatus or devices.

[0177] It should be understood that each part of this application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.

[0178] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk or an optical disc, etc.

[0179] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A watermark embedding method, characterized in that, Including: Obtain the original file; Generate first watermark data according to the original file and the invisible watermark algorithm; Determine a watermark algorithm combination according to the file type of the original file and a preset rule, and generate combined watermark data according to the watermark algorithm combination and the original file, where the watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm; Write the first watermark data and the combined watermark data into the original file to obtain a target file.

2. The watermark embedding method according to claim 1, wherein: The generating first watermark data according to the original file and the invisible watermark algorithm includes: Read the digest information of the original file; Generate a unique distribution serial number, and encrypt the distribution serial number through an encryption algorithm to obtain an encrypted distribution serial number; Convert the encrypted distribution serial number into a zero-width character to obtain first watermark data; Wherein, the first watermark data is used to be written into the digest information.

3. The watermark embedding method according to claim 1 or 2, characterized in that: When the watermark algorithm combination includes the pseudo-row watermark algorithm, the generating combined watermark data according to the watermark algorithm combination and the original file includes: Extract several rows of data from the original file as row samples; Extract the data of each column from the row samples to obtain column data of several columns, and determine the attributes of the column data; Forge pseudo-row data of several rows according to the attributes of the column data to obtain second watermark data; Wherein, each row of pseudo-row data is respectively used to be inserted in the original file at intervals of a preset number of rows.

4. The watermark embedding method according to claim 1 or 2, characterized in that: When the watermark algorithm combination includes the pseudo-column watermark algorithm, the generating combined watermark data according to the watermark algorithm combination and the original file includes: Determine the usage characteristics of the data in the original file, and determine a target attribute as the pseudo-column attribute name from a preset digital dictionary attribute library according to the usage characteristics; Generate pseudo-column data of several columns according to the pseudo-column attribute name and a generation algorithm to obtain third watermark data; Wherein, the pseudo-column data is used to be inserted into the corresponding position of the original file according to a preset insertion position strategy.

5. The watermark embedding method according to claim 1 or 2, characterized in that: When the watermark algorithm combination includes the desensitization watermark algorithm, the generating combined watermark data according to the watermark algorithm combination and the original file includes: Determine the attributes of the column data in the original file; Determine a target column to be desensitized according to the attributes of the column data; Match a sensitive algorithm for corresponding sensitive words from a sensitive word library according to the attributes, and desensitize the target column through the sensitive algorithm to obtain fourth watermark data; Wherein, the fourth watermark data is used to replace the target column in the original file.

6. The watermark embedding method according to claim 2, characterized in that: The method further includes: Record the first watermark data and the combined watermark data row by row, record the unique distribution serial number, the positions and times when the first watermark data and the combined watermark data are written into the original file, to obtain a log file, where the distribution serial number is generated during the generation process of the first watermark data; Store the log file in a specified directory, and save the distribution information of the distribution serial number to a relational database.

7. A traceability method, characterized in that, Including: Obtain the leaked content, where the leaked content is the target file or part of the content in the target file, and the target file is obtained by the watermark embedding method described in any one of claims 1-6; Determine the distribution serial number of the leaked content through a preset traceability algorithm; Perform traceability according to the distribution serial number.

8. A watermark embedding device, characterized in that Comprising: An acquisition module, configured to acquire the original file; A first generation module, configured to generate first watermark data according to the original file and the invisible watermark algorithm; A second generation module, configured to determine a watermark algorithm combination according to the file type of the original file and a preset rule, and generate combined watermark data according to the watermark algorithm combination and the original file, where the watermark algorithm combination includes at least one of a pseudo-row watermark algorithm, a pseudo-column watermark algorithm, and a desensitization watermark algorithm; A writing module, configured to write the first watermark data and the combined watermark data into the original file to obtain a target file.

9. An electronic device, characterized in that, Comprising: A processor and a memory, where instructions are stored in the memory, and the instructions are loaded and executed by the processor to implement the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, the method described in any one of claims 1-7 is implemented.

Citation Information

Cited By

  • Watermark embedding method and device, storage medium and computer readable storage medium

    CN120805114A

  • Data transmission method and device applied to data security and electronic equipment

    CN122205007A