File watermark embedding and extracting method, system and equipment based on null pseudo column and medium
By clearing the empty cells at the end of the CSV file and embedding the watermark using binary encoding, the problems of watermark error detection and format non-compliance in the CSV file are solved, and the accurate transmission of watermark information and the compatibility and concealment of the file are achieved.
Patent Information
- Application Number
- CN202510654187.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-23
AI Technical Summary
The existing technology lacks an effective error detection mechanism in CSV files, which leads to errors in watermark information during transmission. When processing empty cells, the original data structure may be destroyed or the file format may not meet the standards, affecting the versatility and compatibility of the file.
Through the method based on empty pseudo-columns, the empty cells at the end of each row of the CSV document are cleared, the valid data columns are retained, the watermark content is converted into a binary ASCII code string, and empty cells are added or not added at the end of the row to represent 0 or 1. Combined with the parity bit and cyclic redundancy check code, a CSV file with an invisible watermark is generated.
It realizes the accurate transmission and detection of watermark information, improves the versatility and compatibility of files, enhances the concealment and anti-attack ability of watermarks, and ensures that files can be opened and displayed normally in various CSV processing software.
Smart Images

Figure CN120689188A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital watermarking, and in particular relates to a file watermark embedding and extraction method, system, device and medium based on empty pseudo-columns. Background Art
[0002] CSV is a data input format widely used in fields and scenarios such as data exchange, data analysis, database import / export, and internet data sharing. With the development of a digital and data-driven economy and the increase in data leaks and cyberattacks, data has become a critical enterprise asset, and the demand for data copyright protection is also increasing.
[0003] Related technologies, such as CSV file watermark embedding methods based on empty pseudo-columns, often lack error detection mechanisms during watermark transmission. During file storage and transmission, errors in the watermark binary string may occur due to factors such as disk failure and network interference. Existing technologies are unable to effectively identify and correct these errors, resulting in inaccurate or even completely erroneous extracted watermark information, affecting the proper use of the watermark.
[0004] Related technologies, when processing CSV documents, lack precise handling of empty cells. Some methods may mistakenly delete empty cells other than the end of the file when clearing empty cells, destroying the original data structure. Alternatively, when generating CSV files containing watermarks, the file format is not optimized, resulting in non-compliant content. For example, escape characters are not added to cells containing quotation marks or commas, and line break formats are inconsistent across different rows. This can cause the generated file to not open properly or display incorrectly in some CSV processing software, reducing the file's versatility and compatibility. Summary of the Invention
[0005] The present invention provides a file watermark embedding and extraction method based on empty pseudo columns. The present invention embeds and extracts watermarks in CSV files based on empty pseudo columns. Unlike traditional watermark adding methods that change original data, the accuracy of the data in the CSV file after watermarking using this method can still be guaranteed.
[0006] Methods include; S101: Clear the empty cells at the end of each row of the CSV document according to the number of columns in the document, and retain the columns with valid data; S102: Convert the watermark content to be embedded into a binary ASCII code string, where each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end; S103: Calculate the number of valid lines N of the pre-processed document. If the length L of the watermark binary string satisfies L≤N, execute step 104. Otherwise, trigger the mechanism of rejecting embedding or supplementing blank lines. S104: traverse each bit of the binary ASCII code string in order, and if the current bit is 1, add at least one empty cell at the end of the corresponding row; If it is 0, the row will be kept without adding empty cells; S105: Recombining the processed data row with the original table header to generate a CSV file containing an invisible watermark.
[0007] Preferably, step S101 specifically includes: Determine the number of valid data columns M according to the header of the CSV document; Parse each row of data and traverse forward from the last cell in each row. If multiple consecutive cells are empty, only the empty cells after the last non-empty cell are retained as part of the valid column; Clear all the consecutive empty cells at the end that exceed the valid column number M, while retaining the empty cells at non-end positions to maintain the original data structure; The method for determining the effective number of columns M includes: By counting the maximum number of non-empty cells in all rows of the CSV document; Alternatively, it is directly determined by the number of columns in the table header; For each row of data, record the number of original data columns and clear only the consecutive empty cells at the end; If there are empty cells in a row that are not at the end, the empty cells are retained to maintain the integrity of the original data.
[0008] Preferably, in step S102, the conversion of the binary ASCII code string includes the following methods: The binary encoding bit number of each character is configured to be 7 or 8 bits; Non-ASCII characters are converted to Unicode and automatically filled to a fixed number of digits; The control character added to the end of the binary string includes at least one of a line break character, a terminator, or a custom identifier; After generating the binary ASCII code string, generating a check bit, wherein the generated check bit is an 8-bit binary code of each character plus a 1-bit parity check bit, and generating a 9-bit extended code; Add at least 4 bits of cyclic redundancy check code to the end of the complete binary string to form the final watermark binary string; The check bit and the original watermark binary code are embedded in the empty cell in step S104.
[0009] Preferably, the calculation of the number of valid rows N in step S103 includes the following steps: Exclude the header row of the CSV document and count the number of rows with more than 0 non-empty cells in the remaining data rows; If all cells of a row of data are empty, they are not counted in the valid row number N; The blank row supplement mechanism includes: adding at least one blank row at the end of the document, with the number of columns in the blank row being the same as the number of columns in the table header, and filling the blank row with a preset number of empty cells; In step S103, when L>N, the watermark is enhanced by: Divide the watermark binary string into M substrings, where M is an integer greater than 1; Perform the embedding operation of step S104 on each substring, and the embedding position of different substrings is determined by a pseudo-random sequence; The substring length satisfies Σ≤N×K, where K is the preset redundancy coefficient; The blank line supplement mechanism in step S103 includes: According to the length L of the watermark binary string and the current number of valid lines N, the number of blank lines to be added is calculated as N 补 =ceil(L / N×S)-N, where S is the preset safety redundancy factor; The supplementary blank lines generate pseudo-random content through an encryption algorithm.
[0010] Preferably, the method of adding empty cells in step S104 includes: When the binary bit is 1, the number of empty cells added at the end of the row is dynamic, and the specific addition method is: k=(current bit index mod m)+1k, where m is the preset maximum redundancy value; The embedding positions of empty cells are scattered in the following way: Map binary bits to non-continuous rows based on a hash function, and the hash seed is generated by the watermark content; Each binary bit occupies n rows, n ≥ 1, forming a redundant storage structure; The embedding process uses a symmetric encryption algorithm to obfuscate the binary string. The encryption key is generated by the hash value of a specific column in the first row of the document; the encrypted binary bits are mapped to the corresponding rows in random order.
[0011] Preferably, the reassembly process in step S105 includes header watermark marking, adding a preset number of empty cells at the end of the original header row to identify the watermark document version, and the number of empty cells is mapped to the version number; When generating the CSV file in step S105 , format optimization processing is performed, and empty cells at non-end positions in the data row are deleted, escape characters are automatically added to cells containing quotation marks or commas, and the format of the line break at the end of the row is unified.
[0012] Preferably, in step S105, generating a CSV file containing an invisible watermark includes the following steps: Add a hidden checksum column at the end of the original header, which contains a hash value based on the header content; Combine the processed data rows with the original header in the order of original header, check column, and data rows; If there are merged cells in the original table header, redundant blank rows are inserted at the boundaries of the merged cells and filled with the preset encrypted data; The CSV file generation process includes: Encrypt the original header in blocks, using a different key for each block; Alternate the encrypted headers and data rows to form a hierarchical structure; A digital signature is appended to the end of the file. The signature is generated based on the XOR result of the hash value of the original header and the watermark binary string.
[0013] The present application also provides a file watermark embedding and extraction system based on empty pseudo-columns, the system comprising: The data column processing module is used to clear the empty cells at the end of each row of the CSV document according to the number of columns in the document, and retain the valid data columns; A conversion module, used to convert the watermark content to be embedded into a binary ASCII code string, where each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end; The valid line calculation module is used to calculate the number of valid lines N of the pre-processed document. When the length L of the watermark binary string satisfies L≤N, the end-of-line processing module is executed. Otherwise, the embedding rejection or blank line supplement mechanism is triggered. The end-of-row processing module is used to traverse each bit of the binary ASCII code string in order. If the current bit is 1, at least one empty cell is added to the end of the corresponding row; if it is 0, the row is kept without adding any new empty cells. The file generation module is used to recombine the processed data rows with the original table header to generate a CSV file with an invisible watermark.
[0014] According to another embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the file watermark embedding and extraction method based on empty pseudo-columns are implemented.
[0015] According to another embodiment of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the file watermark embedding and extraction method based on empty pseudo-columns are implemented.
[0016] It can be seen from the above technical solutions that the present invention has the following advantages: The empty pseudo-column-based file watermark embedding and extraction method provided in this application uses parity bits and cyclic redundancy check codes to effectively detect errors in the watermark binary string during transmission. Once an error is detected, prompts can be given or corrective measures can be taken to ensure the accuracy of the extracted watermark information. The substring redundancy embedding strategy ensures that even if some data is damaged, the watermark information can still be recovered from other substrings, thereby improving the reliability of the watermark. Diverse encoding methods make the watermark binary string more difficult to analyze and identify. The complex embedding strategy makes the distribution of empty cells lose its obvious pattern. Combined with encryption and obfuscation processing, it improves the concealment of the watermark and effectively prevents the watermark from being easily discovered and cracked by attackers.
[0017] When removing empty cells at the end of each row in a CSV document, the number of valid data columns, M, is determined based on the table header. Only the trailing consecutive empty cells that exceed this number are removed, while all empty cells at the end are retained. When generating a CSV file, format optimization is performed to remove empty cells at the end of the data row, automatically add escape characters to cells containing quotation marks or commas, and unify the format of end-of-row line breaks. This prevents file content errors caused by data structure corruption. This format optimization ensures that the generated watermarked CSV file strictly adheres to the CSV file format standard, ensuring that the file can be properly opened, displayed, and edited in various CSV processing software, improving the file's versatility and compatibility. The number of valid columns, M, can be determined by counting the maximum number of non-empty cells or directly based on the number of columns in the table header, providing two flexible determination methods. Supporting dynamic configuration of the binary encoding bit count makes this method applicable to CSV documents of varying formats and data volumes. Whether documents have sufficient or insufficient rows, efficient and accurate watermark embedding is achieved, enhancing the applicability of the watermark embedding method. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of the file watermark embedding and extraction method based on empty pseudo-columns; Figure 2 Schematic diagram of the file watermark embedding and extraction system based on empty pseudo-columns; Figure 3 Schematic diagram of an electronic device. DETAILED DESCRIPTION
[0020] This application provides a file watermark embedding and extraction method based on empty pseudo-columns, which enables CSV watermark embedding without modifying the original data or affecting the user experience. This method embeds watermarks by adding a number of empty cells at the end of each row in the CSV document, using the number of empty cells to represent 0 or 1. The watermark is then extracted by identifying the number of empty cells at the end of each row in the CSV document.
[0021] The following describes in detail the specific steps of the file watermark embedding and extraction method based on empty pseudo-columns involved in this application. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are provided to facilitate a thorough understanding of the embodiments of this application. However, it should be clear to those skilled in the art that this application can also be implemented in other embodiments without these specific details.
[0022] It should be understood that when used in this specification, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their collections. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0023] The phrases "one embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the application. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in other embodiments," etc. that appear in different places in this application do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] See also Figure 1 FIG2 is a flowchart of a method for embedding and extracting a file watermark based on an empty pseudo-column in a specific embodiment. The method includes: S101: Clear the empty cells at the end of each row of the CSV document according to the number of columns in the document, and retain the columns with valid data.
[0026] In some embodiments, a CSV document can be read and the data can be parsed line by line. By traversing from the right side of each row of data to the left, when the first non-empty cell is detected, the cell and all data to its left are retained as valid data, the remaining empty cells to the right are deleted, and the total number of columns in the document is recorded.
[0027] In some specific embodiments, the number of valid data columns M is determined based on the header of the CSV document. For each row of data, the data may be traversed forward from the last cell in the row. If multiple consecutive cells are empty, only the empty cells after the last non-empty cell are retained as part of the valid columns. All consecutive empty cells beyond the valid number of columns M are removed, while the empty cells at non-last positions are retained to maintain the original data structure.
[0028] In this embodiment, the effective number of columns M is determined by counting the maximum number of non-empty cells in all rows of the CSV document, or directly determining the number of columns in the header.
[0029] When clearing empty cells in step S101 of this embodiment, accidental deletion of data is avoided in the following manner: for each row of data, the number of original data columns is recorded, and only the last consecutive empty cells are cleared; if there are empty cells at non-last positions in a row, that is, non-continuous last empty cells, the empty cells are retained to maintain the integrity of the original data.
[0030] This embodiment also parses the table header. Specifically, the header row of the CSV file is read to determine the number of columns, M, of the original data. For example, if the header row contains three columns, then M = 3. If there are any empty cells at the end of the header row, these empty cells are ignored, and only the number of columns in the header is used as the basis.
[0031] In this embodiment, the data rows are processed row by row, and each row of data is scanned forward from the last cell to find the position of the first non-empty cell.
[0032] For example, if the original data in a row is [1,"a",A, ,], the last non-empty cell is in column 3 (A), and the two empty cells at the end (columns 4 and 5) are cleared. If the number of non-empty cells in a row is less than the number of header columns M, only the cell up to the last non-empty cell in the row is retained, and the remaining empty cells at the end are cleared.
[0033] For example, if the table header has 3 columns and the data in a row is [4,"d",D], it will be retained as [4,"d",D] with no empty cell at the end.
[0034] In the non-end empty cell retention mode of this embodiment, if there are non-end empty cells in a row of data, these empty cells are retained to maintain the original data structure.
[0035] This embodiment also implements a dynamic column number adaptation mechanism, introducing dynamic column number detection. When the number of columns in the table header is inconsistent with the maximum number of columns in the actual data row, the maximum number of columns in the data row is used as the valid columns M. This method enhances robustness by dynamically adjusting to data fluctuations. This prevents data structure corruption caused by clearing empty cells. After clearing empty cells, an integrity check is performed on each row of data. If the number of columns in a row is less than the number of columns in the table header, it is marked as an abnormal row and triggers user confirmation. In this way, combined with data verification, the reliability of pre-processing before watermark embedding is enhanced, avoiding watermark embedding failures due to incomplete data.
[0036] S102: Convert the watermark content to be embedded into a binary ASCII code string, wherein each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end.
[0037] In some embodiments, each character in the watermark content is read sequentially, the decimal value corresponding to the character is searched in the ASCII code table, the decimal value is converted to an 8-bit binary number, and the converted binary numbers are concatenated in sequence. After all characters are converted, the binary code corresponding to the line break character is added to the end of the binary string.
[0038] In this embodiment, the binary-ASCII code string conversion method includes: the binary encoding bit number of each character can be dynamically configured to 7 or 8 bits. Non-ASCII characters are converted through Unicode encoding and automatically padded to a fixed number of bits. The control character appended to the end of the binary string includes at least one of a line feed character, a terminator, or a custom identifier.
[0039] The conversion process of this embodiment involves generating a check bit by appending a parity check bit after the binary encoding of each character; appending a CRC check code at the end of the complete binary string; and embedding the check bit and watermark data together in the CSV document.
[0040] The conversion process in this embodiment supports dynamic switching between multiple encoding rules. Specifically, ASCII, UTF-8, or GBK encoding can be automatically selected based on the watermark content character set, with different encoding rules applied to mixed character types. An encoding identifier is added to the beginning of the binary string. For example, "00" represents ASCII, and "01" represents Unicode.
[0041] This embodiment converts text-based watermark content into a standardized binary stream through configurable encoding rules, and simultaneously embeds control characters and verification information to construct watermark carrier data that matches the number of rows in the CSV document, ensuring the bidirectional reversibility of the watermark information at the encoding layer and the physical storage layer.
[0042] This embodiment also implements adaptive encoding, dynamically selecting an encoding strategy based on the CSV document's attributes. For documents written in pure English, 7-bit ASCII encoding is used to compress the watermark length. For documents containing multilingual characters, UTF-8 encoding is automatically switched and the code table offset is recorded. This embodiment can predict the optimal encoding combination using a machine learning model.
[0043] As can be seen, this embodiment converts recognizable watermark content into computer-processable binary data, establishes a correspondence between characters and binary data, and facilitates subsequent embedding of watermark information into CSV documents. Furthermore, the clear end marker helps accurately extract the watermark content and avoids data confusion.
[0044] S103: Calculate the number of valid lines N of the pre-processed document. When the length L of the watermark binary string satisfies L≤N, execute step 104. Otherwise, trigger the embedding rejection or blank line supplement mechanism.
[0045] In some embodiments, the number of valid lines N (i.e., the number of non-empty data lines) in the document after processing in step S101 is counted. The length L of the watermark binary string generated in step S102 is calculated and compared with N. If L ≤ N, then the number of document lines can accommodate the watermark information, and step 104 is continued. If L > N, according to preset rules, the watermark addition is rejected and the user is prompted, or a blank line is added to the end of the document until the number of document lines meets the conditions for embedding the watermark information. This ensures that the watermark information can be fully embedded in the CSV document, avoiding watermark loss or incompleteness due to insufficient document lines.
[0046] As an embodiment of step S103, the number of valid rows N is calculated by first excluding the header row of the CSV document and counting the number of rows in the remaining data rows whose number of non-empty cells is greater than 0. If all cells in a row of data are empty, it is not counted in the number of valid rows N. The blank row supplementation mechanism includes: adding at least one blank row at the end of the document, where the number of columns in the blank row matches the number of columns in the header, and the blank row is filled with a preset number of empty cells.
[0047] When L > N in step S103, the robustness of watermark embedding is enhanced as follows. The watermark binary string is divided into M substrings, where M is an integer greater than 1. The embedding operation of step S104 is performed on each substring, and the embedding position of each substring is determined using a pseudo-random sequence. The substring length satisfies Σ(substring length) ≤ N × K, where K is a preset redundancy factor, optionally K = 1.5 or 1.8.
[0048] Step S103 also involves a blank line supplement mechanism that can calculate the number of blank lines to be supplemented as N based on the length L of the watermark binary string and the current number of valid lines N. 补=ceil(L / N×S)-N, where ceil(L / N×S) represents rounding up the result of (L / N×S). S is a preset safety redundancy factor, optionally 1.2 or 1.5.
[0049] Optionally, the supplemented blank lines generate pseudo-random content through an encryption algorithm to increase the difficulty for attackers to identify and delete the content.
[0050] In this embodiment, the number of valid rows N is calculated by skipping the first row of the CSV document and counting only the rows. For each row of data, if all cells are empty, it is considered an invalid row and is not counted in N. The remaining rows are counted as N, ensuring that the watermark is embedded only in rows containing valid data.
[0051] This embodiment compares the watermark length with the number of lines, and compares the length L of the watermark binary string with N: if L≤N, step S104 is directly executed to embed the watermark; if L>N, the blank line supplement mechanism is triggered or the embedding is rejected.
[0052] The blank line supplement mechanism of this embodiment is to add a blank line at the end of the document, and the number of columns in the blank line is consistent with the table header. Dynamic supplementation is to calculate the number of blank lines to be supplemented based on the difference between L and N, ensuring that the total number of lines ≥ L.
[0053] As can be seen, the watermark binary string is split into multiple substrings, each of which is independently embedded in a different row interval. The substring's position is determined by a pseudo-random sequence. Even if some rows are deleted or tampered with, the remaining substrings can still restore the watermark, enhancing resistance to deletion attacks. When filling empty rows, an encryption algorithm is used to generate pseudo-random data to fill the empty cells, preventing direct deletion of these rows and destroying the watermark.
[0054] The redundancy coefficient K is dynamically adjusted according to the ratio of the number of document lines N and the watermark length L: while ensuring the watermark capacity, the damage to the document structure is reduced.
[0055] S104: Traverse each bit of the binary ASCII code string in order. If the current bit is 1, add at least one empty cell at the end of the corresponding row. If it is 0, keep the row without adding any empty cells.
[0056] In some embodiments, the traversal is performed sequentially, starting from the first bit of the binary ASCII code string generated in step S102. During the traversal process, each line of the document is mapped to each bit of the binary code string. When a bit in the binary code string is 1, at least one empty cell is added to the end of the corresponding line; when a bit is 0, no new empty cell is added to the line, and the line remains unchanged.
[0057] In this way, the presence or absence of empty cells at the end of each row in a CSV document represents the 0s and 1s in the binary code string, allowing the watermark information to be embedded in the document in an invisible manner. The state of the empty cell at the end of each row corresponds to one bit of data in the binary code string, thus ensuring the storage of the watermark information and ensuring the integrity and usability of the document. Furthermore, the watermark information is embedded through a simple rule for adding empty cells.
[0058] S105: Recombining the processed data row with the original table header to generate a CSV file containing an invisible watermark.
[0059] In some embodiments, after the watermark embedding operation in step S104 is completed, each row of processed data is arranged in its original order and then combined with the original header information of the document. The combined data is saved according to the format requirements of the CSV file to generate the final CSV file containing the invisible watermark.
[0060] The CSV file here consists of a header and data rows. The header identifies the meaning of each column of data. After the watermark is embedded, the processed data rows are reassembled with the original header to restore the original structure of the CSV file, making it readable and usable by relevant software. This ensures that the generated file still conforms to the CSV file format and can be opened and processed normally by various software that supports the CSV format. At the same time, the watermark information is fully integrated into the CSV file, achieving the purpose of watermark embedding.
[0061] In one embodiment of the present invention, based on step S104, a possible example is provided below to provide a non-limiting explanation of its specific implementation. The rule for adding empty cells in step S104 includes: when the binary bit is 1, the number of empty cells added at the end of the row is a dynamic value, specifically calculated using the formula k = (current bit index mod m) + 1k, where m is a preset maximum redundancy value. When the binary bit is 0, no empty cells are left at the end of the row.
[0062] In this embodiment, the embedding positions of empty cells are dispersed in the following manner: Based on the hash function, binary bits are mapped to non-contiguous rows. The hash seed is generated by the watermark content. Each binary bit occupies n rows (n ≥ 1), forming a redundant storage structure.
[0063] In step S104, differential embedding is also used for specific marked rows. Specifically, if the first cell of a row contains a preset identifier, the row embedding is skipped; and the number of empty cells in the comment row is forced to zero.
[0064] This embodiment establishes a dynamic bit-row mapping relationship between the binary watermark stream and the CSV row sequence, achieving watermark steganography by controlling the existence and number of empty cells at the end of the row, while maintaining the legality of the CSV document structure. A single binary bit is expanded to 2 bits of information, and the parity of the number of empty cells represents 1 bit, odd = 1, even = 0. The total number of empty cells represents another 1 bit, such as 3 = 1, 2 = 0, enabling a single row to store multiple bits of watermark information, increasing capacity. Multi-level encoding is used to increase the single-row capacity to 2 bits, breaking through the traditional 1-bit / row limit; this embodiment proposes multi-level encoding and dynamic redundancy; breaking through the limitation of conventional steganography that only modifies data content, and embedding through the dual dimensions of the number and position of empty cells.
[0065] Furthermore, as a refinement and expansion of the specific implementation of step S105 in the above embodiment, in order to fully illustrate the specific implementation process of step S105 in this embodiment, step S105 includes: generating a CSV file containing an invisible watermark, specifically including the following steps: A hidden checksum column containing a hash value based on the header contents is added to the end of the original header. The processed data rows are combined with the original header in the following order: original header - checksum column - data row. If the original header contains merged cells, redundant blank rows are inserted at the boundaries of the merged cells and filled with the preset encrypted data.
[0066] The file generation process in step S105 of this embodiment includes encrypting the original header in blocks, using a different key for each block. The encrypted header and data rows are arranged alternately to form a hierarchical structure. The specific structure is encrypted header block - data row - encrypted header block - data row. A digital signature is appended to the end of the file. This signature is generated by XORing the hash value of the original header with the watermark binary string.
[0067] In the combined CSV file format, if the original document contains multiple levels of headers, a hidden watermark row is inserted below each level of headers. The number of columns in the watermark row matches the number of columns in the header, and its content is the binary-encoded watermark length L, encoded using Base64 to avoid parsing errors.
[0068] This embodiment directly retains the first row of the original CSV to ensure that the data column names are not modified. The processed data rows are arranged in the original order and aligned with the number of columns in the header. Ensure that the generated CSV file conforms to the standard format to avoid parsing errors caused by the addition of empty cells. The empty cells added in step S104 will be retained at the end of the corresponding row to form an invisible watermark mark. It can be seen that a hidden watermark mark row is inserted below the header, and the mark row content is the Base64 encoding of the watermark length L, and the position is determined by a pseudo-random sequence. Even if some data rows are deleted, the watermark range can still be quickly located through the mark row. Redundant blank rows filled with encrypted data are inserted into the boundaries of the merged cells, and the number of columns is consistent with the header. Redundant blank rows are visually no different from ordinary blank rows, but the encrypted content can prevent attackers from destroying the watermark or header structure by deleting blank rows.
[0069] In some specific embodiments, as an alternative implementation of step S105, distinct from the above, the reassembly process in step S105 includes a header watermark, adding a preset number of empty cells at the end of the original header row to identify the watermarked document version, with the number of empty cells mapped to the version number. When generating a CSV file, format optimization is performed to remove empty cells at non-terminal positions in the data row; automatically add escape characters to cells containing quotation marks or commas; and unify line breaks at the end of rows to either CRLF or LF.
[0070] Embed invisible metadata in the final CSV file. That is, insert a hidden comment line after the first row header to record the watermark embedding time, hash value, and encoding parameters; the number of empty cells at the end of the comment line is forced to zero to avoid interfering with watermark extraction.
[0071] This embodiment also performs integrity verification on the generated file. It can calculate a lightweight hash value of the processed document, convert it into a binary string, and scatter-embed it into the empty cells at the end of the last three rows. The verification information is independent of the watermark content and is used to verify the integrity of the file during extraction.
[0072] In this way, the watermarked data rows and the original headers are reorganized in a standard CSV format. Through structural correction and metadata insertion, the machine readability and concealment of the watermarked document are ensured, while also providing traceable verification information. The invisible metadata provides watermark source, time, and integrity verification information.
[0073] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0074] As an implementation of the above method, this embodiment adds some empty cells at the end of each row of the CSV file, and uses the number of empty cells to represent 0 or 1, thereby realizing the watermark function.
[0075] For embedding, empty cells at the end of each row are removed based on the number of columns in the document. The content to be embedded is converted into a 01 string. Based on the length of the 01 string, the number of rows in the document is evaluated. If the document has too few rows, the watermark is rejected or padded with blank rows. Based on the number of rows and the 01 string, empty cells are added at the end of each row to represent the watermark information.
[0076] For extraction, the number of empty cells in each row is calculated based on the number of columns containing data in the watermark document. This number is converted into a 01 string. This 01 string is then converted into a watermark string. The watermark content is determined based on the watermark string (which is often multiple repetitions of the watermark content).
[0077] The first row of a CSV file typically serves as a header. This implementation does not use the header to store watermark information. In this example, the watermark is added only once, eliminating redundancy and attack resistance. If there are no empty cells at the end of a row, the value is considered 0; if there are empty cells at the end of a row, the value is considered 1 (regardless of the number of empty cells).
[0078] Known: The binary ASCII code of letter C is 01100011.
[0079] The binary ASCII code of the letter o is 01101111.
[0080] The binary ASCII code of the letter m is 01101101.
[0081] The binary ASCII code for the line feed character is 00001010.
[0082] Assume that the watermark content to be embedded is com (ending with a newline character), and assume that a CSV document is as follows (a total of 33 lines, including the header): first column, second column, third column: 1,"a",A,, 2,"b",B,, 3,"c",C, 4,"d",D 5,"e",E 6,"f",F,, 7,"g",G 8,"h",H 9,"i",I, 10,"j",J 11,"k",K, 12,"l",L 13,"m",M,, 14,"n",N 15,"o",O, 16,"p",P 17,"q",Q,,,, 18,"r",R 19,"s",S, 20,"t",T 21,"u",U,, 22,"v",V 23,"w",W,,,,, 24,"x",X 25,"y",Y 26,"z",Z, 27,"a",A 28,"b",B, 29,"c",C 30,"d",D 31,"e",E,, 32,"f",F Embedding is based on the number of columns in the document. The empty cells at the end of each row of the document are cleared, and the CSV document becomes: first column, second column, third column: 1,"a",A 2,"b",B 3,"c",C 4,"d",D 5,"e",E 6,"f",F 7,"g",G 8,"h",H 9,"i",I 10,"j",J 11,"k",K 12,"l",L 13,"m",M 14,"n",N 15,"o",O 16,"p",P 17,"q",Q 18,"r",R 19,"s",S 20,"t",T 21,"u",U 22,"v",V 23,"w",W 24,"x",X 25,"y",Y 26,"z",Z 27,"a",A 28,"b",B 29,"c",C 30,"d",D 31,"e",E 32,"f",F In this embodiment, the content to be embedded is converted into a 01 string, and the "com line break" becomes: "01100011011011110110110100001010".
[0083] Based on the length of the 01 string, the number of lines in the document is evaluated. If the number of lines in the document is too small, the watermark is rejected or supplemented with blank lines. In this example, the length of the 01 string is 32 and the length of the document is 33. The watermark can only be added once, and blank lines are no longer used for expansion in this embodiment. Based on the number of lines in the document and the 01 string, an empty cell is added at the end of each line to represent the watermark information. The CSV document becomes: First column, second column, third column 1,"a",A 2,"b",B, 3,"c",C,, 4,"d",D 5,"e",E 6,"f",F 7,"g",G, 8,"h",H,, 9,"i",I 10,"j",J, 11,"k",K,, 12,"l",L 13,"m",M,,,, 14,"n",N,,, 15,"o",O,, 16,"p",P, 17,"q",Q 18,"r",R,, 19,"s",S,, 20,"t",T 21,"u",U,, 22,"v",V,, 23,"w",W 24,"x",X, 25,"y",Y 26,"z",Z 27,"a",A 28,"b",B 29,"c",C, 30,"d",D 31,"e",E,,, 32,"f",F It should be noted that: 1. ASCII code is only one way to convert watermark content into 01 string, other conversion methods are also possible; 2. The example uses a line break as the end of the watermark, but other marks or symbols can also be used; 3. During the watermark embedding and extraction process, it is optional to use the first line to store the watermark information; 4. In the example shown in this embodiment, only one watermark is added, which has a poor ability to resist attacks. In real environments, there is usually a large amount of data, which has a certain resistance to data tampering. 5. In this embodiment, if there is no empty cell at the end of a row, the value is considered to be 0, and if there is an empty cell at the end of a row, the value is considered to be 1 (regardless of the number of empty cells). Other methods of representing 0 and 1 are also within the scope of protection of this patent.
[0084] As for the extraction method, 1. Calculate the number of empty cells in each row based on the number of columns containing data in the watermark document.
[0085] From the header, we can see that the number of columns in this watermark document is 3. The number of empty cells in each row is calculated as follows: 0,1,2,0,0,0,1,2,0,1,2,0,4,3,2,1,0,2,2,0,2,2,0,1,0,0,0,0,1,0,3,0.
[0086] 2. Convert to a 01 string based on the number of empty cells in each row. Since the agreed rule is "if there is no empty cell at the end of the row, it is considered 0, and if there is an empty cell at the end of the row, it is considered 1 (regardless of the number of empty cells)", the 01 string results are as follows: 0110 0011 0110 1111 0110 1101 0000 1010.
[0087] 3. Convert the 01 string to the watermark string. In this embodiment, the ASCII code table is used as the conversion method for the watermark content to the 01 string. The watermark content can be found as follows by querying the ASCII code table: "com newline character".
[0088] 4. Determine the watermark content based on the watermark string (the watermark string is often multiple repetitions of the watermark content).
[0089] In this embodiment, a line break is used as the end of the watermark, and it can be seen that there is only one complete watermark content in the watermark string. Therefore, the watermark content is: "com".
[0090] The following is an embodiment of a file watermark embedding and extraction system based on empty pseudo-columns provided by an embodiment of the present disclosure. This system and the file watermark embedding and extraction methods based on empty pseudo-columns in the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiment of the file watermark embedding and extraction system based on empty pseudo-columns, please refer to the embodiment of the file watermark embedding and extraction method based on empty pseudo-columns.
[0091] like Figure 2 As shown, the system includes: a data column processing module, which is used to clear the empty cells at the end of each row of the CSV document according to the number of columns in the document and retain the valid data columns.
[0092] The conversion module is used to convert the watermark content to be embedded into a binary ASCII code string, wherein each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end.
[0093] The valid line calculation module is used to calculate the number of valid lines N of the preprocessed document. When the length L of the watermark binary string satisfies L≤N, the end-of-line processing module is executed, otherwise the embedding rejection or blank line supplement mechanism is triggered.
[0094] The end-of-row processing module is used to traverse each bit of the binary ASCII code string in order. If the current bit is 1, at least one empty cell is added to the end of the corresponding row; if it is 0, the row is kept without adding any new empty cells.
[0095] The file generation module is used to recombine the processed data rows with the original table header to generate a CSV file with an invisible watermark.
[0096] like Figure 3 As shown, the present application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored in the memory and executable on the processor 101. When the processor 101 executes the program, the steps of the power transmission engineering GIM model parsing and loading method are implemented.
[0097] In the embodiments of the present invention, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0098] In the embodiment of the present application, the processor 101 can be implemented by using at least one of a special purpose integrated circuit, a programmable logic device, a field programmable gate array, a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to perform the functions described herein. In some cases, such an embodiment can be implemented in a controller. For software implementation, an embodiment such as a process or function can be implemented with a separate software module that allows the execution of at least one function or operation. The software code can be implemented by a software application (or program) written in any appropriate programming language, and the software code can be stored in a memory and executed by a controller.
[0099] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0100] The memory 102 can be used to store software programs and various data. The memory 102 can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0101] The present application also provides a storage medium having a computer program stored thereon, which implements the steps of the file watermark embedding and extraction method based on empty pseudo-columns when the computer program is executed by a processor.
[0102] The storage medium can be any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0103] In the context of storage media, a readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0104] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A file watermark embedding and extraction method based on empty pseudo-columns, characterized in that: Methods include; S101: Clear the empty cells at the end of each row of the CSV document according to the number of columns in the document, and retain the columns with valid data; S102: Convert the watermark content to be embedded into a binary ASCII code string, where each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end; S103: Calculate the number of valid lines N of the pre-processed document. If the length L of the watermark binary string satisfies L≤N, execute step 104. Otherwise, trigger the mechanism of rejecting embedding or supplementing blank lines. S104: traverse each bit of the binary ASCII code string in order, and if the current bit is 1, add at least one empty cell at the end of the corresponding row; If it is 0, the row will be kept without adding empty cells; S105: Recombining the processed data row with the original table header to generate a CSV file containing an invisible watermark.
2. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: Step S101 specifically includes: Determine the number of valid data columns M according to the header of the CSV document; Parse each row of data and traverse forward from the last cell in each row. If multiple consecutive cells are empty, only the empty cells after the last non-empty cell are retained as part of the valid column; Clear all the consecutive empty cells at the end that exceed the valid column number M, while retaining the empty cells at non-end positions to maintain the original data structure; The method for determining the effective number of columns M includes: By counting the maximum number of non-empty cells in all rows of the CSV document; Alternatively, it is directly determined by the number of columns in the table header; For each row of data, record the number of original data columns and clear only the consecutive empty cells at the end; If there are empty cells in a row that are not at the end, the empty cells are retained to maintain the integrity of the original data.
3. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: In step S102, the conversion of the binary ASCII code string includes the following methods: The binary encoding bit number of each character is configured to be 7 or 8 bits; Non-ASCII characters are converted to Unicode and automatically filled to a fixed number of digits; The control character added to the end of the binary string includes at least one of a line break character, a terminator, or a custom identifier; After generating the binary ASCII code string, generating a check bit, wherein the generated check bit is an 8-bit binary code of each character plus a 1-bit parity check bit, and generating a 9-bit extended code; Add at least 4 bits of cyclic redundancy check code to the end of the complete binary string to form the final watermark binary string; The check bit and the original watermark binary code are embedded in the empty cell in step S104.
4. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: The calculation of the number of valid rows N in step S103 includes the following steps: Exclude the header row of the CSV document and count the number of rows with more than 0 non-empty cells in the remaining data rows; If all cells of a row of data are empty, they are not counted in the valid row number N; The blank row supplement mechanism includes: adding at least one blank row at the end of the document, with the number of columns in the blank row being the same as the number of columns in the table header, and filling the blank row with a preset number of empty cells; In step S103, when L>N, the watermark is enhanced by: Divide the watermark binary string into M substrings, where M is an integer greater than 1; Perform the embedding operation of step S104 on each substring, and the embedding position of different substrings is determined by a pseudo-random sequence; The substring length satisfies Σ≤N×K, where K is the preset redundancy coefficient; The blank line supplement mechanism in step S103 includes: According to the length L of the watermark binary string and the current number of valid lines N, the number of blank lines to be added is calculated as N 补 =ceil(L / N×S)-N, where S is the preset safety redundancy factor; The supplementary blank lines generate pseudo-random content through an encryption algorithm.
5. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: The method of adding an empty cell in step S104 includes: When the binary bit is 1, the number of empty cells added at the end of the row is dynamic, and the specific addition method is: k=(current bit index mod m)+1k, where m is the preset maximum redundancy value; The embedding positions of empty cells are scattered in the following way: Map binary bits to non-continuous rows based on a hash function, and the hash seed is generated by the watermark content; Each binary bit occupies n rows, n ≥ 1, forming a redundant storage structure; The embedding process uses a symmetric encryption algorithm to obfuscate the binary string. The encryption key is generated by the hash value of a specific column in the first row of the document; the encrypted binary bits are mapped to the corresponding rows in random order.
6. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: The reassembly process in step S105 includes header watermark marking, adding a preset number of empty cells at the end of the original header row to identify the watermark document version, and the number of empty cells is mapped to the version number; When generating the CSV file in step S105 , format optimization processing is performed, and empty cells at non-end positions in the data row are deleted, escape characters are automatically added to cells containing quotation marks or commas, and the format of the line break at the end of the row is unified.
7. The file watermark embedding and extraction method based on empty pseudo-columns according to claim 1 is characterized in that: In step S105, generating a CSV file containing an invisible watermark includes the following steps: Add a hidden checksum column at the end of the original header, which contains a hash value based on the header content; Combine the processed data rows with the original header in the order of original header, check column, and data rows; If there are merged cells in the original table header, redundant blank rows are inserted at the boundaries of the merged cells and filled with the preset encrypted data; The CSV file generation process includes: Encrypt the original header in blocks, using a different key for each block; Alternate the encrypted headers and data rows to form a hierarchical structure; A digital signature is appended to the end of the file. The signature is generated based on the XOR result of the hash value of the original header and the watermark binary string.
8. A file watermark embedding and extraction system based on empty pseudo-columns, characterized in that: The system is used to implement the file watermark embedding and extraction method based on empty pseudo-columns as described in any one of claims 1 to 7; The system includes; The data column processing module is used to clear the empty cells at the end of each row of the CSV document according to the number of columns in the document, and retain the valid data columns; A conversion module, used to convert the watermark content to be embedded into a binary ASCII code string, where each character corresponds to an 8-bit binary number and a binary code of a line break is appended at the end; The valid line calculation module is used to calculate the number of valid lines N of the pre-processed document. When the length L of the watermark binary string satisfies L≤N, the end-of-line processing module is executed. Otherwise, the embedding rejection or blank line supplement mechanism is triggered. The end-of-row processing module is used to traverse each bit of the binary ASCII code string in order. If the current bit is 1, at least one empty cell is added to the end of the corresponding row; if it is 0, the row is kept without adding any new empty cells. The file generation module is used to recombine the processed data rows with the original table header to generate a CSV file with an invisible watermark.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the file watermark embedding and extraction method based on empty pseudo-columns as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the file watermark embedding and extraction method based on empty pseudo-columns as claimed in any one of claims 1 to 7 are implemented.