Watermark processing method and system for comma separation value file and medium

By clearing unnecessarily escaped characters and adjusting word frequency embedding watermarks, the problem of traditional CSV file watermark addition method destroying data availability is solved, and the watermark concealment and data integrity are achieved.

CN120068028APending Publication Date: 2025-05-30SHANDONG ZHICHUANG DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225640.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional watermark addition method of modifying the original CSV file consisting of English and numbers will destroy the availability of the original data and cannot be directly used in actual production processes such as machine learning.

Method used

By clearing preset non-escaping characters in the file, counting and adjusting word frequency, and embeding watermarks according to the preset character table, ensuring the concealment of the watermark and data integrity.

Benefits of technology

It implements the embedding of watermarks to maintain data integrity and availability without changing the actual value and structure of the original data, and solves the problem of traditional methods destroying data availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068028A_ABST
    Figure CN120068028A_ABST
Patent Text Reader

Abstract

The invention discloses a watermark processing method and system for a comma separation value file and a medium, mainly relates to the technical field of watermark processing, and is used for solving the problems that the availability of original data can be damaged and the original data cannot be directly used as input in the actual production process such as machine learning when the original file is modified traditionally. Comprising the following steps: clearing a preset unnecessary escape symbol in a file; counting the occurrence word frequency of words except the content contained in the preset necessary escape characters in the file, and determining the total word frequency corresponding to each character according to the sequence of the characters in the preset character table; determining the word frequency of the watermark character; the word frequencies of the watermark characters are adjusted, so that the word frequencies of the watermark characters are arranged from large to small according to the sequence of the watermark characters; and according to the word frequency of the adjusted watermark character and the word corresponding to the watermark character, adding a preset unnecessary escape character on the word with the adjusted word frequency in the current file to obtain a comma separation value file with the English watermark added.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of comma-separated value file applications, and in particular, to a watermark processing method, system, and medium for comma-separated value files. Background Art

[0002] Due to the development of machine learning and the requirements of data security, the need to add watermarks to symbol-separated value file data such as comma-separated value files (csv) composed of English and numbers is becoming increasingly strong. Among numerous symbol-separated value file formats, comma-separated value files composed of English and numbers are the most widely used.

[0003] Traditional csv watermark addition methods modify the original file composed of English and numbers, mainly including: adding pseudo rows (usually inserting some rows that do not belong to the real business data into the original CSV file data), adding pseudo columns (involving adding one or more columns based on the original column structure of the CSV file. The newly added columns will be filled with corresponding pseudo data, such as encrypted strings generated according to a certain algorithm, or codes representing information such as the file source and version), floating numerical attribute values (operating on the original numerical data in the file and making small increases or decreases within a reasonable range that does not affect the use and understanding of the numerical value in the normal business scenario), and adding spaces before and after text attribute values (adding spaces at the beginning or end of the text content in the CSV file. These spaces may not be added randomly and may be executed according to specific rules such as the number and order of spaces).

[0004] However, the traditional operation of modifying the original file composed of English and numbers will damage the usability of the original data and cannot be directly used as input in actual production processes such as machine learning. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, this application provides a watermark processing method, system, and medium for comma-separated value files to solve the problem that the traditional operation of modifying the original file composed of English and numbers will damage the usability of the original data and cannot be directly used as input in actual production processes such as machine learning.

[0006] In a first aspect, this application provides a watermark processing method for comma-separated value files, and the method includes: When performing the watermark addition operation, traverse the current file and clear the preset non-essential escape characters in the file; count the word frequencies of the words outside the content included in the preset essential escape characters in the file, and obtain the word sorting; based on the word sorting, determine the characters corresponding to each word in the order of the characters in the preset character table, and then determine the total word frequency corresponding to each character; obtain the watermark characters and the watermark character order involved in the English watermark, and determine that the total word frequency corresponding to the same characters of the watermark characters in the preset character table is the word frequency of the watermark characters; according to the watermark character order, adjust the word frequencies of the watermark characters so that the word frequencies of the watermark characters are arranged from largest to smallest in the watermark character order; according to the adjusted word frequencies of the watermark characters and the words corresponding to the watermark characters, add the preset non-essential escape characters to the adjusted word frequency number of words in the current file to obtain a comma-separated value file with the English watermark added; when performing the watermark extraction operation, traverse the current file, determine the correspondence between the words and the characters in the preset character table, and then read the word frequencies of each word included in the non-essential escape characters in the file. Based on the correspondence between the words and the characters, obtain the total word frequency corresponding to the characters, and arrange the characters in the order of the total word frequency to obtain the watermark.

[0007] In an implementation manner of the present application, the preset non-essential escape character is that when any field does not contain a comma but is enclosed by double quotes, it is considered that the double quotes enclosing the field are non-essential escape characters.

[0008] In an implementation manner of the present application, when performing the watermark addition operation, traversing the current file and clearing the preset non-essential escape characters in the file specifically include: Traverse the current file, determine whether the content enclosed by double quotes contains a comma, and when it does not contain a comma, clear the double quotes.

[0009] In an implementation manner of the present application, based on the word sorting, determining the characters corresponding to each word in the order of the characters in the preset character table, and then determining the total word frequency corresponding to each character specifically include: Obtain the preset character table, and then obtain the order of the characters in the preset character table. According to the order of the characters, poll to determine the characters corresponding to each word in the word sorting; According to the word frequencies of all the words corresponding to the current character, sum them up to obtain the total word frequency corresponding to the current character.

[0010] In an implementation manner of the present application, according to the watermark character order, adjusting the word frequencies of the watermark characters so that the word frequencies of the watermark characters are arranged from largest to smallest in the watermark character order specifically include: According to the watermark character order, obtain the word frequency corresponding to the next watermark character; when the word frequency corresponding to the next watermark character is greater than or equal to the current watermark character, modify the word frequency corresponding to the next watermark character to be less than the word frequency of the current watermark character.

[0011] In one implementation of the present application, according to the word frequency of the adjusted watermark characters and the words corresponding to the watermark characters, add a preset non-essential escape character to the adjusted word frequency number of words in the current file to obtain a comma-separated value file with the English watermark added. Specifically, it includes: Obtain the total word frequency of the words in the current file that do not include the content included in the preset essential escape character corresponding to the watermark character; according to the word frequency of the adjusted watermark character, add double quotes to the adjusted word frequency number of words from the total word frequency number of words to obtain a comma-separated value file with the English watermark added.

[0012] In a second aspect, the present application provides a watermark processing system for a comma-separated value file. The system includes: A watermark adding component, including a clearing module, an obtaining module, a determining module, a watermark frequency module, an adjusting module, and an obtaining module; Among them, the clearing module is used to traverse the current file and clear the preset non-essential escape characters in the file when performing the watermark adding operation; The obtaining module is used to count the word frequency of the words that do not include the content included in the preset essential escape character in the file and obtain the word sorting; The determining module is used to determine the characters corresponding to each word according to the order of the characters in the preset character table based on the word sorting, and further determine the total word frequency corresponding to each character; The watermark frequency module is used to obtain the watermark characters involved in the English watermark and the order of the watermark characters, and determine the total word frequency corresponding to the same characters of the watermark characters in the preset character table as the word frequency of the watermark characters; The adjusting module is used to adjust the word frequency of the watermark characters according to the order of the watermark characters so that the word frequency of the watermark characters is arranged from large to small according to the order of the watermark characters; The obtaining module is used to add a preset non-essential escape character to the adjusted word frequency number of words in the current file according to the adjusted word frequency of the watermark characters and the words corresponding to the watermark characters to obtain a comma-separated value file with the English watermark added; The watermark extracting component is used to traverse the current file when performing the watermark extracting operation, determine the correspondence between the words and the characters in the preset character table, and then read the word frequency of each word included in the non-essential escape character in the file. Based on the correspondence between the words and the characters, obtain the total word frequency corresponding to the characters, and arrange the characters according to the total word frequency size to obtain the watermark.

[0013] In one implementation of the present application, the clearing module includes a clearing unit, which is used to traverse the current file and determine whether the content enclosed by double quotes contains a comma. When it does not contain a comma, the double quotes are cleared.

[0014] In an implementation manner of the present application, the obtaining module includes an obtaining unit, which is used to obtain the total word frequency of words other than the content included in the preset necessary escape characters corresponding to the watermark characters in the current file; according to the word frequency of the adjusted watermark characters, add double quotes to the adjusted word frequency of words from the total word frequency of words to obtain a comma-separated value file with English watermark added.

[0015] In a third aspect, the present application provides a non-volatile computer storage medium, on which computer instructions are stored, and when the computer instructions are executed, they implement a watermark processing method for comma-separated value files as described in any one of the above.

[0016] Those skilled in the art can understand that the present application has at least the following beneficial effects: The present application provides a watermark processing method, system and medium for comma-separated value files, which uses the word frequency statistical data of the entire document to carry the watermark content and has a certain ability to resist modification. Different from traditional watermark addition methods, the present application embeds the watermark by adjusting the word frequency of words and adding non-necessary escape characters, without changing the actual values and structures of the original data, thus maintaining the integrity and availability of the data, and solving the problem that the traditional operation of modifying the original file composed of English and numbers will destroy the availability of the original data and cannot be directly used as input in actual production processes such as machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 FIG. is a flowchart of a watermark processing method for comma-separated value files provided by an embodiment of the present application.

[0019] Figure 2 FIG. is a schematic internal structure diagram of a watermark processing system for comma-separated value files provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, which do not mean that the present disclosure can only be implemented through the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.

[0021] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0022] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0023] The embodiment provides a watermark processing method for a comma-separated value file, as Figure 1 shown, the method provided in the embodiment of the present application mainly includes the following steps: Step 110: When performing the watermark addition operation, traverse the current file and clear the preset unnecessary escape characters in the file.

[0024] In some embodiments, the preset unnecessary escape character is that if a double quote encloses a field that does not contain a comma in any field, then the double quote enclosing the field is considered an unnecessary escape character.

[0025] This step can be specifically: Traverse the current file, determine whether the content enclosed by double quotes contains a comma, and when it does not contain a comma, clear the double quotes.

[0026] Further specifically, assume that the watermark content to be added is com, and assume that a certain csv document is as follows: "my name is zhangsan, i like beer.", "a football football new", a newfootball, "my name is lisi football", "football", a a new football.

[0027] After traversing the entire file and clearing the unnecessary escape characters in the file, the csv document becomes: "my name is zhangsan, i like beer.", a football football new, a newfootball, my name is lisi football, football, a a new football.

[0028] Step 120: Count the word frequencies of words other than the content included in the preset necessary escape characters in the file, and obtain the word ranking.

[0029] It should be noted that the preset necessary escape character can specifically be: Determine whether the content enclosed in double quotes contains a comma. When it contains a comma, the current double quote is the preset necessary escape character.

[0030] Using the file in Step 110, perform word frequency analysis on the entire file to count the word frequencies of each word in the file. This step does not count the content included in the preset necessary escape characters, that is, "my name is zhangsan, i likebeer.".

[0031] The statistical results are as follows: football: 6; a: 4; new: 3; name: 1; my: 1; is: 1; lisi: 1.

[0032] Step 130: Based on the word ranking, determine the characters corresponding to each word according to the order of the characters in the preset character table, and then determine the total word frequency corresponding to each character.

[0033] As an example, this step can specifically be: Obtain the preset character table, and then obtain the order of the characters in the preset character table. According to the order of the characters, poll to determine the characters corresponding to each word in the word ranking; According to the word frequencies of all the words corresponding to the current character, sum them up to obtain the total word frequency corresponding to the current character.

[0034] It should be further noted that the preset character table can be an alphabet composed of any 26 letters, or an ordered sequence of characters composed of any letters, as long as it can provide a mapping relationship with the words. The preset character table needs to contain watermark characters.

[0035] Step 140: Obtain the watermark characters and the order of the watermark characters involved in the English watermark, and determine that the total word frequency corresponding to the same characters of the watermark characters in the preset character table is the word frequency of the watermark characters.

[0036] Step 150: Adjust the word frequencies of the watermark characters according to the order of the watermark characters, so that the word frequencies of the watermark characters are arranged from largest to smallest according to the order of the watermark characters.

[0037] This step can specifically be: Obtain the word frequency corresponding to the next watermark character according to the watermark character order; when the word frequency corresponding to the next watermark character is greater than or equal to the current watermark character, modify the word frequency corresponding to the next watermark character to be less than the word frequency of the current watermark character.

[0038] More specifically, this step uses the simplest a - z method for correspondence, and the word frequencies of the original watermark characters are as follows: a -- append 400 times; b -- iphone 300 times; c -- add 280 times; d -- attach 250 times; e -- touch 200 times; f -- batch 190 times; g -- branch 150 times; h -- banana 110 times; i -- apple 100 times √; j -- brand 90 times; k -- both 85 times; l -- brother 80 times; m -- other 70 times; n -- egg 60 times √; o -- fire 45 times; p -- born 30 times √; q -- bear 25 times; r -- man 20 times √; s -- body 10 times √; t -- baby 7 times; u -- girl 3 times √.

[0039] When the watermark content is inspur, according to the order of the letters in inspur, the frequencies of p and r are restricted to lower levels.

[0040] i apple 100 times; n egg 60 times; p born 30 times (before adjustment); r man 20 times (before adjustment); s body 10 times; u girl 3 times.

[0041] It should be further noted that according to the rule that when the frequency of the next watermark character is greater than or equal to the current watermark character, the frequency of the next watermark character is modified to be less than the frequency of the current watermark character, it can be modified as follows: p born 7 times (after adjustment), r man 1 time (after adjustment). At this time, p (7 times) is less than 10 times of the previous s, and r (1 time) is less than 3 times of the previous u. In addition, when the frequency of the next watermark character is less than the current watermark character, the number of times remains unchanged.

[0042] Step 160: According to the adjusted frequency of the watermark character and the word corresponding to the watermark character, add a preset non-essential escape character to the adjusted frequency of words in the current file to obtain a comma-separated value file with the English watermark added.

[0043] This step can be specifically as follows: Obtain the total frequency of words in the current file that are outside the content included in the preset essential escape character corresponding to the watermark character; according to the adjusted frequency of the watermark character, add double quotes to the adjusted frequency of words from the total frequency of words to obtain a comma-separated value file with the English watermark added.

[0044] Further specifically, the specific implementation process of "adding double quotes to the adjusted frequency of words from the total frequency of words" can be exemplified as follows: When p is the watermark character, born (assuming the total frequency is 30 times) is the corresponding word, and the adjusted frequency of the watermark character is 7 times, select 7 borns from 30 borns in the full text and add double quotes. The rule for selecting 7 borns can be to add 7 double quotes in the order of appearance in the article, or to randomly select 7 from 30 borns and add double quotes.

[0045] In addition, it should be further noted that when there are multiple words corresponding to one watermark character (for example, no and nobody), and the total frequency (the sum of the frequencies of no and nobody) is 100, and the adjusted frequency of the watermark character is 90, double quotes can be added to no and nobody in any feasible way as long as the total number of double quotes for no + nobody is equal to 90.

[0046] Step 170: When performing the watermark extraction operation, traverse the current file to determine the correspondence between words and characters in the preset character table, and then read the frequency of each word included in the non-essential escape character in the file. Based on the correspondence between words and characters, obtain the total frequency corresponding to the characters, and arrange the characters in descending order of the total frequency to obtain the watermark.

[0047] It should be noted that the determination of the correspondence between words and characters in the preset character table is the same as the above steps of "traversing the current file to clear the preset unnecessary escape characters in the file; counting the word frequencies of words other than the content included in the preset necessary escape characters in the file to obtain the word sorting; and determining the corresponding characters for each word according to the order of characters in the preset character table", and this application will not elaborate here.

[0048] Based on the description above, this application can be specifically as follows: Assume that the watermark character to be added is off. Assume that the preset character table adopted is afffbcdon. Adding watermark: First step, traverse the entire file to clear the unnecessary escape characters in the file.

[0049] Assume a certain csv file is as follows: i2,i2,"i3","i3","i3",i4,i4,i4; i4,i5,i5,i5,i5,i5,i6,i6; i6,i6,i6,i6,i7,i7,i7,i7; i8,i8,i8,i9,i9,i9,i9,i9; i9,i9,i9,i9,i10,i10,i10,i10; i10,i10,i10,i10,i10,i10,i11,i11; i11,i11,i11,i11,i11,i11,i11,i11; i11,i12,i12,i12,i12,i12,i12,i12; i12,i12,i12,i12,i12,i13,i13,i13; i13,i13,i13,i13,i13,i13,i13,i13; i13,i13,i14,i14,i14,i14,i14,i14; i14,i14,i14,i14,i14,i14,i14,i14; i15,i15,i15,i15,i15,i15,i15,i15; i15,i15,i15,i15,i15,i15,i15,i16; i16,i16,i16,i16,i16,i16,i16,i16; i16, i16, i16, i16, i16, i16, i16, i17; i17, i17, i17, i17, i17, i17, i17, i17; i17, i17, i17, i17, i17, i17, i17, i17; i18, i18, i18, i18, i18, i18, i18, i18; i18, i18, i18, i18, i18, i18, i18, i18; i18, i18, i19, i19, i19, i19, i19, i19; i19, i19, i19, i19, i19, i19, i19, i19; i19, i19, i19, i19, i19.

[0050] It should be noted that for the convenience of reading, ";" and "." are added to the above csv text, and the original file does not contain ";" and ".".

[0051] Clear the preset unnecessary escape characters in the file, count the word frequencies of words other than the content included in the preset necessary escape characters in the file, obtain the word sorting, and the mapping with the preset character table (afffbcdon) is as follows: a - i19, i10, a total of 29 times; f (the first f) - i18, i9, a total of 27 times; f (the second f) - i17, i8, a total of 25 times; f (the third f) - i16, i7, a total of 23 times; b - i15, i6, a total of 21 times; c - i14, i5, a total of 19 times; d - i13, i4, a total of 17 times; o - i12, i3, a total of 15 times; n - i11, i2, a total of 13 times.

[0052] Based on the word sorting, in accordance with the order of the characters in the preset character table, determine the characters corresponding to each word, and then determine the total word frequency corresponding to each character, obtain the watermark characters and watermark character order involved in the English watermark, and determine that the total word frequency corresponding to the same character in the preset character table for the watermark character is the word frequency of the watermark character.

[0053] A feasible number of times each letter is marked is: o (mapped with i12, i3) 15 times; f (the first f, the one mapped with i18, i9) 13 times; f (the second f, the one mapped with i17, i8) 10 times.

[0054] It should be supplemented and explained that when the word frequency corresponding to the next watermark character is greater than or equal to the current watermark character, the word frequency corresponding to the next watermark character is modified to be less than the word frequency of the current watermark character. According to this rule, when the word frequency corresponding to the next watermark character is already less than the current watermark character, the number of times remains unchanged.

[0055] Based on the adjusted word frequencies of the watermark characters and the words corresponding to the watermark characters, add a preset non-essential escape character to the adjusted word frequency number of words in the current file to obtain a comma-separated value file with the English watermark added. The document becomes: i2, i2, "i3", "i3", "i3", i4, i4, i4; i4, i5, i5, i5, i5, i5, i6, i6; i6, i6, i6, i6, i7, i7, i7, i7; i7, i7, i7, i8, i8, i8, i8, i8; i8, i8, i8, i9, i9, i9, i9, i9; i9, i9, i9, i9, i10, i10, i10, i10; i10, i10, i10, i10, i10, i10, "i11", "i11"; "i11", "i11", "i11", i11, i11, i11, i11, i11; i11, "i12", "i12", "i12", "i12", "i12", "i12", "i12"; "i12", "i12", "i12", "i12", "i12", i13, i13, i13; i13, i13, i13, i13, i13, i13, i13, i13; i13, i13, i14, i14, i14, i14, i14, i14; i14, i14, i14, i14, i14, i14, i14, i14; i15, i15, i15, i15, i15, i15, i15, i15; i15, i15, i15, i15, i15, i15, i15, i16; i16, i16, i16, i16, i16, i16, i16, i16; i16, i16, i16, i16, i16, i16, i16, "i17"; "i17", "i17", "i17", "i17", "i17", "i17", "i17", "i17"; "i17", i17, i17, i17, i17, i17, i17, i17; "i18", "i18", "i18", "i18", "i18", "i18", "i18", "i18"; "i18", "i18", "i18", "i18", "i18", i18, i18, i18; i18, i18, i19, i19, i19, i19, i19, i19; i19, i19, i19, i19, i19, i19, i19, i19; i19, i19, i19, i19, i19。

[0056] It should be noted that for the convenience of reading, ";" and "." are added to the above csv text, and ";" and "." do not exist in the content of the file itself.

[0057] When performing the watermark extraction operation, the specific correspondence between words and characters in the preset character table can be as follows: Perform a word frequency analysis on the entire file to count the word frequencies of each word in the file: i19 appears 19 times; i18 appears 18 times; i17 appears 17 times; i16 appears 16 times; i15 appears 15 times; i14 appears 14 times; i13 appears 13 times; i12 appears 12 times; i11 appears 11 times; i10 appears 10 times; i9 appears 9 times; i8 appears 8 times; i7 appears 7 times; i6 appears 6 times; i5 appears 5 times; i4 appears 4 times; i3 appears 3 times; i2 appears 2 times.

[0058] Map each word that appears in the document to a preset character table according to the established rules (afffbcdon, corresponding from a to n, polling from a to n, and reaching the last word), so that a certain word in the document represents a certain letter (character) in the watermark.

[0059] Since the preset character table is afffbcdon, according to the established mapping rules, the mapping results are as follows: a - i19, i10; f - i18, i9; f - i17, i8; f - i16, i7; b - i15, i6; c - i14, i5; d - i13, i4; o - i12, i3; n - i11, i2.

[0060] Read the word frequencies of each word contained in the non-essential escape characters in the file, specifically: i19 0 times; i18 18 times; i17 17 times; i16 0 times; i15 0 times; i14 0 times; i13 0 times; i12 12 times; i11 0 times; i10 0 times; i9 9 times; i8 8 times; i7 0 times; i6 0 times; i5 0 times; i4 0 times; i3 3 times; i2 0 times.

[0061] Based on the correspondence between words and characters (for example, i12 corresponds to o, i3 corresponds to o), obtain the total word frequency corresponding to the character (the total word frequency of o = the word frequency of i12 + the word frequency of i3 = 12 + 3 = 15), and arrange the corresponding characters in descending order of the total word frequency corresponding to the character to obtain the watermark, specifically as follows.

[0062] Arrange the characters in descending order of the total word frequency as follows: a - i19, i10 - 0 times; f - i18,i9 - 13 times; f - i17,i8 - 10 times; f - i16,i7 - 0 times; b - i15,i6 - 0 times; c - i14,i5 - 0 times; d - i13,i4 - 0 times; o - i12,i3 - 15 times; n - i11,i2 - 0 times; Arrange o - i12,i3 - 15 times, f - i18,i9 - 13 times, f - i17,i8 - 10 times according to the word frequency size, and obtain the watermark string off.

[0063] In addition, this application Figure 2 is a watermark processing system for comma - separated value files provided by an embodiment of this application. As Figure 2 shown, the system provided by the embodiment of this application mainly includes: The watermark - adding component 200 includes a clearing module 210, an obtaining module 220, a determining module 230, a watermark frequency module 240, an adjusting module 250, and an obtaining module 260.

[0064] Among them, the clearing module 210 is used to traverse the current file and clear the preset unnecessary escape characters in the file when performing the watermark - adding operation.

[0065] The clearing module 210 includes a clearing unit for traversing the current file to determine whether the content enclosed by double quotes contains a comma. When it does not contain a comma, the double quotes are cleared.

[0066] The obtaining module 220 is used to count the word frequencies of the words outside the content included in the preset necessary escape characters in the file and obtain the word sorting.

[0067] The obtaining module 220 includes an obtaining unit for obtaining the total word frequency of the words outside the content included in the preset necessary escape characters corresponding to the watermark characters in the current file; according to the adjusted word frequency of the watermark characters, adding double quotes to the adjusted word frequency of words from the total word frequency of words to obtain a comma - separated value file with English watermarks added.

[0068] The determining module 230 is used to determine the characters corresponding to each word based on the word sorting in the order of the characters in the preset character table, and further determine the total word frequency corresponding to each character.

[0069] The watermark frequency module 240 is used to obtain the watermark characters and the order of watermark characters involved in the English watermark, and determine that the total word frequency corresponding to the same characters of the watermark characters in the preset character table is the word frequency of the watermark characters.

[0070] The adjustment module 250 is used to adjust the word frequency of the watermark characters according to the order of the watermark characters, so that the word frequency of the watermark characters is arranged from large to small according to the order of the watermark characters.

[0071] The obtaining module 260 is used to add a preset non-essential escape character to the adjusted word frequency number of words in the current file according to the adjusted word frequency of the watermark characters and the words corresponding to the watermark characters, so as to obtain a comma-separated value file with the English watermark added.

[0072] The watermark extraction component 300 is used to traverse the current file when performing the watermark extraction operation, determine the correspondence between the words and the characters in the preset character table, and then read the word frequency of each word included in the non-essential escape character in the file. Based on the correspondence between the words and the characters, obtain the total word frequency corresponding to the characters, arrange the characters in descending order of the total word frequency, and obtain the watermark.

[0073] In addition, the embodiments of the present application also provide a non-volatile computer storage medium, on which executable instructions are stored. When the executable instructions are executed, a watermark processing method for a comma-separated value file as described above is implemented.

[0074] So far, the technical solutions of the present disclosure have been described in combination with multiple foregoing embodiments. However, it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principle of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.

Claims

1. A watermark processing method for comma separated value files, characterized in that: The method comprises: When adding a watermark, traverse the current file and clear the preset unnecessary escape characters in the file; The frequency of words other than those in the content with the preset necessary escape characters in the statistics file is calculated to obtain the word ranking; Based on the word sorting, the characters corresponding to each word are determined according to the order of the characters in the preset character table, and then the total word frequency corresponding to each character is determined; Acquire watermark characters and watermark character sequences involved in an English watermark, and determine that the total word frequency corresponding to the same character of the watermark character in a preset character table is the word frequency of the watermark character; wherein the preset character table includes the watermark character; According to the sequence of watermark characters, the word frequencies of the watermark characters are adjusted so that the word frequencies of the watermark characters are arranged from large to small according to the sequence of the watermark characters; According to the adjusted word frequency of the watermark character and the word corresponding to the watermark character, a preset non-essential escape character is added to the adjusted word frequency words in the current file to obtain a comma-separated value file with the English watermark added; When extracting watermarks, the current file is traversed to determine the correspondence between words and characters in the preset character table, and then the word frequency of each word contained in the non-essential escape characters in the file is read. Based on the correspondence between words and characters, the total word frequency corresponding to the characters is obtained, and the characters are arranged according to the total word frequency to obtain the watermark.

2. The watermark processing method for comma separated value files according to claim 1, characterized in that: The default non-essential escape character is that when any field does not contain a comma but is enclosed in double quotes, the double quotes enclosing the field are considered non-essential escape characters.

3. The watermark processing method for comma separated value files according to claim 2, characterized in that: When adding a watermark, traverse the current file and clear the preset unnecessary escape characters in the file, including: Traverse the current file and determine whether the content enclosed by double quotes contains commas. If it does not contain commas, clear the double quotes.

4. The watermark processing method for comma separated value files according to claim 1, characterized in that: Based on the word sorting, the characters corresponding to each word are determined according to the order of the characters in the preset character table, and then the total word frequency corresponding to each character is determined, specifically including: Obtain a preset character table, and then obtain the order of characters in the preset character table, and according to the order of characters, poll to determine the characters corresponding to each word in the word sorting; According to the word frequency of all words corresponding to the current character, add up and get the total word frequency corresponding to the current character.

5. The watermark processing method for comma separated value files according to claim 1, characterized in that: According to the sequence of watermark characters, the word frequencies of the watermark characters are adjusted so that the word frequencies of the watermark characters are arranged from large to small according to the sequence of the watermark characters, specifically including: According to the sequence of watermark characters, the word frequency corresponding to the next watermark character is obtained; when the word frequency corresponding to the next watermark character is greater than or equal to the current watermark character, the word frequency corresponding to the next watermark character is modified to be less than the word frequency of the current watermark character.

6. The watermark processing method for comma separated value files according to claim 1, characterized in that: According to the adjusted word frequency of the watermark character and the word corresponding to the watermark character, a preset non-essential escape character is added to the adjusted word frequency word in the current file to obtain a comma-separated value file with the English watermark added, specifically including: The total frequency of words other than the preset necessary escape characters corresponding to the watermark character in the current file is obtained; according to the frequency of the adjusted watermark character, double quotation marks are added to the adjusted frequency words from the total frequency words to obtain a comma-separated value file with the English watermark added.

7. A watermark processing system for comma separated value files, characterized in that: The system comprises: Add watermark components, including clearing module, obtaining module, determining module, watermark frequency module, adjusting module, and obtaining module; The clearing module is used to traverse the current file and clear the preset unnecessary escape characters in the file when adding a watermark. An acquisition module is used to count the frequency of words in the file that do not contain the preset necessary escape characters, and obtain the word ranking; A determination module, for determining the characters corresponding to each word based on the word sorting and in the order of the characters in the preset character table, and then determining the total word frequency corresponding to each character; The watermark frequency module is used to obtain the watermark characters and the watermark character sequence involved in the English watermark, and determine the total word frequency corresponding to the same character of the watermark character in the preset character table as the word frequency of the watermark character; An adjustment module, used for adjusting the word frequencies of the watermark characters according to the watermark character sequence, so that the word frequencies of the watermark characters are arranged from large to small according to the watermark character sequence; An acquisition module is used to add a preset non-essential escape character to the adjusted word frequency words in the current file according to the adjusted word frequency of the watermark character and the word corresponding to the watermark character, so as to obtain a comma-separated value file with the English watermark added; The watermark extraction component is used to traverse the current file when performing the watermark extraction operation, determine the correspondence between words and characters in the preset character table, and then read the word frequency of each word contained in the non-essential escape characters in the file, based on the correspondence between words and characters, obtain the total word frequency corresponding to the characters, arrange the characters according to the total word frequency, and obtain the watermark.

8. The watermark processing system for comma separated value files according to claim 7, characterized in that: The clearing module includes a clearing unit, which is used for traversing the current file, determining whether the content enclosed by double quotes contains commas, and clearing the double quotes when the content does not contain commas.

9. The watermark processing system for comma separated value files according to claim 7, characterized in that: The acquisition module includes an acquisition unit, which is used to obtain the total word frequency of words outside the preset necessary escape characters corresponding to the watermark character in the current file; according to the word frequency of the adjusted watermark character, double quotation marks are added to the adjusted word frequency words from the total word frequency words to obtain a comma-separated value file with the English watermark added.

10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, the watermark processing method for comma separated value files according to any one of claims 1 to 6 is implemented.