A method and system for embedding and extracting data fingerprints

The convolutional neural network processes hashing, reordering and grouping of table data, and the robustness and resistance of fingerprint embedding are improved without affecting data statistics, and the problem of fingerprint easy to miss or distortion in the existing technology is solved.

CN115510053BActive Publication Date: 2025-07-11CHONGQING SHUDA INFORMATION SECURITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210889591.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-07-11
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

While reducing the impact on data, existing data fingerprint embedding technologies are prone to missing fingerprint data or causing data distortion, making it difficult to maintain resistance to attacks when facing attacks.

Method used

The convolutional neural network encoder and decoder are used to hash calculation, reorder and group table data, and fingerprint embedding and extraction are used to ensure that the fingerprint coverage is wide and does not affect the statistical properties of the data.

Benefits of technology

It improves the robustness and aggressiveness of fingerprints, ensuring that fingerprints can still be effectively extracted after partially deleting or modifying the data, without affecting the statistical usability of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510053B_ABST
    Figure CN115510053B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for embedding and extracting data fingerprints, belonging to the technical field of data security. For data fingerprint embedding, the data table is first split and grouped, then for each group of data, a convolutional neural network encoder and fingerprint data are used for fingerprint embedding, and finally the grouped data embedded with fingerprints are combined in the original grouping order to obtain the tabular data embedded with fingerprint data. For data fingerprint extraction, grouping is performed in the same way as during embedding, then for each group of data, a convolutional neural network decoder is used to remove the fingerprint, and then the difference between the data before and after fingerprint removal for each group is averaged to obtain the fingerprint data. Compared with the prior art, the present invention embeds fingerprints in the entire domain data, has stronger anti-attack ability, and improves the robustness of the embedded fingerprints; starting from the global statistics of the data, encoding the fingerprints according to the overall characteristics of the data to ensure that the embedding of fingerprints does not affect the usability related to data statistics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and system for embedding and extracting data fingerprints, belonging to the technical field of data security. Background Art

[0002] A data fingerprint refers to embedding corresponding information in data (generally having a row-column structure). This information can be used as a judgment identifier for personnel related to the data such as data users and owners, and can also be used as a basis for determining whether the data has been tampered with. In case of disputes over the ownership of data, it can be used as a criterion for determining copyright ownership. Data fingerprints are widely used, including audio data fingerprints, video data fingerprints, image data fingerprints, etc. proposed for various data. Data fingerprints can embed information of relevant participants such as handlers and participants into the data during the process of data circulation and release. In case of a security incident, the entire leakage process can be quickly traced based on the leaked data, and it can also play a certain role in restricting the security issues of data users during data use. In case of disputes over data, data fingerprints can also play a certain role in determining copyright ownership.

[0003] Data fingerprint technology can currently be classified into two categories according to the degree of modification of data: one is distorted data fingerprints, and the other is undistorted data fingerprints. Undistorted data fingerprints are based on relevant attributes of the data such as the hash value and data arrangement characteristics of the original data without modifying the original data as the basis for fingerprint calculation. The advantage of this method is that it does not modify the data at all and will not have any impact on the data, but the anti-attack ability of this method is relatively poor. For example, the fingerprint method based on the hash value will completely fail when facing data tampering attacks. Distorted data fingerprints modify the original data to embed fingerprint data into the data table. This method will modify the data, resulting in a certain degree of difference between the data after embedding the fingerprint and the original data, and the data will be affected during use. Compared with the undistorted method, this method embeds the fingerprint into the data by modifying the data, and the fingerprint data will intersect with the original data. When facing some common attack means such as deletion attacks and modification attacks, although the data is damaged, there will still be fingerprint residues. Therefore, the distorted method has a stronger anti-attack ability. In the research of data fingerprints, distortion and anti-attack ability are inherently contradictory. Generally speaking, the higher the distortion degree, the stronger the anti-attack ability. Currently, the mainstream data fingerprint embedding practices are all proposed based on distorted data fingerprints. However, in order to reduce the impact on the data after embedding the fingerprint, existing methods modify as few data units as possible. But a small number of modifications are extremely easy to be missed during fingerprint extraction, resulting in incomplete fingerprint data and being difficult to distinguish. However, if most of the original data is modified, it will lead to serious data distortion and affect the use. Summary of the Invention

[0004] The object of the present disclosure is to propose a method and a software system for embedding and extracting data fingerprints for the above-mentioned part or all of the problems.

[0005] The object of the present disclosure is achieved by the following technical solutions:

[0006] In a first aspect, an embodiment of the present disclosure provides a method for embedding fingerprints of tabular data, including the following:

[0007] 1) Perform a hash calculation on the unique identifier K of the tabular data T to obtain the hash value HK of each row of K;

[0008] 2) Reorder the rows of the T in ascending or descending order according to the HK;

[0009] 3) Group the numerical data of the T to obtain n groups of data, denoted as T = {T1, T2,...T i ...T n}, where each group contains k rows, and the number of numerical data columns of the T except the unique identifier is k;

[0010] 4) Respectively use a trained convolutional neural network encoder and fingerprint data to perform fingerprint embedding on each group of data of the T to obtain E1, E2,...E i ...E n ;

[0011] 5) Combine E1, E2,...E i ...E n in the grouping order to obtain the tabular data T', and T' is the data table after fingerprint embedding.

[0012] In a second aspect, an embodiment of the present disclosure provides a method for extracting fingerprints of tabular data, including:

[0013] 1) Perform a hash calculation on the unique identifier K of the tabular data T' to obtain the hash value HK of each row of K;

[0014] 2) Reorder the rows of the T' in ascending or descending order according to the HK;

[0015] 3) Group the numerical data of the T' to obtain n groups of data, denoted as T' = {E1, E2,...E i ...E n}, where each group contains k rows, and the number of numerical data columns of the T' except the unique identifier is k;

[0016] 4) Respectively use a trained convolutional neural network decoder to remove fingerprints from each group of data of the T' to obtain {S1, S2,...S i ...Sn};

[0017] 5) Calculate the fingerprint data DM through the following formula:

[0018]

[0019] In a third aspect, an embodiment of the present disclosure provides a fingerprint embedding system for tabular data, including:

[0020] A hash value calculation module, configured to perform a hash calculation on the unique identifier K of the tabular data T to obtain the hash value HK of each row of K;

[0021] A reordering module, configured to reorder the rows of the T in ascending or descending order according to the HK;

[0022] A grouping module, configured to group the numerical data of the T to obtain n groups of data, denoted as T = {T1, T2,... T i ... T n}, where each group contains k rows, and the number of columns of the numerical data of the T except the unique identifier is k;

[0023] A fingerprint embedding module, configured to perform fingerprint embedding on each group of data of the T using a trained convolutional neural network encoder and fingerprint data to obtain E1, E2,... E i ... E n ;

[0024] A table integration module, configured to combine E1, E2,... E i ... E n in the grouping order to obtain the tabular data T', and T' is the data table after fingerprint embedding.

[0025] In a fourth aspect, an embodiment of the present disclosure provides a fingerprint extraction system for tabular data, including:

[0026] A hash value calculation module, configured to perform a hash calculation on the unique identifier K of the tabular data T' to obtain the hash value HK of each row of K;

[0027] A reordering module, configured to reorder the rows of the T' in ascending or descending order according to the HK;

[0028] A grouping module, configured to group the numerical data of the T' to obtain n groups of data, denoted as T' = {E1, E2,... E i ... E n}, where each group contains k rows, and the number of columns of the numerical data of the T' except the unique identifier is k;

[0029] A fingerprint removal module, which is used to separately remove fingerprints from each group of data of the T by using a trained convolutional neural network decoder to obtain {S1, S2,... S i ... S n};

[0030] A fingerprint extraction module, which is used to calculate fingerprint data DM by the following formula:

[0031]

[0032] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor; and

[0033] a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the foregoing first aspect or any implementation manner of the first aspect or the second aspect or any implementation manner of the second aspect.

[0034] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect or any implementation manner of the first aspect or the second aspect or any implementation manner of the second aspect.

[0035] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the method described in the foregoing first aspect or any implementation manner of the first aspect or the second aspect or any implementation manner of the second aspect.

[0036] Advantageous effects

[0037] Compared with the prior art, the present disclosure has the following effects:

[0038] (1) Fingerprints are embedded in the global data, and the fingerprint coverage is wide. When the data is partially deleted or modified, the fingerprint data still exists, which has stronger anti-attack ability and improves the robustness of the embedded fingerprints;

[0039] (2) Starting from the global statistics of the data, encoding the fingerprints according to the overall characteristics of the data to ensure that the embedding of the fingerprints does not affect the usability related to data statistics. Description of the drawings

[0040] Figure 1It is a schematic flowchart of a fingerprint embedding method for tabular data provided by an embodiment of the present disclosure;

[0041] Figure 2 It is a schematic diagram of the convolutional neural network encoder structure provided by an embodiment of the present disclosure;

[0042] Figure 3 It is a schematic flowchart of a fingerprint extraction method for tabular data provided by an embodiment of the present disclosure;

[0043] Figure 4 It is a schematic diagram of the convolutional neural network decoder structure provided by an embodiment of the present disclosure;

[0044] Figure 5 It is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;

[0045] Figure 6 It is a schematic diagram of the performance of various fingerprint embedding methods during deletion attacks. Detailed implementation manners

[0046] The following further illustrates and describes the present disclosure in conjunction with the accompanying drawings and embodiments.

[0047] The following illustrates the implementation manners of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0048] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or this method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.

[0049] It should also be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0050] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0051] With the development of information technology, the emergence and widespread use of a large number of information systems have made tabular data ubiquitous. All walks of life have accumulated a large amount of tabular data, which is crucial for technologies such as artificial intelligence, big data, and data mining. Through such technologies, the analysis of a large amount of tabular data can enable people to understand the laws behind the data, thus better guiding people's life practices. The rise of these data analysis and utilization technologies inevitably brings data circulation, and in the process of data circulation, how to ensure that the data is not misused and the data can be traced requires effective fingerprint embedding and extraction means. Existing commonly used fingerprint embedding technologies are all proposed based on distorted data fingerprints. The common point of these methods is that in a data table, specific data units are selected, and in the selected data units, corresponding modifications are made according to a pre-agreed method. For example, in a data table, according to a specific rule, a certain proportion of rows to be embedded with fingerprints are selected from it. Generally, about 20% of the row data is selected as the standard for fingerprint embedding. In the selected row data, according to a specific rule, one or several data units are selected from the data row, and fixed binary bits are embedded in the selected data units, modified into fixed values, specific numbers are inserted after the decimal point, etc. The selected rows, selected data units, and the embedded data, etc., are all encoded by the data fingerprint. The data fingerprint embeds the fingerprint into the data based on the characteristics of row selection, data unit selection, and data embedding. When extracting the fingerprint, according to the same method and agreement as when embedding the fingerprint, the data in the data unit is extracted, and based on the completeness of the extracted fingerprint data, it is judged whether the data has been tampered with to achieve purposes such as data circulation traceability. In order to reduce the impact on the data after embedding the fingerprint, these methods modify as few data units as possible. However, a small number of modifications are very likely to be missed during fingerprint extraction, resulting in incomplete fingerprint data and making it difficult to distinguish. However, if most of the original data is modified, it will cause serious data distortion and affect the use. To address this problem, the present disclosure provides a method for fingerprint embedding and extraction of tabular data.

[0052] A method for fingerprint embedding and extraction of tabular data provided by an embodiment of the present disclosure can be executed by a computing device, which can be implemented as software, or as a combination of software and hardware. The computing device can be integrally provided in a server, a terminal device, etc.

[0053] The terminal in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, file servers, database servers, etc.

[0054] Figure 1 A method for fingerprint embedding of tabular data provided by the present disclosure, as shown in the figure, the method includes the following steps:

[0055] 1) Perform a hash calculation on the unique identifier K of the tabular data T to obtain the hash value HK of each row of K;

[0056] Preferably, the unique identifier can be a primary key or a composite primary key.

[0057] A row of data in the table can be determined by the unique identifier. A primary key consisting of one column or a composite primary key consisting of multiple columns can be used as the unique identifier to uniquely determine a row of data. In practice, it can be selected according to the specific tabular data. For the selected unique identifier, perform a hash calculation on each row of data to obtain the hash value HK of each row of K. As shown in Table 1, it is a table, and the ID column is the primary key of this table. The present disclosure selects ID as the unique identifier. As shown in Table 2, it is the result obtained by performing a hash calculation on the numerical values in the ID column of each row. Except for the numerical values in the ID column, the other columns remain unchanged.

[0058] Table 1

[0059] ID name money ...... ...... ...... 312022001 jack 3242 ...... ...... ...... 312022002 tom 34543 ...... ...... ...... 312022003 K.J 3456 ...... ...... ...... 312022004 Jonh 456456 ...... ...... ...... 312022005 merry 3452 ...... ...... ...... 312022006 So. 123 ...... ...... ...... 312022007 Jmmay 978 ...... ...... ...... 312022008 Eit 456 ...... ...... ...... 312022009 KK 4567 ...... ...... ......

[0060] Table 2

[0061]

[0062] 2) Reorder the rows of T in ascending or descending order according to HK;

[0063] As shown in Table 3, it is the table result after reordering Table 2.

[0064] Table 3

[0065]

[0066] 3) Group the numerical data of T to obtain n groups of data, denoted as T = {T1, T2,... Ti ...T n},where each group contains k rows, and the number of numerical data columns of T except the unique identifier is k;

[0067] Reordering and grouping based on the unique identifier in a certain fixed manner can, to a certain extent, prevent deletion and insertion attacks on the data and provide the robustness of the method of the present disclosure.

[0068] When grouping, k rows of data are extracted as a group based on the number of columns k except the unique identifier column until the number of remaining rows in the table is less than a group. Since statistical operations such as average and variance are mainly based on numerical data, the present disclosure mainly embeds fingerprints for numerical data, so that although the numerical data in the table is changed due to watermark embedding, it will not affect the use of such data in statistical operations.

[0069] 4) Use the trained convolutional neural network encoder and fingerprint data for each group of data of T to perform fingerprint embedding to obtain {E1, E2,... E i ...E n};

[0070] Each operator (such as data publisher, handler, etc.) of the table data has its own unique fingerprint data, and embeds its fingerprint data into the table data group by group using the encoder. The encoder can use any coding algorithm that does not affect the statistical characteristics of the table data, such as MLP multi-layer perceptron, PCL algorithm, DNN (deep neural network), etc. A convolutional neural network can also be used.

[0071] As a specific example, the present disclosure uses a convolutional neural network encoder, and the network structure of the encoder model is as Figure 2 shown, including the following four layers:

[0072] The first layer is a neural network with 32 3x3 convolutional kernels, expressed as:

[0073] a = Conv 1→32 (C)

[0074] where Conv represents a convolutional neural network, 1→32 means: the input C is changed from the initial 1 layer to 32 layers after convolution;

[0075] This layer uses 32 convolutional kernels to expand the features of the original data into 32 layers, so that the subsequent process can make full use of the features of the data at different levels.

[0076] The second layer is a neural network with 32 3x3 convolutional kernels, expressed as:

[0077] b = Conv 33→32(Cat(a, M))

[0078] Among them, M is fingerprint data; cat() is data concatenation operation. Here: a is the output of the first-layer convolutional neural network, with 32 layers, M is fingerprint data. After performing the Cat operation on a and M, a 33-layer data is obtained. It can be understood that a is data with 32 sheets of paper, and M is data with one sheet of paper. M is placed at the bottom of a;

[0079] This layer fuses the fingerprint data M with the data a after feature expansion, so that the fingerprint data and the original data have consistent statistical characteristics in the 32-layer features. M is a square matrix. When its dimension is inconsistent with that of a, cat() will perform an equi-ratio scaling on M according to the dimension of a.

[0080] The third layer is a neural network with 32 convolutional kernels of size 3x3, expressed as:

[0081] E b = Flatten 32→1 (Conv 32→32 (b))

[0082] Among them, the flatten operation compresses the 32-layer expanded features into one layer;

[0083] This layer further extracts features from the output b of the second-layer network embedded with fingerprint data using 32 convolutional kernels, and then compresses the 32-layer features into one layer through the Flatten operation.

[0084] The fourth layer embeds fingerprints through the following expression:

[0085] E = C + E b

[0086] Among them, E is the fingerprint data after embedding the fingerprint M in C.

[0087] In the above encoding (fingerprint embedding) process, the first layer is an encoding of the original data C. When the second layer encodes the fingerprint data, it takes into account the output of the first layer, that is, the basis for encoding M refers to the original data C. Therefore, the output b of the second layer is an output after weighing the features of C and M. The third layer encodes b again and then flattens it to obtain E b , it can be considered that E b takes into account the overall and local features of the original data C. Therefore, the fourth layer directly adds E b to the original data C. At this time, it may have an impact on individual data in C, but for global statistical features such as calculating the mean and variance of the whole, there is almost no impact.

[0088] 6) Combine E1, E2,... Ei ...E n Combine them in the grouping order to obtain the tabular data T’, which is the data table after embedding the fingerprint.

[0089] For the grouped E1, E2,...E of the fingerprint-embedded data i ...E n The data table T’ obtained by successively splicing and combining according to the grouping order of the original table data T in step 3) is the data table embedded with fingerprint data. This table has the same structure as the original table T, and the content of the numerical column is slightly different, but the statistical properties are consistent with those of the numerical column of the original table T.

[0090] Figure 2 A fingerprint extraction method for tabular data provided by the present disclosure is shown in the figure. The method includes the following steps:

[0091] 1) Perform a hash calculation on the unique identifier K of the tabular data T’ to obtain the hash value HK of each row of K;

[0092] 2) Reorder the rows of T’ in ascending or descending order according to HK;

[0093] 3) Group the numerical data of T’ to obtain n groups of data, denoted as T’ = {E1, E2,...E i ...E n}, where each group contains k rows, and the number of numerical data columns of T’ except the unique identifier is k;

[0094] The above process corresponds to steps 1-3 of the fingerprint embedding process, and the meaning is the same as that, so it will not be elaborated here.

[0095] 4) Use the trained convolutional neural network decoder to remove the fingerprint from each group of data in T’ to obtain {S1, S2,...S i ...S n};

[0096] 5) Calculate the fingerprint data DM through the following formula:

[0097]

[0098] Erase the fingerprint data of each group in the table through the decoder, that is, obtain the original data after decoding. Then, calculate the difference between the fingerprint data and the original data to obtain the fingerprint data, and then average the fingerprint data in each group to obtain the fingerprint data embedded in the table T’. Then, operations such as fingerprint comparison can be performed on the fingerprint data extracted through the above process to confirm the data source or other operations. The decoder can use any decoding algorithm corresponding to the foregoing encoder, such as the regular autoencoder algorithm, the DNN encoder algorithm, the restricted Boltzmann machine algorithm, etc. A convolutional neural network can also be used.

[0099] As a specific example, corresponding to the aforementioned neural network encoder, the present disclosure uses a convolutional neural network decoder, and the network structure of the decoder model is as Figure 4 shown, which is a neural network with 32 convolutional kernels of size 3x3, expressed as:

[0100] S = Flatten 32→1 (Conv 1→32 (E))

[0101] where Conv represents a convolutional neural network, 1→32 means that the input E is changed from the initial 1 layer to 32 layers after convolution; the flatten operation compresses the features of the 32 layers into one layer.

[0102] This decoder uses 32 convolutional kernels to restore the data with fingerprint information, that is, to erase it without any processing on the fingerprint itself; therefore, the expectation of S is the original data without fingerprints. Finally, the fingerprint data embedded in the table T’ is extracted by taking the average of the difference between the fingerprint data and the original data for each group.

[0103] Since the above encoder and decoder models are both neural networks, they need to be trained using training data. To enable the model to fully learn the fingerprint encoding and fingerprint erasure features, in this example, the encoder and decoder models are jointly trained; each table in the training data used has more than 10,000 rows, has a primary key or a composite primary key, the number of numerical columns other than the primary key or composite primary key is not less than 10, and the number of groups is not less than 30; each table is regarded as a batch; a random fingerprint data M is assigned to each table, so that all groups under the same table use the same M for fingerprint embedding; the following loss function L d Calculate the loss of each group T i :

[0104] L d = CrossEntropy(D(E(T i , M)), T i )

[0105] where CrossEntropy() is the cross-entropy loss function, E(T i , M) is the result after the i-th group T i of the table T is embedded with fingerprints by the aforementioned convolutional neural network encoder, and D(E(T i , M)) is E(T i, M) The result after erasing the fingerprint by the aforementioned convolutional neural network decoder. For a batch, after summing up the losses of all groups, backpropagation is performed to update the model parameters of the encoder and the decoder. Using all the tables in the training dataset to train the encoder and decoder models once is considered as one round. The training is repeated for multiple rounds until the loss no longer decreases. At this point, this model (encoder and decoder) can be used for the aforementioned fingerprint embedding and extraction methods.

[0106] In the above fingerprint embedding and extraction method of the present disclosure, different from the existing fingerprint algorithms that focus on the fingerprint itself and work on how to embed and decode the fingerprint, it makes full use of the advantages of feature learning of convolutional neural networks, uses convolutional neural networks to embed fingerprints and restore data. In this way, as long as two data are simply subtracted, the fingerprint can be obtained. Moreover, since the above method embeds fingerprints in the global data, the fingerprint coverage is wide. When the data is partially deleted or modified, the fingerprint data still exists, with strong anti-attack ability and good robustness of the embedded fingerprint. Furthermore, when embedding fingerprints, starting from the global statistics of the data and encoding the fingerprints based on the overall characteristics of the data, although individual data in the table is modified, the embedding of the fingerprint does not affect the usability related to the statistics of the table data.

[0107] A fingerprint embedding system for tabular data provided by the present disclosure includes the following modules:

[0108] A hash value calculation module, configured to perform hash calculation on the unique identifier K of the tabular data T to obtain the hash value HK of each row of K;

[0109] A reordering module, configured to reorder the rows of the T in ascending or descending order according to the HK;

[0110] A grouping module, configured to group the numerical data of the T to obtain n groups of data, denoted as T = {T1, T2,... T i ... T n}, where each group contains k rows, and the number of numerical data columns of the T except the unique identifier is k;

[0111] A fingerprint embedding module, configured to perform fingerprint embedding on each group of data of the T respectively using a trained convolutional neural network encoder and fingerprint data to obtain E1, E2,... E i ... E n ;

[0112] A table integration module, configured to combine E1, E2,... E i ... E n in the grouping order to obtain tabular data T', and T' is the data table after embedding fingerprints.

[0113] A fingerprint extraction system for tabular data provided by the present disclosure includes the following modules:

[0114] A hash value calculation module, configured to perform a hash calculation on the unique identifier K of the tabular data T' to obtain the hash value HK of each row of K;

[0115] A reordering module, configured to reorder the rows of the T' in ascending or descending order according to the HK;

[0116] A grouping module, configured to group the numerical data of the T' to obtain n groups of data, denoted as T' = {E1, E2,... E i ... E n}, where each group contains k rows, and the number of numerical data columns of the T' except the unique identifier is k;

[0117] A fingerprint removal module, configured to use a trained convolutional neural network decoder to remove fingerprints from each group of data of the T to obtain {S1, S2,... S i ... S n};

[0118] A fingerprint extraction module, configured to calculate fingerprint data DM through the following formula:

[0119]

[0120] For the specific implementation methods of each module, refer to the relevant content of the foregoing fingerprint embedding method for tabular data and fingerprint extraction method for tabular data, which will not be elaborated here.

[0121] Experimental results

[0122] To verify the advancement and effectiveness of the method described in the present disclosure, we designed data distortion experiments and anti-attack experiments for verification, as follows:

[0123] 1. Dataset. Select the U.S. Forest Service Resource Information System data (RIS), which is commonly used in this field, as the dataset for the experiment and the comparative experiment. This dataset is a data table with 581,012 rows, and each row contains 54 numerical data.

[0124] 2. Distortion experiment. The distortion experiment is to verify the impact of fingerprints on the original data after the fingerprints are embedded. We use the changes in the mean and variance of the data before and after the fingerprint embedding as the criterion for judging whether the data is severely distorted, and compare the method proposed in this paper with the currently relatively advanced and mainstream methods. The results are shown in Table 4:

[0125] PE method: Thodi, Diljith M., and Jeffrey J. Rodriguez. "Prediction-error based reversible watermarking." 2004 International Conference on Image Processing, 2004. ICIP’04.. Vol. 3. IEEE, 2004.

[0126] APD method: Siledar, Seema Babusingh, and Sharvari Tamane. "Quadratic difference expansion based Reversible Watermarking for relational database." Journal of Integrated Science and Technology 9.2 (2021): 107 - 112.

[0127] Chin's method: Chang, Chin-Chen, Thai-Son Nguyen, and Chia-Chen Lin. "A Reversible Database Watermark Scheme for Textual and Numerical Datasets." 2021 IEEE / ACIS 22 nd International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel / Distributed Computing (SNPD). IEEE, 2021.

[0128] Ali's method: Hamadou, Ali, et al. "Reversible fragile watermarking scheme for relational database based on prediction-error expansion." Mathematical Problems in Engineering 2020 (2020).

[0129] Table 4 Distortion Comparison

[0130]

[0131] Among them, the PE method is a prediction error method borrowed from image fingerprints. After grouping the data, four adjacent data are taken as a group for fingerprint embedding. This method effectively reduces the distortion of the data. However, since the encoding and decoding of fingerprint data both depend on quadruple data, its anti-attack performance is poor. The APD method is based on a statistical histogram and makes partial changes to the data. This method is relatively traditional and has a high data distortion. The method of Ali is an improvement of the PE method. Although there is an improvement, it still cannot match the performance of the method of the present disclosure in terms of mean and variance.

[0132] 3. Anti-attack experiment. The anti-attack experiment is to verify the effectiveness of the fingerprint when the data is attacked by malicious tampering, etc. After embedding the fingerprint in the data, we respectively perform deletion attacks (DeletionAttack) on it ranging from 10% to 90%, and at the same time extract the fingerprint from the tampered data, using the integrity of the extracted fingerprint as the evaluation criterion.

[0133] As Figure 6 shown, when the deletion attack (attack) reaches 50%, the detection integrity of other algorithms drops significantly. When it reaches 70%, the fingerprints of some algorithms are completely lost. The method of the present disclosure (ours) still has 70% integrity when other related methods fail. At the same time, when the attack area reaches 90%, it still has 40% integrity.

[0134] In summary, the present disclosure: 1. Embed fingerprints in the global data. The fingerprint coverage is wide. When the data is partially deleted or modified, the fingerprint data still exists, with stronger anti-attack performance and improved robustness of the embedded fingerprint; 2. Start from the global statistics of the data, encode the fingerprint according to the overall characteristics of the data, so that the embedding of the fingerprint does not affect the usability related to data statistics.

[0135] See Figure 5 , the present disclosure also provides an electronic device 60, which includes:

[0136] At least one processor; and,

[0137] A memory communicatively connected to the at least one processor; wherein,

[0138] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the foregoing embodiments.

[0139] The present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions for causing the computer to execute any of the foregoing embodiments.

[0140] The present disclosure also provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute any of the foregoing embodiments.

[0141] Reference is made below to Figure 5 , which shows a schematic structural diagram of an electronic device 60 suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0142] As Figure 5 shown, the electronic device 60 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 60 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0143] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 60 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 5 shows an electronic device 60 having various devices, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.

[0144] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0145] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0146] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.

[0147] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: obtain at least two Internet Protocol addresses; send a node evaluation request including the at least two Internet Protocol addresses to a node evaluation device, wherein the node evaluation device selects an Internet Protocol address from the at least two Internet Protocol addresses and returns it; receive the Internet Protocol address returned by the node evaluation device; wherein the obtained Internet Protocol addresses indicate edge nodes in a content delivery network.

[0148] Alternatively, the above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: receive a node evaluation request including at least two Internet Protocol addresses; select an Internet Protocol address from the at least two Internet Protocol addresses; return the selected Internet Protocol address; wherein the received Internet Protocol addresses indicate edge nodes in a content delivery network.

[0149] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0151] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not, in some cases, constitute a limitation on the unit itself. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0152] It should be understood that the various parts of the present disclosure may be implemented in hardware, software, firmware, or a combination thereof.

[0153] As described above, these are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A fingerprint embedding method for tabular data, characterized in that: It includes the following: 1) Calculate the hash value HK of each row of the unique identifier K of the tabular data T through hash calculation; 2) Re - sort the rows of the T in ascending or descending order according to the HK; 3) Group the numerical data of the said T to obtain n groups of data, denoted as , where each group contains k rows, and the number of numerical data columns of the said T except for the unique identifier is k; 4) For each set of data of the said T, use the trained convolutional neural network encoder and fingerprint data respectively to perform fingerprint embedding to obtain ; 5) Combine in the order of the said grouping to obtain tabular data , which is the data table after embedding fingerprints.

2. The method according to claim 1, wherein: The convolutional neural network encoder includes the following 4 layers: The first layer is a neural network with 32 convolutional kernels of size 3x3, expressed as: a = Conv 1→32 (C) Among them, Conv represents the convolutional neural network, 1→32 means: the input C is changed from the initial 1 layer to 32 layers after convolution; The second layer is a neural network with 32 convolutional kernels of size 3x3, expressed as: b = Conv 32→32 (Cat(a, M)) Among them, M is the fingerprint data; cat() is the data concatenation operation; The third layer is a neural network with 32 convolutional kernels of size 3x3, expressed as: E b = Flatten 32→1 (Conv 32→32 (b)) Among them, the flatten operation compresses the features unfolded from 32 layers into one layer; The fourth layer embeds the fingerprint through the following expression: E = C + E b Among them, E is the fingerprint data after embedding the fingerprint M in C.

3. The method according to claim 2, characterized in that: When training the convolutional neural network encoder, the following loss function is used: L d = CrossEntropy(D(E(T i ,M)),T i ) Among them, CrossEntropy() is the cross entropy loss function, E(T i ,M) is T i The result after the fingerprint is embedded by the convolutional neural network encoder, D(E(T i ,M)) is E(T i ,M) The result after the fingerprint is erased by the convolutional neural network decoder; one T is regarded as a batch.

4. A method for extracting fingerprints of tabular data, characterized in that: It includes the following: 1) Calculate the hash value HK of each row of the unique identifier K of the tabular data T' through hash calculation; 2) Re - sort the rows of the T' in ascending or descending order according to the HK; 3) Group the numerical data of the T', obtaining n groups of data, denoted as , where each group contains k rows, and the number of numerical data columns of the T' except for the unique identifier is k; 4) For each set of data of the T', use the trained convolutional neural network decoder to remove fingerprints to obtain ; 5) Calculate the fingerprint data DM through the following formula: 。 5. The method according to claim 4, wherein: The convolutional neural network decoder is a neural network with 32 convolutional kernels of size 3x3, expressed as: S= Flatten 32→1 (Conv 1→32 (E)) Among them, Conv represents the convolutional neural network, 1→32 means: the input E is changed from the initial 1 layer to 32 layers after convolution; the flatten operation compresses the features unfolded from 32 layers into one layer.

6. The method according to claim 5, characterized in that: When training the convolutional neural network decoder, the following loss function is used: L d = CrossEntropy(D(E(T i ,M)),T i ) Among them, CrossEntropy() is the cross-entropy loss function, E(T i , M) is the i-th group T of the table T i after being embedded with fingerprints by the convolutional neural network encoder, and D(E(T i , M)) is the result after erasing the fingerprints of E(T i , M) by the convolutional neural network decoder; one T serves as a batch.

7. A fingerprint embedding system for tabular data, characterized in that: It includes: A hash value calculation module, used to calculate the hash value HK of each row of the unique identifier K of the tabular data T through hash calculation; A re - sorting module, used to re - sort the rows of the T in ascending or descending order according to the HK; Grouping module, used to group the numerical data of T to obtain n groups of data, denoted as , where each group contains k rows, and the number of numerical data columns of T except for the unique identifier is k; A fingerprint embedding module, which is used to perform fingerprint embedding on each group of data of the T by using a trained convolutional neural network encoder and fingerprint data respectively to obtain ; The table integration module is used to Combine them in the order of the described grouping to obtain table data , That is, the data table after embedding fingerprints.

8. A fingerprint extraction system for tabular data, characterized in that: It includes: A hash value calculation module, used to calculate the hash value HK of each row of the unique identifier K of the tabular data T' through hash calculation; A re - sorting module, used to re - sort the rows of the T' in ascending or descending order according to the HK; A grouping module for grouping the numerical data of the T' to obtain n groups of data, denoted as , where each group contains k rows, and the number of numerical data columns of the T' except the unique identifier is k; A fingerprint removal module, which is used to separately perform fingerprint removal on each group of data of the T by using a trained convolutional neural network decoder to obtain ; A fingerprint extraction module, used to calculate the fingerprint data DM through the following formula: 。 9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the method according to any one of claims 1 - 6.

10. A non-transitory computer-readable storage medium, characterized in that, This non - transitory computer - readable storage medium stores computer instructions, and these computer instructions are used to make the computer execute the method according to any one of claims 1 - 6.

Citation Information

Patent Citations

  • Digital fingerprint traitor tracing watermark embedding method based on reference sharing mechanism

    CN113704711A

  • Fragile watermarks

    US20060095775A1