Ownership verification methods, processing methods, equipment and media for structured data sets
By using secret information and specific mathematical properties in structured data sets to calculate check values, watermark data can be identified and ownership confirmed, which solves the problem of easy peeling of watermark marks in existing technologies, achieves the concealment and security of watermark data, and protects the legitimate rights and interests of the data set owners.
Patent Information
- Application Number
- CN202311467146.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-11-06
AI Technical Summary
In the existing technology, the watermark marking method of structured data sets is easy to identify and strip off, making it difficult to effectively confirm ownership, resulting in malicious use by data set theft parties and failing to effectively protect the legitimate rights and interests of data set owners.
The check value is calculated using secret information and specific mathematical properties, the watermark data is identified through preset mathematical rules, and the ownership is confirmed based on the proportion of the watermark data in the structured data set. The watermark data is consistent in appearance with the business data and is difficult to be removed by machines or manual means.
The watermark data is concealed and secure in the structured dataset. Any dataset verifier cannot distinguish the watermark data without obtaining secret information, which increases the difficulty of protecting the ownership of the structured dataset and ensures the legitimate rights and interests of the dataset owner.
Smart Images

Figure CN117521038B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information security, and in particular to a method, processing method, device, and medium for verifying ownership of a structured data set. Background Art
[0002] The concept of watermarking is common in multimedia copyright-related technologies. For example, additional data such as the creator's identity, used for copyright identification, is embedded in multimedia content files such as images, audio, and video, either in a visible or invisible form. This helps determine the copyright ownership of these contents and protect the legitimate rights and interests of the creators. Embedded watermarking technology is widely used to confirm the ownership of unstructured data.
[0003] Structured data composed of fields like phone numbers and ID numbers doesn't support embedded watermarks, so other methods are needed for watermarking. In related technologies, structured data is typically watermarked using column or row watermarks. A column watermark is an additional, meaningless (or marginally meaningful) data field, or simply a decorative markup added to the existing data. A row watermark, on the other hand, involves synthesizing multiple sets of fabricated data based on the original structured data and incorporating them into the dataset, thereby watermarking the structured data. The drawback of column watermarking is that the added meaningless (or marginally meaningful) data fields are easily discernible. Once other individuals or organizations (using dataset theft as an example below) obtain the structured data, they can easily identify and remove the column watermarks through machine learning. Once the column watermarks are removed, the ownership of the structured data becomes difficult to determine. A drawback of row watermarking is that the fabricated structured data often differs significantly from the business data in terms of format or content, making it difficult to fully integrate it. Removing the row watermarks is also relatively easy for dataset theft. After the dataset thief removes the watermark, he or she can maliciously exploit the structured data, making it difficult to guarantee the legitimate rights and interests of the dataset owner. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a method, processing method, device, and medium for verifying ownership of a structured data set.
[0005] One aspect of the present application provides a method for verifying ownership of a structured dataset, comprising the following steps:
[0006] Acquire a structured data set; the structured data set includes a plurality of structured data, each of the structured data is business data or watermark data, and the business data and the watermark data meet the same predetermined data format;
[0007] Obtaining secret information corresponding to the structured data set, a proportional label of the watermark data, and a specific mathematical property corresponding to the watermark data from a target object to be verified; the specific mathematical property is used to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold; and the proportional label is greater than the first threshold;
[0008] identifying the watermark data from the structured data set according to the secret information and the specific mathematical property;
[0009] The proportion of the watermark data in the structured data set is calculated, and the ownership relationship between the target object and the structured data set is determined according to the proportion result and the proportion label.
[0010] Furthermore, in some embodiments, identifying the watermark data from the structured data set based on the secret information and the specific mathematical property includes:
[0011] Calculating the secret information and the structured data using the preset mathematical rule to obtain a first check value;
[0012] Determining, based on the specific mathematical property, whether the first verification value meets the preset mathematical characteristic;
[0013] If the first check value meets the preset mathematical characteristic, the structured data is determined to be watermark data.
[0014] Furthermore, in some embodiments, determining the ownership relationship between the target object and the structured dataset based on the ratio result and the ratio label includes:
[0015] Calculating a difference between the ratio result and the ratio label;
[0016] If the difference value is less than a second threshold, it is determined that the target object is the owner of the structured data set.
[0017] Furthermore, in some embodiments, calculating the difference between the ratio result and the ratio label includes:
[0018] Calculating a difference between the ratio result and the ratio label, and determining an absolute value of the difference as a difference value;
[0019] Alternatively, a difference between the ratio result and the ratio label is calculated, and the ratio of the absolute value of the difference to the ratio label is determined as the difference value.
[0020] Another aspect of the present application discloses a method for processing a structured data set, comprising the following steps:
[0021] Obtaining an original data set and ownership marking information; wherein the original data set is used to store structured data, and the structured data satisfies a predetermined data format; the ownership marking information includes secret information, a proportional label, and a specific mathematical property; the specific mathematical property is used to constrain a check value calculated using the secret information and watermark data according to a preset mathematical rule to conform to a preset mathematical feature; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical feature is less than a first threshold; and the proportional label is greater than the first threshold;
[0022] determining watermark data from data satisfying the predetermined data format according to the secret information and the specific mathematical property;
[0023] Determining a target amount of watermark data to be added to the original data set according to the amount of business data contained in the original data set and the ratio label;
[0024] The target amount of watermark data is added to the original data set to obtain a target data set.
[0025] Furthermore, in some embodiments, obtaining ownership mark information includes:
[0026] Acquire association information corresponding to the original data set; the association information is used to characterize the ownership of the original data set;
[0027] The secret information is generated based on the associated information.
[0028] Furthermore, in some embodiments, adding the target amount of watermark data to the original data set to obtain a target data set includes:
[0029] determining an insertion position of the target quantity in the original data set;
[0030] Each watermark data is added to an insertion position in the original data set to obtain a target data set.
[0031] Another aspect of the present application discloses a device for verifying ownership of a structured data set, comprising:
[0032] A first acquisition unit is configured to acquire a structured data set, wherein the structured data set includes a plurality of structured data, each of which is business data or watermark data, and the business data and the watermark data satisfy the same predetermined data format;
[0033] a second acquisition unit, configured to acquire, from a target object to be verified, secret information corresponding to the structured data set, a proportional label of the watermark data, and a specific mathematical property corresponding to the watermark data; the specific mathematical property being configured to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold; and the proportional label is greater than the first threshold;
[0034] a processing unit, configured to identify the watermark data from the structured data set based on the secret information and the specific mathematical property;
[0035] A statistical unit is used to count the proportion of the watermark data in the structured data set, and determine the ownership relationship between the target object and the structured data set according to the proportion result and the proportion label.
[0036] Another aspect of the present application discloses an electronic device, including a processor and a memory;
[0037] The memory is used to store programs;
[0038] The processor executes the program to implement the method for verifying ownership of a structured data set or the method for processing a structured data set.
[0039] On the other hand, the present application discloses a computer-readable storage medium, which stores a program. The program is executed by a processor to implement the ownership verification method of a structured data set or the processing method of a structured data set.
[0040] The embodiments of the present application have the following beneficial effects: In the ownership verification method, processing method, device and medium of a structured data set of the present application, the watermark data has an appearance consistent with the business data, is difficult to be stripped off by machines or manual means, and can be well hidden in the business data of the structured data set; any data set verification party cannot distinguish between the watermark data and the business data without obtaining secret information from the data set owner, so that the watermark data of the present application has good security and is difficult to be maliciously used. On the other hand, the present application confirms the ownership of the structured data set based on the proportional characteristics formed by the watermark data and the business data, and introduces secret information, which is more difficult to crack than the method of directly confirming ownership through watermark data, and can better protect the legitimate rights and interests of the data set owner. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 It is a schematic diagram of a watermark in the prior art;
[0043] Figure 2 It is a schematic diagram of a watermark in the prior art;
[0044] Figure 3 This is a flow chart of a method for verifying ownership of a structured data set provided in an embodiment of the present application;
[0045] Figure 4 This is a schematic diagram of a process for determining watermark data from structured data provided in an embodiment of the present application;
[0046] Figure 5 This is a flow chart of a method for processing a structured data set provided in an embodiment of the present application;
[0047] Figure 6 This is a schematic diagram of a predetermined data format of business data provided in an embodiment of the present application;
[0048] Figure 7 This is a schematic diagram of a process for generating secret information provided in an embodiment of the present application;
[0049] Figure 8 This is a schematic diagram of inserting watermark data into an original data set provided in an embodiment of the present application;
[0050] Figure 9A schematic diagram of the structure of a device for verifying ownership of a structured data set provided in an embodiment of the present application;
[0051] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0053] The concept of watermarking is common in multimedia copyright-related technologies. For example, additional data such as the creator's identity, used for copyright identification, is embedded in multimedia content files such as images, audio, and video, either in a visible or invisible form. This helps determine the copyright ownership of these contents and protect the legitimate rights and interests of the creators. Embedded watermarking technology is widely used to confirm the ownership of unstructured data.
[0054] In related technologies, watermarking technology for structured data sets is mainly divided into column watermarking and row watermarking.
[0055] Specifically, a column watermark is an additional data field that has no actual meaning (or little actual meaning), or is simply a decorative mark added to the existing data in the format. Figure 1 As shown in the example, the mobile phone number 12300762185 is decorated with {#12300762185#} to identify it as "this is my data." The {# and #} here are column watermarks. The drawback of column watermarks is that the added meaningless (or marginally meaningful) data fields are very easy to identify. Once other individuals or organizations (using the dataset theft as an example below) obtain the structured data, they can easily use machine learning to identify and remove the column watermarks. After the column watermarks are removed, the ownership of the structured data becomes difficult to identify.
[0056] In contrast, row watermarking does not change the data from the column dimension, but generates multiple sets of fake data based on the original structured data and mixes them into the data set. The watermarking of structured data is achieved through these fake data, which is equivalent to inserting "entire fake data records". Figure 2As shown in the example, a group of fake users were inserted into the mobile phone number dataset. Their phone numbers used non-existent numbers, such as 111xxxxxxxx, to identify them as "my data." The drawback of row watermarking is that forged structured data often differs significantly from business data in format or content. For example, numbers like 111xxxxxxxx are easily identified as fake by industry insiders, making them difficult to fully integrate. Removing the row watermarks is also relatively easy for dataset theft.
[0057] The embodiments of the present application improve upon existing watermarking technology and propose a method, processing method, device, and medium for verifying ownership of a structured data set.
[0058] In the embodiments of the present application, the data set owner refers to an object entity that legally enjoys various legal rights and interests related to the structured data set. Ideally, the structured data set of the data set owner will not be maliciously exploited by the data set thief, and the data set owner can normally use the business data in the structured data set for commercial activities and enjoy the legal rights and interests related thereto; however, in the information age, the illegal acquisition and malicious use of structured data sets are common, which seriously affects the information security of the data set owner and hinders the data set owner from enjoying various legal rights and interests. Therefore, it is necessary to use the data watermark technology of the embodiments of the present application to protect the information security of the structured data set.
[0059] In the embodiment of the present application, the dataset verifier refers to the object entity that wants to confirm the ownership of the structured dataset.
[0060] To solve the problem of ownership confirmation of structured data sets, such as Figure 3 As shown, the embodiment of the present application proposes a method for verifying the ownership of a structured data set, which can be applied to the data set verification party. Specifically, the method includes the following steps:
[0061] Step 310: Acquire a structured data set; the structured data set includes a plurality of structured data, each of which is business data or watermark data, and the business data and the watermark data meet the same predetermined data format;
[0062] Step 320: Obtain secret information corresponding to the structured data set, a scale label of the watermark data, and a specific mathematical property corresponding to the watermark data from the target object to be verified; the specific mathematical property is used to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold;
[0063] Step 330: Identify the watermark data from the structured data set based on the secret information and the specific mathematical property;
[0064] Step 340: Count the proportion of the watermark data in the structured data set, and determine the ownership relationship between the target object and the structured data set based on the proportion result and the proportion label.
[0065] In an embodiment of the present application, a method for verifying the ownership of a structured dataset is provided, which can be applied by a dataset verifier. Specifically, first, a structured dataset requiring ownership verification can be obtained and an object to be verified can be determined. In this embodiment of the present application, this object is referred to as a target object. The target object can be a related object claiming to hold the structured dataset requiring ownership verification, or an object providing the structured dataset to the dataset verifier, although this is not limited in this embodiment of the present application.
[0066] In an embodiment of the present application, the structured data set obtained includes multiple structured data, and these structured data meet the same predetermined data format, for example, the number of fields of each structured data is the same and each corresponding field has the same data format. Here, each structured data in the structured data set is business data or watermark data, and business data refers to normal real data, for example, it may include fields such as mobile phone number and ID card number; watermark data is constructed false data. In an embodiment of the present application, the business data and watermark data meet the same predetermined data format, and there is no difference between the two in format and content. In other words, in an embodiment of the present application, the watermark data is not intuitively marked (such as the aforementioned setting of the synthesized mobile phone number to start with a prefix that does not actually exist, such as 111), and the watermark data in the embodiment of the present application does not have any warning function. In this way, in the absence of corresponding secret information, no one can determine whether the structured data is business data or watermark data, that is, no one can distinguish between business data and watermark data.
[0067] Then, the secret information corresponding to the structured dataset, the ratio label of the watermark data, and the specific mathematical properties of the watermark data can be obtained from the target object. In the embodiment of the present application, for the owner of the structured dataset, it is necessary to pre-select the secret information, the ratio label of the watermark data, and the specific mathematical properties of the watermark data for each structured dataset with a claim requirement, wherein the secret information must be kept confidential and authorized only to the dataset verifier (such as the regulatory department or law enforcement agency) when necessary. Based on the secret information, the dataset owner can generate watermark data with specific mathematical properties; a watermark data is indistinguishable from a business data in appearance, but for all possible data that meet the predetermined data format, only a small proportion of the synthetic data can meet the specific mathematical properties pre-set by the dataset owner and thus become watermark data. Therefore, in the embodiment of the present application, the dataset owner can control the ratio of the watermark data inserted in the structured dataset and record the ratio value as the ratio label. After obtaining the secret information, the dataset verifier can use the secret information to identify each watermark data, and then determine whether the dataset belongs to the target object by checking the appearance ratio of the watermark data in the structured dataset.
[0068] In the embodiments of the present application, the specific mathematical property corresponding to the watermark data refers to the fact that the check value calculated using the secret information and the watermark data, through a preset mathematical rule, conforms to the preset mathematical characteristics, that is, the check value has a specific mathematical regularity or characteristic. Here, there are multiple ways to calculate the check value using the secret information and the watermark data, which can be implemented using relevant cryptographic rules. For example, a message authentication code (Message Authentication Code) can be used as the check value. The message authentication code is a special type of algorithm in the field of cryptography. This type of algorithm inputs a specified key and data (the former is secretly selected in advance; both can be readable strings or arbitrary bit strings) and outputs a data fingerprint (also called a data digest or simply a message authentication code) of standard length (e.g., 256 bits) in the form of a random number. Although the formula of the message authentication code is not complex, its output is unpredictable and can only be determined after both the key and data are input. The mathematical properties of message authentication codes also include: (a) the data fingerprint is sensitive to both the key and the data; a change in either input (even a single bit) results in a completely different output; and (b) calculating the data fingerprint from the key and data is easy, but for sufficiently strong keys and data from a sufficiently large sample space, it is unrealistic to infer either the key or the data from the data fingerprint. In the embodiments of the present application, when using a message authentication code algorithm, secret information can be used as the key and structured data as the data, and the algorithm can be used to calculate the check value corresponding to each piece of structured data.
[0069] In the embodiments of the present application, the check value calculated using the secret information and watermark data conforms to a preset mathematical characteristic. The specific form of the mathematical characteristic is not limited herein. For example, for example, using a message authentication code as the check value, the preset mathematical characteristic may be: in some embodiments, the check value corresponding to the watermark data begins with 16 consecutive 1 bits, i.e., begins with two bytes ffff in hexadecimal; in some embodiments, the check value corresponding to the watermark data ends with 20 consecutive 0 bits; in some embodiments, the second and penultimate bytes of the check value corresponding to the watermark data are both the hexadecimal number aa. It is understood that since the message authentication code is formally a random number of fixed length, each of the above examples is a low-probability event. Therefore, specific mathematical properties can restrict only a small portion of the synthesized data to be watermark data. In other words, it can be assumed that the probability that the check value calculated using any data satisfying the predetermined data format and the secret information, calculated using a preset mathematical rule, conforms to the preset mathematical characteristic is less than a first threshold. The first threshold can be determined based on the total amount of data in the predetermined data format and actual needs. Generally speaking, the value of the first threshold should be as small as possible to reduce potential interference with normal business data. For example, the first threshold can be set to 2 -10 In the embodiment of the present application, there is no restriction on the size of the first threshold.
[0070] It is understandable that in the embodiment of the present application, based on the set specific mathematical properties, after the data set verifier obtains the secret information with the authorization of the target object (and only after obtaining the secret information), it can verify whether the structured data meets the specific mathematical properties one by one, thereby identifying whether each structured data is watermark data. Accordingly, the legitimate data set owner can reveal to the data set verifier those watermark data that appear to belong to business data mixed in the structured data set, and prove the claim of ownership of the structured data set based on the appearance ratio of the watermark data in the structured data set. A key effect of this process is that only the data set owner and the authorized data set verifier can distinguish between business data and watermark data in the structured data set, and therefore only they can count the proportion of watermark data appearance. If the target object is not the legal owner of the dataset, it cannot obtain the correct secret information. Furthermore, if the target object cannot provide the secret information, it can be determined that it is not the legal owner of the structured dataset; or if the target object provides incorrect secret information, then the check value calculated based on the incorrect secret information and the structured data cannot correctly identify whether each structured data is watermark data based on the set specific mathematical properties, and the obtained proportion result will be far different from the correct proportion label, so it can also be determined that it is not the legal owner of the structured dataset.
[0071] Of course, it should be noted that in the embodiments of the present application, in order to more accurately and clearly determine the target object's ownership of the structured dataset, the dataset owner generally needs to add a certain amount of watermark data to the structured dataset, so that the proportion of the watermark data far exceeds the probability that the check value calculated from conventional data that meets the predetermined data format and the secret information meets the preset mathematical characteristics. In other words, the value of the ratio tag must be far greater than the first threshold. This ensures that it can be clearly distinguished that the structured dataset has been processed based on the watermark data. In the embodiments of the present application, there is no specific limitation on the size of the ratio tag; for example, it can be set to 5%.
[0072] It is understandable that compared with similar technologies, this application has the following four main features:
[0073] 1. It can be used to prove the ownership claim of a data set by the owner of the data set, even if the fields in the structured data set contain "unmodifiable data" such as ID card number and mobile phone number;
[0074] 2. Similar to some watermark-related algorithms, this application involves cryptographic technology, but this application is not limited to a specific cryptographic algorithm. As long as the cryptographic type that can generate a check value for structured data based on secret information (such as a message authentication code or a deterministic digital signature) can be used as the underlying algorithm of the mathematical rules preset in this application;
[0075] 3. Unlike image watermarks or audio / video watermarks, which extract information such as the producer's identity from the data to perform copyright identification such as content traceability, this application uses secret information to verify each piece of structured data in the structured data set one by one to identify the watermark data. The ownership of the structured data set is then determined based on the proportion of the watermark data appearing in the structured data set. The one-by-one verification here is to check whether the check value of each piece of structured data satisfies specific mathematical properties.
[0076] 4. When the target object proves its ownership claim of structured data to the dataset verifier (such as regulatory authorities or law enforcement agencies), it needs to authorize the latter to obtain confidential information relative to the structured data; unauthorized persons cannot perform the above verification.
[0077] In some embodiments, reference Figure 4 , identifying the watermark data from the structured data set according to the secret information and the specific mathematical property, comprising:
[0078] Calculating the secret information and the structured data using the preset mathematical rule to obtain a first check value;
[0079] Determining, based on the specific mathematical property, whether the first verification value meets the preset mathematical characteristic;
[0080] If the first check value meets the preset mathematical characteristic, the structured data is determined to be watermark data.
[0081] In an embodiment of the present application, when determining watermark data from structured data, a check value can be obtained by calculating using secret information and structured data according to a preset mathematical rule, and this check value is recorded as a first check value. Subsequently, based on a specific mathematical property, it can be determined whether the first check value meets a preset mathematical characteristic. If the first check value meets the preset mathematical characteristic, it can be determined to be watermark data; conversely, if the first check value does not meet the preset mathematical characteristic, it can be determined to be business data.
[0082] In some embodiments, determining the ownership relationship between the target object and the structured dataset based on the ratio result and the ratio label includes:
[0083] Calculating a difference between the ratio result and the ratio label;
[0084] If the difference value is less than a second threshold, it is determined that the target object is the owner of the structured data set.
[0085] In the embodiment of the present application, when determining the ownership relationship between the target object and the structured data set based on the ratio result and the ratio label, the business data may have certain mathematical properties, and the structured data set may also be modified, added, deleted, etc. by the data set theft party. Therefore, the ratio result and the ratio label here may not be completely consistent. In the embodiment of the present application, the difference value between the ratio result and the ratio label can be calculated, and the difference value here can be flexibly set as needed. For example, in some embodiments, the difference between the ratio result and the ratio label can be calculated, and then the absolute value of the difference is determined as the difference value; in some embodiments, the ratio of the absolute value of the difference to the ratio label (or ratio result) can be calculated, and the ratio is determined as the difference value. It can be understood that in the embodiment of the present application, the greater the difference between the ratio result and the ratio label, the less close the two are, and the less likely the target object is the owner of the structured data set; the smaller the difference between the ratio result and the ratio label, the closer the two are, and the more likely the target object is the owner of the structured data set. In an embodiment of the present application, a threshold can be set, which is recorded as the second threshold. If the calculated difference value is very small and is within the second threshold, it can be determined that the target object is the owner of the structured data set; if the calculated difference value is large and is greater than or equal to the second threshold, it can be determined that the target object is not the owner of the structured data set.
[0086] Reference Figure 5 In an embodiment of the present application, a method for processing a structured data set is also provided. The method can be used to generate a structured data set with watermark data and can be used by the data set owner. Specifically, the processing method includes:
[0087] Step 510: Obtain an original data set and ownership tag information; wherein the original data set is used to store structured data, and the structured data satisfies a predetermined data format; the ownership tag information includes secret information, a proportional tag, and a specific mathematical property; the specific mathematical property is used to constrain a check value calculated using the secret information and watermark data according to a preset mathematical rule to conform to a preset mathematical feature; wherein, according to the preset mathematical rule, the probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical feature is less than a first threshold; and the proportional tag is greater than the first threshold;
[0088] Step 520: Determine watermark data from data that meets the predetermined data format based on the secret information and the specific mathematical property.
[0089] Step 530: Determine the target amount of watermark data to be added to the original data set according to the amount of business data contained in the original data set and the ratio label;
[0090] Step 540: Add the target amount of watermark data to the original data set to obtain a target data set.
[0091] In an embodiment of the present application, a method for processing a structured data set is provided, which can be used by the owner of the data set. Specifically, the owner of the data set can obtain the original data set and ownership mark information, wherein the original data set is a structured data set, which can be used to store related structured data. These structured data meet the predetermined data format, and watermark data will be searched and filtered out based on the predetermined data format. In the embodiment of the present application, there is no limitation on the specific situation of the predetermined data format. For example, Figure 6 As shown, when the original data set is used to store mobile phone numbers, the predetermined data format presented is a data structure of country code + domestic destination code + user number; similarly, when the original data set is used to store user addresses, the predetermined data format presented is a data structure of province + city + county / district + town / street + community. In the embodiments of the present application, understanding the predetermined data format helps to produce watermark data similar to real business data, so that without specific secret information, no party can determine whether the structured data contains business data or watermark data, that is, no one can distinguish business data from watermark data.
[0092] It should be noted that, in the embodiment of the present application, all the data in the target data set may be set to be watermark data. In this case, the obtained original data set may not include any real business data.
[0093] In an embodiment of the present application, the ownership marking information is used to realize the ownership marking of the original data set, which may include secret information, ratio labels and specific mathematical properties. The specific meaning of this information has been introduced in the aforementioned embodiment and will not be repeated here. In an embodiment of the present application, there is usually no coupling relationship between the three types of information in the ownership marking information, so there is no order of precedence when selecting. Among them, secret information must be kept confidential and authorized only to a specific data set verifier (regulatory department, law enforcement agency) when necessary; other information needs to be announced to the data set verifier and can also be made public to the whole society.
[0094] In an embodiment of the present application, based on the set ownership mark information, a candidate data can be selected from the data that meets the predetermined data format according to a certain strategy (random selection or traversal in a certain order), and calculated according to the preset mathematical rules. If the calculation result just meets the selected specific mathematical property, the candidate data is output as watermark data, and if it does not meet the selected specific mathematical property, it is ignored (the candidate data is not watermark data). In this way, the watermark data can be determined from the data that meets the predetermined data format. In an embodiment of the present application, the target number of watermark data that needs to be added to the original data set is also determined based on the number and proportion label of the business data contained in the original data set. For example, assuming that the original data set has 9,500 business data and the proportion label is 5%, 500 watermark data need to be obtained. If the specified number of watermark data has been obtained, the watermark data can be mixed with the business data so that the proportion of watermark data in the original data set is equal to the pre-selected proportion label, so that the processed target data set can be obtained.
[0095] In the embodiments of this application, the watermark data is visually indistinguishable from the service data. Of all possible data, only a small percentage satisfy specific mathematical properties and thus serve as watermark data. This means that a large amount of data must be individually calculated to identify the relatively small number of watermark data that satisfy these specific mathematical properties. The specific search process can be performed randomly or by traversing the data based on certain conditions, and this application does not impose any restrictions on this.
[0096] In some embodiments, reference Figure 7 , obtain ownership mark information, including:
[0097] Acquire association information corresponding to the original data set; the association information is used to characterize the ownership of the original data set;
[0098] The secret information is generated based on the associated information.
[0099] In some cases, in order to facilitate the clarification of the ownership of the relevant data set, in embodiments of the present application, information with natural semantics can be used to generate secret information when processing the data set, thereby facilitating subsequent possible attribute proof operations. Specifically, in embodiments of the present application, for the original data set, associated information corresponding to it can be obtained. The associated information here can be information used to characterize the ownership of the original data set, such as "Company A XXX Exclusive". Then, secret information can be generated based on the associated information. For example, the associated information can be used directly as secret information, or it can be processed and used as secret information. This application does not impose any restrictions on this.
[0100] In some embodiments, adding the target amount of watermark data to the original data set to obtain a target data set includes:
[0101] determining an insertion position of the target quantity in the original data set;
[0102] Each watermark data is added to an insertion position in the original data set to obtain a target data set.
[0103] In the embodiment of the present application, a target amount of watermark data is added to the original data set, which can be achieved by using a mixed method. Figure 8 As shown, the mixing mentioned in the embodiment of the present application refers to inserting watermark data into the original business data relatively evenly, so that it is impossible to judge whether a piece of data is business data or watermark data based on its position in the data set (the row number in the data table). Specifically, a target number of insertion positions can be determined in the original data set, and then each watermark data is added to an insertion position to obtain a target data set. For example, the original data set has 9,500 business data. After randomly inserting 500 watermark data, the resulting structured data set has a total of 10,000 data, of which 5% are watermark data, but it is impossible to determine where they appear, and there is no pattern. In this way, the ownership security of the structured data set can be improved.
[0104] The following introduces and illustrates the technical solution of this application with reference to specific application scenario examples.
[0105] The embodiments of this application are applicable to the case where business data contains only a single field, and are also applicable to the case where business data contains multiple fields. One embodiment is described below for each case. The examples are intended to illustrate the concepts and mathematical calculation processes involved in this application, and do not mean that the situation is like this in the real world (for example, the mobile phone numbers in the 123 network segment used in the embodiments are imaginary). To make the description concise and facilitate the verification of the embodiments, the following technical conventions are uniformly adopted in the examples below:
[0106] Strings are encoded according to the internationally common UTF-8 rule. For example, the encoding result of the string "数学" composed of two Chinese characters is an array composed of 6 bytes, and its hexadecimal representation is e695b0e5ada6. To be compatible with most programming languages, the watermark data verification uses the internationally popular Message Authentication Code HMAC-SHA-256, and the calculation result is an array composed of 32 bytes. See IETF RFC 4231 for relevant test vectors.
[0107] The above technical conventions are only for better describing the embodiments, and do not mean that this application is restricted in any way in terms of generality. In actual applications, this application is neither limited to a specific encoding nor to a specific cryptographic algorithm. For example, the character encoding can adopt the Chinese standard GB 18030, etc. The underlying message authentication code for the mathematical properties of the watermark data can adopt the Chinese standard HMAC-SM3 or CMAC-SM4, or KMAC based on SHA-3 internationally, etc.
[0108] Embodiment 1: Provide attribute proof for a data set containing only a single field.
[0109] Suppose a certain operator needs to use a batch of mobile phone numbers in the 123 network segment in a test project carried out in February 2024. For differentiation, the company only uses mobile phone numbers whose "message authentication code starts with 16 consecutive 1 bits, that is, starts with ffff in hexadecimal", and this batch of numbers will no longer be assigned to normal services. In other words, the mobile phone numbers whose corresponding message authentication code starts with ffff are the watermark data (only the company itself knows the secret information required for calculating the message authentication code). The company only uses such mobile phone numbers in the test project, which means that the proportion of watermark data in the generated structured data set is 100%. According to the processing method of the structured data set provided by this application, as the owner of the data set, the company can generate a data set composed entirely of watermark data, such as 12300023180, 12300034919, 12300078978, 12300088650, 12300393151, 12300421487, 12300600146, 12300814686, 12300857998, 12301037953... by incrementally traversing starting from 12300000000.
[0110] Assume that after the project begins, these phone numbers are exposed. To allay potential concerns, the company declares to regulators that all phone numbers are reserved for testing purposes (not assigned to real individuals). The company also discloses to them (and only to the regulators) the secret information used to calculate the message authentication code, which was pre-selected before the project began: "China X Co., Ltd., February 2024 Testing Only" (this key is for example purposes only; the actual secret information must be unguessable).
[0111] As the data set verifier, the regulatory authorities encoded the secret information and each mobile phone number as a string into UTF-8, and substituted it into the HMAC-SHA-256 formula to calculate the message authentication code shown in Table 1 below (due to space limitations, only 10 data items are shown):
[0112] Table 1
[0113]
[0114]
[0115] The company also informed the regulatory authorities that the watermark data has a specific mathematical property that the message authentication code obtained begins with 16 consecutive 1 bits. For any random data, the probability of meeting this mathematical property is 2 -16 power; that is, the probability of a piece of data being watermarked is only approximately 1.5 in 100,000. As shown in Table 1, regulatory authorities verified that all exposed mobile phone numbers were watermarked data. Therefore, there is reason to believe that these numbers were pre-selected test numbers by the company, rather than being assigned to real individuals and then leaked in a security incident. If a data set were leaked in a security incident, it would be impossible to later identify the secret information that ensures that all data satisfy the specific mathematical properties (thus confirming it as watermarked data). Furthermore, the secret information provided by the company, "China X Co., Ltd., February 2024, for testing purposes only," itself limits the ownership of the data set, further proving that this structured data set was pre-selected test numbers.
[0116] Example 2: Providing attribute proof for a dataset containing multiple fields.
[0117] When verifying data containing multiple fields, all key fields need to participate in the calculation (to determine whether they meet specific mathematical properties). From the source, when all parties generate watermark data, all key fields must participate in the logical operation process and obtain authentication information (all HMAC-SHA-256 message authentication codes in the embodiments), and one or more fields in all key fields may be composite values depending on the situation. In this process, a simple way to allow all key fields to participate in the logical operation is to directly concatenate the values of all key fields as strings and then encode them (all UTF-8 encoding in the embodiments).
[0118] Suppose two e-commerce companies, A and B, are competitors. Company A's customer data (consisting of at least three fields: ID number, mobile phone number, and name) has been illegally stolen by Company B for a long time. To provide attribute verification for its customer dataset, Company A selects a secret message each quarter and generates and inserts watermark data into 5% of its transaction data (one watermark is inserted for every 19 transactions, with the insertion position randomly selected). Without access to the secret message, Company B cannot identify these watermarks and is even unaware of their existence.
[0119] Assume that the following Table 2 is a sample of watermark data inserted by Company A into its customer dataset in the first quarter of 2025 (each field can be synthetic). Each row of data can be concatenated and then encoded to calculate the message authentication code according to the aforementioned "simple approach":
[0120] Table 2
[0121] ID number Phone number Name 110101190212280049 12300762185 Boss Liu 110102190112290078 12303833627 Wang Xiaoer 110103190012300255 12300144692 Zhang San 11010418991231081X 12301137414 Li Si
[0122] When the law enforcement agency arrests Company B for illegal activities, Company A discloses to the law enforcement agency the secret information corresponding to a certain data set, which is "China Party A Co., Ltd. 2025 First Quarter Test Special", in which the specific mathematical property of the watermark data is that the obtained message authentication code ends with 20 consecutive 0 bits. For any data, the probability of meeting this mathematical property is 2 -20 power; that is, the probability of a random data becoming watermark data is less than 1 in a million.
[0123] After obtaining the above-mentioned secret information with Company A's authorization, the law enforcement agency verified the seized data set and found that approximately 5% of the data was indeed watermark data that met the specific mathematical properties claimed by Company A. An example is shown in Table 3 below:
[0124] Table 3
[0125]
[0126] Understandably, if Company B had not stolen Company A's data, the probability of Company A's watermark appearing in the seized dataset would be less than 1 in a million. Because Company A's watermarked data accounted for approximately 5% of the seized dataset, law enforcement authorities confirmed the ownership of the data and acknowledged Company A's claim that Company A's data was stolen by Company B.
[0127] It is understandable that in the embodiments of the present application, the watermark data is hidden in the business data, making the two indistinguishable in terms of format, content, and other aspects. Without specific secret information, no one can determine whether the data is business data or watermark data. The present application confirms the ownership of a structured data set based on the proportional characteristics formed by the watermark data and the business data, and introduces secret information. Compared with the method of directly confirming ownership through watermark data, it is more difficult to crack and can better protect the legitimate rights and interests of the data set owner.
[0128] The following describes an apparatus for verifying ownership of a structured data set according to an embodiment of the present application with reference to the accompanying drawings.
[0129] Reference Figure 9 The ownership verification device for a structured data set proposed in the embodiment of the present application includes:
[0130] A first acquiring unit 910 is configured to acquire a structured data set, wherein the structured data set includes a plurality of structured data, each of which is business data or watermark data, and the business data and the watermark data satisfy the same predetermined data format;
[0131] A second acquisition unit 920 is configured to acquire, from the target object to be verified, the secret information corresponding to the structured data set, the proportional label of the watermark data, and a specific mathematical property corresponding to the watermark data; the specific mathematical property is configured to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold; and the proportional label is greater than the first threshold;
[0132] a processing unit 930, configured to identify the watermark data from the structured data set according to the secret information and the specific mathematical property;
[0133] The statistical unit 940 is configured to calculate a proportion of the watermark data in the structured data set, and determine the ownership relationship between the target object and the structured data set based on the proportion result and the proportion label.
[0134] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0135] Reference Figure 10 , an embodiment of the present application provides an electronic device, including:
[0136] at least one processor 1010;
[0137] at least one memory 1020, configured to store at least one program;
[0138] When the at least one program is executed by the at least one processor 1010 , the at least one processor 1010 implements a method for verifying ownership of a structured data set or a method for processing a structured data set.
[0139] Similarly, the contents of the above method embodiments are applicable to the present electronic device embodiment. The functions specifically implemented by the present electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0140] An embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by the processor 1010. When executed by the processor 1010, the program executable by the processor 1010 is used to execute the above-mentioned structured data set ownership verification method or structured data set processing method.
[0141] Similarly, the contents of the above method embodiments are applicable to the computer-readable storage medium embodiments. The functions specifically implemented by the computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0142] In some optional embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, the two boxes shown in succession may actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logic flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0143] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0144] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0145] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0146] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0147] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0148] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.
[0149] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0150] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A method for verifying ownership of a structured data set, characterized in that: The following steps are involved: Acquire a structured data set; the structured data set includes a plurality of structured data, each of the structured data is business data or watermark data, and the business data and the watermark data meet the same predetermined data format; Obtaining secret information corresponding to the structured data set, a proportional label of the watermark data, and a specific mathematical property corresponding to the watermark data from a target object to be verified; the specific mathematical property is used to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold; and the proportional label is greater than the first threshold; identifying the watermark data from the structured data set according to the secret information and the specific mathematical property; The proportion of the watermark data in the structured data set is calculated, and the ownership relationship between the target object and the structured data set is determined according to the proportion result and the proportion label.
2. The ownership verification method of a structured data set according to claim 1, characterized in that: The step of identifying the watermark data from the structured data set according to the secret information and the specific mathematical property comprises: Calculating the secret information and the structured data using the preset mathematical rule to obtain a first check value; Determining, based on the specific mathematical property, whether the first verification value meets the preset mathematical characteristic; If the first check value meets the preset mathematical characteristic, the structured data is determined to be watermark data.
3. The method for verifying ownership of a structured data set according to claim 1, wherein: Determining the ownership relationship between the target object and the structured dataset based on the ratio result and the ratio label includes: Calculating a difference between the ratio result and the ratio label; If the difference value is less than a second threshold, it is determined that the target object is the owner of the structured data set.
4. The ownership verification method of a structured data set according to claim 3, characterized in that: The calculating the difference between the ratio result and the ratio label includes: Calculating a difference between the ratio result and the ratio label, and determining an absolute value of the difference as a difference value; Alternatively, a difference between the ratio result and the ratio label is calculated, and the ratio of the absolute value of the difference to the ratio label is determined as the difference value.
5. A method for processing a structured data set, characterized in that: The following steps are involved: Obtaining an original data set and ownership marking information; wherein the original data set is used to store structured data, and the structured data satisfies a predetermined data format; the ownership marking information includes secret information, a proportional label, and a specific mathematical property; the specific mathematical property is used to constrain a check value calculated using the secret information and watermark data according to a preset mathematical rule to conform to a preset mathematical feature; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical feature is less than a first threshold; and the proportional label is greater than the first threshold; determining watermark data from data satisfying the predetermined data format according to the secret information and the specific mathematical property; Determining a target amount of watermark data to be added to the original data set according to the amount of business data contained in the original data set and the ratio label; The target amount of watermark data is added to the original data set to obtain a target data set.
6. The method for processing a structured data set according to claim 5, characterized in that: Obtain ownership mark information, including: Acquire association information corresponding to the original data set; the association information is used to characterize the ownership of the original data set; The secret information is generated based on the associated information.
7. The method for processing a structured data set according to claim 5, wherein: Adding the target amount of watermark data to the original data set to obtain a target data set includes: determining an insertion position of the target quantity in the original data set; Each watermark data is added to an insertion position in the original data set to obtain a target data set.
8. A device for verifying ownership of a structured data set, characterized in that: include: A first acquisition unit is used to acquire a structured data set; The structured data set includes a plurality of structured data, each of the structured data is business data or watermark data, and the business data and the watermark data meet the same predetermined data format; a second acquisition unit, configured to acquire, from a target object to be verified, secret information corresponding to the structured data set, a proportional label of the watermark data, and a specific mathematical property corresponding to the watermark data; the specific mathematical property being configured to constrain a check value calculated using the secret information and the watermark data according to a preset mathematical rule to conform to a preset mathematical characteristic; wherein, according to the preset mathematical rule, a probability that a check value calculated using any data satisfying the predetermined data format and the secret information conforms to the preset mathematical characteristic is less than a first threshold; and the proportional label is greater than the first threshold; a processing unit, configured to identify the watermark data from the structured data set based on the secret information and the specific mathematical property; A statistical unit is used to count the proportion of the watermark data in the structured data set, and determine the ownership relationship between the target object and the structured data set according to the proportion result and the proportion label.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Dynamic watermark embedding and verifying method and system and dynamic watermark processing system
CN109740316A
Data watermark embedding method, watermark tracing method and device
CN112948895A