A data rights confirmation method based on symbol mapping coding and blockchain
By using symbol mapping encoding and blockchain technology in the data sharing system, user fingerprint information is embedded in the data object, and the problems of third-party untrustworthiness, single point of failure and inefficiency in the existing technology are solved, and strong access control and anti-refusal and anti-piracy capabilities that are not related to the content of data rights confirmation are achieved.
Patent Information
- Application Number
- CN202211591493.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-12-12
AI Technical Summary
The existing data rights confirmation methods have problems such as third-party untrustworthiness, single point of failure, inefficiency and data content dependence, which is difficult to effectively prevent data debunk and piracy.
Using a data rights confirmation method based on symbol mapping encoding and blockchain, trusted access control and rebel tracking is achieved by embedding user fingerprint information into data objects and using blockchain to supervise the data sharing process.
It realizes strong access control that is not related to the content of data rights confirmation, enhances the anti-refusal and piracy capabilities of the data sharing system, and avoids single point of failure and inefficiency.
Smart Images

Figure CN116127429B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of data sharing and data right confirmation, and specifically to a data right confirmation method based on symbol mapping coding and blockchain. Background Art
[0002] Existing data right confirmation schemes can be divided into three groups: methods based on a trusted third party (TTP), methods based on blockchain, and other methods.
[0003] (1) Data right confirmation methods based on a trusted third party
[0004] TTP-based schemes usually use a third party to supervise the communication between data sharing participants and provide evidence for data right confirmation. The TTP in these schemes acts as a witness and an evidence verifier. In case of a dispute, the TTP will provide arbitration evidence. The problem is that the so-called TTP may not be that trustworthy. Coffey et al. proposed a TTP-based scheme to solve the general non-repudiation problem. In this scheme, the third party is relied on without reservation, which may increase the possibility of collusion attacks. To solve this problem, Zhu et al. proposed an anti-collusion attack data sharing scheme based on an asymmetric cryptosystem and the Delov-Yao model. However, since the server does not verify user registration, this model cannot resist man-in-the-middle attacks and data tampering attacks. In addition, under a centralized architecture, single point of failure is also a problem that cannot be ignored. Therefore, the TTP-based method does not meet the requirements of data right confirmation, and it has the disadvantages that data is easily tampered with and the third party is not completely trustworthy.
[0005] (2) Data right confirmation methods based on blockchain
[0006] Blockchain has been introduced as an infrastructure into data sharing to provide irrefutable evidence for confirming data permissions. The excellent functions of blockchain are often used to achieve the traceability of data sources and their sharing histories. Some research combines blockchain and digital watermarking to improve the security of copyright protection and data source tracking. Typically, Wang et al. proposed a combined model that first uses a proof of data possession method to audit data integrity and then uses a digital watermarking scheme to confirm the source of shared data. However, this combination is still not sufficient to combat data piracy because not much emphasis is placed on access control.
[0007] The combination of blockchain and encryption systems helps to improve access control in data sharing, but it is usually inefficient when faced with a large amount of data. Wang et al. proposed a big data sharing scheme that uses smart contracts to enforce access rules. However, encrypting and decrypting a large amount of data takes a lot of time, resulting in low efficiency.
[0008] (3) Other methods
[0009] Zero - Knowledge Proof (ZKP) and Non - Fungible Token (NFT) have inspired researchers to develop some characteristic data rights confirmation methods. Some new research has constructed ZKP models to prove the ownership or right of use of assets to third parties. Some other research mainly generates unique digital certificates (such as NFTs) to claim the ownership of specific data assets and realizes secure data sharing by selling these NFTs. However, the above - mentioned technologies have some intolerable defects in data rights confirmation for data sharing. For example, ZKP construction usually leads to low efficiency and poor scalability. So far, NFT is still an immature concept technically.
[0010] Regarding the serious data repudiation and piracy behaviors that occur during the data sharing process, existing data watermarking methods are not only restricted by the limitations of data content but also prone to watermark forgery due to the lack of supervision. Additionally, methods based on trusted third parties may lead to data security and other problems due to the untrustworthiness of the third party. Therefore, in this context, it is particularly important and necessary to invent an efficient data rights confirmation method. Summary of the Invention
[0011] The purpose of the present invention is to provide a data rights confirmation method based on symbol mapping encoding and blockchain, which can effectively confirm the rights of data, achieve trusted access control, and traitor tracing.
[0012] The present invention is implemented as follows: A data rights confirmation method based on symbol mapping encoding and blockchain includes the following steps:
[0013] a. The data owner uses SMC technology to map the fingerprint F D , the private - key hash h(SK D ) and the original data d together into a symbol mapping table SMT D , and then uses the symbol mapping table SMT D to encode the original data to obtain data associated with the fingerprint . After that, the data owner publishes the fingerprint F D , the hash of the data associated with the fingerprint the data description DDes and the hash of the public part of the symbol mapping table together to the blockchain to provide evidence for later data rights confirmation;
[0014] b. The data owner sends the data associated with the fingerprint to the data receiver through an off - chain channel;
[0015] c. The data receiver obtains from the blockchain and makes a hash comparison with the data obtained from the off - chain channel .
[0016] d. When the data receiver applies to the data owner for querying data, the data receiver first uses its own fingerprint F U , the private key hash h(SK U ) and the received data associated with the data owner's fingerprint to map them together into its own symbol mapping table SMT U , and through this symbol mapping table SMT U encode the received data to obtain the data associated with its own fingerprint Then, through the blockchain, the hash of the public part of the symbol mapping table and the query request q are sent to the data owner;
[0017] e. The data receiver sends the public part of the symbol mapping table in step d to the data owner through an off-chain channel ;
[0018] f. After the data owner obtains the query request q sent by the data receiver from the blockchain, it first verifies whether the query request q meets the requirements of the data description and access control; if so, it combines the private part of the symbol mapping table generated by itself and the query request q to generate Then, it combines and the F U of the data receiver to generate a data decoder required for the data receiver to query data; if not, it returns null;
[0019] g. The data owner sends the data decoder to the data receiver through the blockchain;
[0020] h. The data receiver obtains the data decoder and uses the data decoder to restore the applied data, thereby obtaining the applied data;
[0021] In step h, after receiving the data decoder, the data receiver first inputs the data associated with its own fingerprint and the private part of its own symbol mapping table so that the data decoder can verify whether it has the permission to query data; when the verification passes, the data decoder returns the data applied by the data receiver, otherwise the data decoder does not return any data;
[0022] i. When the data owner discovers data in the network, the data owner applies to the arbitrator for rights protection, and the arbitrator determines the data right holder by identifying the fingerprint in the data.
[0023] There are two methods for identifying fingerprints, as follows:
[0024] First, adopt the method of forward verification. By counting the characters in the cipher part of the SMT - and cropping and dividing the characters according to the fingerprint length, compare these divided fields with the fingerprint in turn, count the number of successful comparisons, and when it exceeds a certain proportion, it is considered a successful comparison and its rights are confirmed.
[0025] Secondly, adopt the method of backward verification. It first counts the byte sequence in the data, and then reversely estimates the clear text byte sequence corresponding to the fingerprint in the data, and compares these two byte sequences to confirm its rights.
[0026] Backward verification is to supplement the forward verification in case the forward verification result is inaccurate. As long as one of the forward verification and the backward verification can confirm the rights, it is considered that the rights are confirmed.
[0027] In the forward verification method, first judge whether formula (6) holds; if it holds, then judge whether the number of successful comparisons exceeds the product of ε and When it exceeds, the data rights can be confirmed. If it does not exceed, the rights cannot be confirmed; if formula (6) does not hold, then adjust the ε value and value through formula (7) and formula (8) until formula (6) holds;
[0028]
[0029]
[0030]
[0031] where ε is the lower limit ratio of the test fingerprint redundancy, is the upper limit average value of the byte usage difference, τ is the data storage upper limit, ρ is the clear text length, F is the fingerprint, d F is the data associated with the fingerprint, is the fingerprint redundancy.
[0032] The backward verification method specifically includes the following steps:
[0033] Step 1: First count the byte sequence of the data associated with the fingerprint;
[0034] Step 2: Divide the fingerprint in the data into η equal parts; η is the cipher length;
[0035] Step 3: Find the corresponding clear text sequence from the symbol mapping table according to the divided fingerprint;
[0036] Step 4: Count the byte sequence of the clear text obtained in Step 3;
[0037] Step 5: Compare the number of identical bytes in the byte sequences in Step 1 and Step 4 to obtain the ratio of the number of identical bytes in the entire byte sequence. When this ratio is lower than the upper limit average of byte usage differences the obtained confirmation result is false; otherwise, it is true;
[0038] Step 6: By comparing the byte sequences in Step 1 and Step 4, determine whether these two byte sequences are in a multiple relationship of 1 or -1. If so, the obtained data confirmation result is true; otherwise, it is false;
[0039] Step 7: When the confirmation results of both Step 5 and Step 6 are true, it indicates that the data has been confirmed.
[0040] The symbol mapping table maps each symbol to two different types of digital codes: one is the clear code for encoding the symbol; the other is the hidden code for carrying fingerprint information. The symbol mapping table includes a symbol set, a clear code set, and a hidden code set; each symbol corresponds to at least one clear code and at least one hidden code; the symbol set and the clear code set form the private part of the symbol mapping table, and the clear code set and the hidden code set form the public part of the symbol mapping table. The generation process of the symbol mapping table is as follows:
[0041] First, determine the clear code length ρ according to the fingerprint F and the original data d. The clear code length ρ satisfies the following formula:
[0042]
[0043] After that, determine the symbol length θ. The symbol length θ satisfies the following formula:
[0044]
[0045] In Equation (1) and Equation (2), τ represents the upper limit of data storage, λ represents the parameter affecting the randomness of the symbol length θ, and a represents the number corresponding to the first five-digit hexadecimal string in h(SK), where h(SK) is the private key hash;
[0046] According to the symbol length θ, divide the byte sequence of the original data d to obtain the symbol set S;
[0047] Then determine the hidden code length η and obtain the fingerprint redundancy See the following formulas (3) and (4);
[0048]
[0049]
[0050] Where |S| represents the length of the symbol set, γ represents the redundancy coefficient, and b represents the number corresponding to the second five-digit hexadecimal string in h(SK);
[0051] Then the fingerprints are divided into η equal parts, and the ciphertext set of the symbol mapping table can be obtained;
[0052] After that, the data randomly generated from the plaintext space serves as the plaintext, so that each plaintext corresponds to a ciphertext, thereby constructing the initial symbol mapping table;
[0053] Finally, the redundancy of the symbol mapping table is increased by filling data to form the final symbol mapping table; the redundancy χ of the symbol mapping table satisfies the following formula:
[0054] χ ≤ min(|2 8ρ - |SMT - ||, c900) (5)
[0055] where the value of c900 comes from the third five-digit hexadecimal string in h(SK), which is obtained by adding one to four characters and multiplying by the fifth character, and the maximum result does not exceed 900.
[0056] Existing data rights confirmation schemes have many problems. For example, data security and single-point failure risks caused by untrusted third parties; the inefficiency of complex encryption systems for large volumes of data; problems such as traditional digital watermarks being too dependent on data content. Given that the lack of effective data rights confirmation methods currently restricts the development of data sharing, therefore, the present invention provides a novel and practical data rights confirmation method, which meets the urgent needs of social development. The present invention uses symbol mapping coding and blockchain to achieve the purpose that users can effectively confirm the rights of data during the data sharing process, thereby realizing trusted access control and traitor tracing, and thus making the data sharing system have the capabilities of non-repudiation and anti-piracy.
[0057] The innovation of the present invention lies in directly embedding the user's fingerprint information into the data object by using symbol mapping coding, so as to achieve the effects of being independent of data content and strong access control. At the same time, by identifying these embedded fingerprints, the rights of the data can be confirmed, thereby realizing the effects of non-repudiation and anti-piracy; in addition, the present invention uses blockchain to replace the endorsement role of a trusted third party in the data sharing process, making the data sharing process not be tampered with, and thus being more convenient for trusted traceability. Brief Description of the Drawings
[0058] Figure 1 is a schematic diagram of the module in the present invention.
[0059] Figure 2 is a data rights confirmation architecture diagram in the present invention.
[0060] Figure 3 This is the access control policy diagram in the present invention.
[0061] Figure 4 This is the data structure diagram of the confirmation chain in the present invention.
[0062] Figure 5 This is the fingerprint recognition result of the method of the present invention under five types of attacks.
[0063] Figure 6 This is the performance test result of the confirmation chain in the present invention. Detailed implementation manners
[0064] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0065] The present invention first designs a symbol mapping coding technology (SMC). By dividing the byte sequence of a data object into symbols, a symbol mapping table (SMT) is generated, and the generated symbol mapping table is used to re-encode the data. The SMT maps each symbol to two different types of digital codes: one is the clear code, which is used to encode the symbol; the other is the hidden code, which is used to carry fingerprint information.
[0066] The generation process of the symbol mapping table SMT will be introduced in detail below.
[0067] First, this method needs to represent the original data in symbol form. To ensure that there is enough symbol coding space (ASCⅡ coding rule) to complete the mapping of symbols to data, it is necessary to ensure that the clear code space is not too small, and in addition, the coding space cannot be too large to prevent memory waste. Therefore, in the present invention, the clear code length ρ satisfies formula (1). In addition, when representing the original data as symbols, the corresponding relationship between the symbols and the original data needs to be considered, that is, how many original data bytes a symbol corresponds to, then the calculation formula for the symbol length θ is formula (2). According to formula (2), the symbol length can be obtained, and the symbol set S can be obtained by dividing the byte sequence of the original data.
[0068]
[0069]
[0070] In the formula, d represents the original data, F represents the fingerprint of the data stakeholder, τ represents the upper limit of data storage, λ represents the parameter affecting the randomness of the symbol length θ, which takes the value 1024 here, ρ represents the clear code length, |d| and |F| represent the corresponding lengths, a represents the number corresponding to the first five-digit hexadecimal string in h(SK), and h(SK) is the private key hash, which forms a key pair with the fingerprint F.
[0071] Then, through formula (3), the length η of the cipher can be obtained, and thus the fingerprint redundancy can be obtained (see formula (4)). Then the fingerprints are divided into η equal parts, and the cipher set H of the symbol mapping table can be obtained
[0072]
[0073]
[0074] where |S| represents the length of the symbol set, |F| represents the byte length of the fingerprint of the data stakeholder, γ represents the redundancy coefficient, and b represents the number corresponding to the second five-digit hexadecimal string in h(SK).
[0075] Next, data is sequentially taken from the symbol set as the symbols of the symbol mapping table, and data is taken from the cipher set as the ciphers of the symbol mapping table. The plaintext of the symbol mapping table is filled with randomly generated data from the plaintext space. Each plaintext corresponds to a cipher, thus constructing the initial symbol mapping table SMT
[0076] Finally, in order to keep θ confidential and make the symbol mapping table SMT more difficult to forge, the redundancy of the symbol mapping table SMT is increased. However, the increase in the redundancy of the SMT cannot be too large to avoid wasting storage space. Therefore, this method designs formula (5) to ensure the security of θ. The specific operation of increasing the redundancy of the symbol mapping table is as follows: randomly generate data from the plaintext space, then randomly select the data corresponding to the generated plaintext from the cipher set, and add it to the initial symbol mapping table to obtain the final symbol mapping table SMT with redundancy
[0077] χ ≤ min(|2 8ρ -|SMT - ||, c900) (5)
[0078] where the value of c900 comes from the third five-digit hexadecimal string in h(SK), which is obtained by adding one to four characters and multiplying by the fifth character, so the maximum result will not exceed 900. χ represents the redundancy of the symbol mapping table
[0079] In the final symbol mapping table SMT with redundancy, the number of plaintexts and ciphers is the same and they correspond one by one
[0080] Finally, the SMT includes a symbol set, a plain code set, and a cipher code set. Moreover, each symbol will correspond to at least one plain code and at least one cipher code. Suppose S represents the symbol set, P represents the plain code set, and H represents the cipher code set. Among them, the symbol set and the plain code set constitute the private part SMT+ of the symbol mapping table, and the plain code set and the cipher code set constitute the public part SMT- of the symbol mapping table. Moreover, the SMT should satisfy the following two irreversible assumptions:
[0081] δ 1 : p → s, p ∈ P, s ∈ S
[0082] δ 2 . p → h, p ∈ p, h ∈ H
[0083] After generating the SMT, use the SMT to re-encode the symbols corresponding to the data. The encoding process is to make each symbol in the SMT correspond to at least one plain code and one cipher code. The following Table 1 gives an example of an SMT. Given the string "WORLD", after encoding, the plain code corresponding to the symbol obtained from Table 1 is "0x00410xE4520xE4C10xE4F00x004D", and the hidden string it contains is "HUANG". The obtained plain code is the data associated with the fingerprint. In this way, the present invention can generate a corresponding SMT for any given data content, where the cipher code represents the fingerprint information.
[0084] Table 1 Symbol mapping table corresponding to the string "WORLD"
[0085]
[0086] As Figure 1 shown, a data rights confirmation method based on symbol mapping encoding and blockchain provided by the present invention includes the following modules: Module 1 maps the user fingerprint to the data object in combination with symbol mapping encoding, so that the data can be rights-confirmed by identifying the fingerprint in the later stage; Module 2 designs data query primitives and encapsulation algorithms to ensure the user's access control to the data, and at the same time further support the data security of the data rights confirmation model in Module 1; Module 3 combines blockchain to supervise the process of data sharing and data access, so that in the later stage, when piracy is found, the data can be traced, so as to track the pirated data and protect the rights of the data owner.
[0087] The working process of the present invention is as Figure 2 shown. The following will be introduced in detail in combination with an actual data sharing case.
[0088] Step 1: The data owner D uses the SMC technology to map the fingerprint F D , the private key hash h(SK D ) and the original data d into the SMT D , and then uses this SMTD Encode the original data to obtain the data associated with the fingerprint After that, data owner D sends the fingerprint F D , the data hash associated with the fingerprint , the data description DDes and the public part hash of SMT to the blockchain together so that data recipient U can obtain them. The above process can be formally expressed as:
[0089] SMT D ←φ 1 (d, h(SK D ), F D , γ, τ, λ)
[0090]
[0091] where λ represents the parameter affecting the randomness of the symbol length θ, τ represents the upper limit of data storage, and γ represents the redundancy coefficient.
[0092] Step 2: Data owner D sends the data associated with the fingerprint to data recipient U through an off-chain channel.
[0093] Step 3: Data recipient U obtains from the blockchain and compares its hash with the data obtained off-chain to prevent data tampering.
[0094] Step 4: When data recipient U applies to data owner D to query data, he needs to first use the SMC technology to map his own fingerprint F U , the private key hash h(SK U ) and the data associated with the fingerprint of the data owner into his own SMT U (in this process, the data associated with the fingerprint of the data owner is regarded as the original data), and encodes the received data U using this SMT to obtain the data associated with his own fingerprint Then send and the query request q to data owner D through the blockchain. The above process can be formally expressed as:
[0095]
[0096]
[0097] Step 6: Data recipient U sends to data owner D through an off-chain channel. The public part in the symbol mapping table generated for the data recipient.
[0098] Step 6: When the data owner D obtains the query request q sent by the data recipient U from the blockchain, it should combine the data encapsulation algorithm to obtain a data decoder that only contains the part applied by the data recipient, and implement it through the smart contract running in the TEE environment. At the same time, record this process in the blockchain. The above process can be formally expressed as:
[0099]
[0100]
[0101] The access control in this part is as Figure 3 shown on the data owner side. When the data owner D receives the query request q, it first verifies whether the query request meets the requirements of the data description DDes and the access control. When the verification passes, it should combine the query request q and the private part of the symbol mapping table it generates to generate the data needed for querying for the data recipient. and the F U Generate a data decoder for the data recipient to query data and send it to the data recipient. This process is carried out through the on-chain channel.
[0102] Step 7: The data owner D sends the above data decoder to the data recipient U through the blockchain.
[0103] Step 8: The data recipient U obtains the data decoder from the blockchain. And use this data decoder to restore the applied data, so as to obtain the applied data. The above process can be formally expressed as:
[0104]
[0105] The data acquisition in this part is as Figure 3 shown on the data recipient side. When the data recipient receives the data decoder, it needs to input the data associated with its own fingerprint and the private part of its own symbol mapping table so that the data decoder can judge whether it has the permission to query data. When the verification passes, the data decoder will combine
[0106] Step 9: When the data owner discovers data in the network, they can apply to the arbiter for rights protection. The arbiter determines the data right holder by identifying the fingerprint in the data and judges the source of the data. The above process can be formally expressed as:
[0107]
[0108]
[0109] ε is the lower limit ratio for testing fingerprint redundancy, is the upper limit average of byte usage differences, Tester U is to detect whether the fingerprint of the data receiver is contained in the data, Tester D is to detect whether the fingerprint of the data owner is contained in the data.
[0110] The above φ 1 -φ 5 represents the formalization process and has no exact meaning.
[0111] Among them, for the fingerprint recognition method, the present invention adopts two methods.
[0112] First, adopt the method of forward verification. By statistically combining the characters in the cipher part of SMT - and cropping and dividing the characters according to the fingerprint length, and then comparing the divided fields with the fingerprint one by one using the Pearson correlation coefficient, and counting the number of successful comparisons. When it exceeds a certain proportion, it is considered a successful comparison and its rights can be confirmed.
[0113] Regarding the problem of how to determine the value of the above "certain proportion", in the present invention, considering better extraction of user fingerprints, the present invention does not confirm rights by comparing whether the number of successful comparisons exceeds a certain threshold. Instead, it first obtains the eligible ε value and value by judging whether formula (6) holds. If formula (6) is not satisfied, then adjust the ε value and value through formula (7) and formula (8). Then when the eligible ε value is obtained, it only needs to judge whether the number of successful comparisons exceeds the product of ε and to confirm the data rights. The advantage of doing this is that it can dynamically adjust the probability of the fingerprint being detected according to the size of the data with the embedded fingerprint, so as to achieve a better data right confirmation effect.
[0114]
[0115]
[0116]
[0117] where ε is the lower limit ratio for testing fingerprint redundancy, is the upper limit average of byte usage differences, τ is the upper limit of data storage, and D is the length of the plain code.
[0118] Secondly, in order to prevent the situation where the forward verification method may not work well, the present invention also designs a backward verification method. Using these two methods to identify fingerprints in data objects has a good effect. The backward verification method mainly proceeds according to the following steps:
[0119] Step 1: First, count the data d F of the associated fingerprint;
[0120] Step 2: Divide the fingerprints that may be embedded in the (pirated) data into η equal parts;
[0121] Step 3: According to the divided fingerprints, find the corresponding plain code sequences from the SMT;
[0122] Step 4: Count the byte sequences of the plain codes obtained in Step 3;
[0123] Step 5: Compare the number of identical bytes in the byte sequences in Step 1 and Step 4 to obtain the proportion of the number of identical bytes in the entire byte sequence. When this proportion is lower than the upper limit average of byte usage differences it can be obtained that the right confirmation result is false; otherwise it is true;
[0124] Step 6: By comparing the byte sequences in Step 1 and Step 4, determine whether these two byte sequences are in a multiple (1 or -1) relationship. If so, the data right confirmation result is true; otherwise it is false;
[0125] Step 7: When the right confirmation results of Step 5 and Step 6 are both true, this method believes that the rights of the data can be confirmed.
[0126] As Figure 4 shown, this figure describes the data structure of the right confirmation chain. In this structure, Transaction 1 records the data information that the data owner D needs to send to the data receiver U, including F D , DDes, and In Transaction 2, some data information sent by the data receiver U to the data owner D is recorded, including q and and other information. In Transaction 3, what needs to be recorded is the data decoder So that the data receiver U can obtain the data. These transactions are connected together in chronological order to form a chain of rights confirmation. It is precisely because of the chain of rights confirmation that people can view the data destination and data source at any time and place, thus achieving data traceability.
[0127] Currently, the technical methods for data rights confirmation at home and abroad mainly involve encrypting data in stages and embedding watermarks in the data. Table 2 below compares the method of the present invention with the existing methods in terms of different technical indicators. It is easy to find that the method of the present invention is significantly superior to the existing methods in various technical indicators.
[0128] Table 2 Comparison of the method of the present invention and the existing methods in different technical indicators
[0129]
[0130] At the same time, the technical indicators of the present invention itself have also achieved the expected good results, as described below:
[0131] 1. By combining symbol mapping coding, the fingerprints of data stakeholders are encoded into the byte sequence of shared data, so that the data rights statement can be independent of the data content, achieving data rights confirmation independent of the data content; the data sets in Table 3 below mainly come from UCI-MLR and have different data types and volumes. By embedding fingerprints in data of different data types and then conducting rights confirmation on the fingerprinted data, it can be found that the accuracy rate of detecting fingerprints in non-tampered data objects can reach 100%.
[0132] Table 3 Recognition results after embedding fingerprints in data of different data types
[0133]
[0134] 2. No matter what kind of illegal processing the data receiver conducts on the fingerprinted data, the method of the present invention can perform fingerprint recognition on it. The present invention mainly measures whether the fingerprints in the data object can be correctly recognized when there are potential malicious users from five types of simulated attacks in Table 4. First, use the simulated data tampering attack in Table 4 to process the data associated with the user's fingerprint, and then import the method of the present invention to extract fingerprints from the tampered data, so as to judge whether the fingerprint extraction method can effectively extract fingerprints from the tampered data. From Figure 5 the experimental results, it can be found that the method of the present invention still has good recognition efficiency for the fingerprints in the data object under five main types of expression attacks, thus can better achieve traitor tracing and achieve the effects of anti-repudiation and anti-piracy.
[0135] Table 4 Explanation of simulated attacks
[0136]
[0137] 3. In the present invention, a rights confirmation chain (a blockchain underlying layer) is designed to replace the third party in the traditional rights confirmation scheme, realizing the detailed record of the data sharing process, which is the underlying basis for performing data rights confirmation operations. Due to the introduction of the rights confirmation chain, data sharing becomes more secure and reliable. In the present invention, the average latency and throughput performance of the rights confirmation chain are mainly simulated and tested. Among them, the average latency mainly represents the confirmation speed of a single transaction, while the throughput reflects the number of transactions completed per unit time. First, the blockchain network is implemented using the SpringBoot framework, and a simulation experiment is carried out using three servers and a PC; then, the written program is run on the servers and the PC, and these nodes are used for transaction consensus; finally, the efficiency of the blockchain network is reflected by recording the average latency of these transaction confirmations and the throughput of the blockchain. The specific test cases are shown in Table 5, where the number of nodes represents the number of users added to the blockchain network, and the number of transactions represents the number of transactions simulated in the experiment. Figure 6 The performance test results of the rights confirmation chain are shown, and it can be seen that when a large number of transactions are concurrent, the system's latency and throughput both reach the expected good results in practical applications.
[0138] Table 5 Blockchain Test Cases
[0139] Parameter Parameter value Number of nodes 3、4、5 Number of transactions 5K, 10K, 15K, 20K, 25K Average transaction size 1Kb
Claims
1. A data right confirmation method based on symbol mapping coding and blockchain, characterized in that, it includes the following steps: a. The data owner uses the SMC technology to map the fingerprint F D , the private key hash h(SK D ), and the original data d into a symbol mapping table SMT D . Then, the original data is encoded using the symbol mapping table SMT D to obtain the data associated with the fingerprint . After that, the data owner publishes the fingerprint F D , the hash of the data associated with the fingerprint , the data description DDes, and the hash of the public part of the symbol mapping table together to the blockchain; b. The data owner sends the data associated with the fingerprint to the data recipient through an off-chain channel ; c. The data recipient obtains from the blockchain and compares its hash with the data obtained from the off-chain channel for a hash comparison; d. When the data receiver applies to the data owner for querying data, the data receiver first uses the SMC technology to map its own fingerprint F U , the private key hash h(SK U ) and the data associated with the data owner's fingerprint received together into its own symbol mapping table SMT U , and encodes the received data U through the symbol mapping table SMT to obtain the data associated with its own fingerprint Then, it sends the hash of the public part of the symbol mapping table and the query request q to the data owner through the blockchain; e. The data recipient sends the public part of the symbol mapping table in step d to the data owner through an off-chain channel ; f. After the data owner obtains the query request q sent by the data recipient from the blockchain, it first verifies whether the query request q meets the requirements of the data description and access control; if so, it combines the private part of the symbol mapping table generated by itself and the query request q to generate Then combine and the F U Generate a data decoder required for the data recipient to query data; if not, return null; g. The data owner sends the data decoder to the data receiver through the blockchain; h. The data receiver obtains the data decoder and restores the applied data through the data decoder, thereby obtaining the applied data.
2. The data right confirmation method based on symbol mapping coding and blockchain according to claim 1, characterized in that, In step h, after the data receiver receives the data decoder, it first inputs the data associated with its own fingerprint and the private part of its own symbol mapping table so that the data decoder can verify whether it has the permission to query the data; when the verification passes, the data decoder returns the data requested by the data receiver, otherwise the data decoder does not return any data.
3. The data right confirmation method based on symbol mapping coding and blockchain according to claim 1, characterized in that, it further includes the following steps: i. When the data owner discovers data in the network, the data owner applies to the arbitrator for rights protection, and the arbitrator determines the data right holder by identifying the fingerprint in the data.
4. The data right confirmation method based on symbol mapping coding and blockchain according to claim 3, characterized in that, There are two methods for identifying fingerprints, as follows: First, adopt the method of forward verification. By counting the characters in the cipher part of SMT - and cropping and dividing the characters according to the fingerprint length, compare these divided fields with the fingerprint one by one, count the number of successful comparisons, and then confirm its rights; Secondly, the method of backward verification is adopted. It first counts the byte sequence in the data, then reversely estimates the clear code byte sequence corresponding to the fingerprint in the data, compares these two byte sequences, and then confirms its rights.
5. The data right confirmation method based on symbol mapping coding and blockchain according to claim 4, characterized in that, In the forward verification method, first, it is judged whether formula (6) holds; if it holds, it is judged whether the number of successful comparisons exceeds the product of ε and When it exceeds, the data right can be confirmed; if formula (6) does not hold, the ε value and value are adjusted through formula (7) and formula (8) until formula (6) holds; where ε is the lower limit ratio for testing fingerprint redundancy, is the upper limit average of byte usage differences, τ is the data storage upper limit, ρ is the clear code length, F is the fingerprint, d F is the data associated with the fingerprint, is the fingerprint redundancy.
6. The data right confirmation method based on symbol mapping coding and blockchain according to claim 4, characterized in that, The backward verification method specifically includes the following steps: Step 1. First count the byte sequence of the data associated with the fingerprint; Step 2. Divide the fingerprint in the data into η equal parts; η is the length of the ciphertext; Step 3. Find the corresponding clear code sequence from the symbol mapping table according to the divided fingerprint; Step 4. Count the byte sequence of the clear code obtained in Step 3; Step 5. Compare the number of identical bytes in the byte sequences in Step 1 and Step 4 to obtain the ratio of the number of identical bytes in the entire byte sequence. When this ratio is lower than the upper limit average of byte usage differences the obtained result of right confirmation is false; otherwise it is true; Step 6. By comparing the byte sequences in Step 1 and Step 4, judge whether these two byte sequences are in a multiple relationship of 1 or -1. If so, the data right confirmation result is true; otherwise, it is false; Step 7. When the confirmation results of Step 5 and Step 6 are both true, it means that the data has been confirmed.
7. The data right confirmation method based on symbol mapping coding and blockchain according to claim 1, characterized in that, The symbol mapping table maps each symbol into two different types of digital codes: one is the clear code, which is used to encode the symbol; the other is the ciphertext, which is used to carry fingerprint information.
8. The data right confirmation method based on symbol mapping coding and blockchain according to claim 7, characterized in that, The symbol mapping table includes a symbol set, a clear code set and a ciphertext set; each symbol corresponds to at least one clear code and at least one ciphertext; the symbol set and the clear code set form the private part of the symbol mapping table, and the clear code set and the ciphertext set form the public part of the symbol mapping table.
9. The data right confirmation method based on symbol mapping coding and blockchain according to claim 8, characterized in that, The generation process of the symbol mapping table is as follows: First, determine the clear code length ρ according to the fingerprint F and the original data d. The clear code length ρ satisfies the following formula: After that, determine the symbol length θ. The symbol length θ satisfies the following formula: In Formula (1) and Formula (2), τ represents the upper limit of data storage, λ represents the parameter affecting the randomness of the symbol length θ, and a represents the number corresponding to the first five-digit hexadecimal string in h(SK); According to the symbol length θ, the byte sequence of the original data d is divided to obtain the symbol set S; Next, determine the length η of the password and obtain the fingerprint redundancy See the following formulas (3) and (4); where |S| represents the length of the symbol set, γ represents the redundancy coefficient, and b represents the number corresponding to the second five-digit hexadecimal string in h(SK); Then divide fingerprints into η equal parts to obtain the cipher set of the symbol mapping table; After that, the data randomly generated from the plaintext space serves as the plaintext, and each plaintext corresponds to a ciphertext, thereby forming an initial symbol mapping table; Finally, the redundancy of the symbol mapping table is increased by padding data to form the final symbol mapping table; the redundancy χ of the symbol mapping table satisfies the following formula: χ ≤ min(|2 8ρ - |SMT - ||, c900) (5) Among them, the value of c900 comes from the third five-digit hexadecimal string in h(SK), which is obtained by adding one to four characters and multiplying by the fifth character, and the maximum result does not exceed 900.