A privacy-enhanced semi-blind digital fingerprinting method for data privacy protection and a detection method thereof

By combining differential privacy technology and digital fingerprinting technology, and embedding noise into specific locations in numerical datasets, the problem of privacy protection and traitor tracking in text-based data carriers is solved, achieving efficient privacy protection and traitor tracking functions, suitable for high-speed data interaction environments.

CN115828194BActive Publication Date: 2026-06-23TIBET UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIBET UNIVERSITY FOR NATIONALITIES
Filing Date
2022-11-21
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing digital fingerprint technology cannot simultaneously protect the privacy of text-based data carriers and track traitors, and traditional solutions are inefficient and cannot adapt to high-speed, real-time data interaction environments.

Method used

By combining differential privacy technology and digital fingerprinting technology, noise is embedded into specific locations in a numerical aggregated dataset. A fingerprint coordinate matrix is ​​generated through a permutation encryption algorithm, achieving a balance between privacy protection and traitor tracking. Random noise blocks are generated using a Gaussian mechanism and embedded into the dataset. Hash calculation and Gaussian fitting are combined to detect illegally spread user information.

Benefits of technology

It achieves a unified approach to protecting the privacy of carrier data and tracking traitors, ensuring the security of data privacy and the effectiveness of tracking. It can resist new types of attacks, is suitable for text-based data carriers, and simplifies computation while meeting different privacy protection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115828194B_ABST
    Figure CN115828194B_ABST
Patent Text Reader

Abstract

The application discloses a data privacy protection method and a detection method for a privacy-enhanced semi-blind digital fingerprint, combines a differential privacy technology and a digital fingerprint technology, embeds carefully designed noise into specific positions of a numerical aggregate data set, can realize fingerprint embedding and noise interference operation on carrier data in one step, and ensures protection of private information in the carrier data set and tracking function on traitors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data privacy protection technology in information security, and specifically relates to a privacy-enhanced semi-blind digital fingerprint carrier data privacy protection method. Background Technology

[0002] Access control and encryption algorithms are considered primary methods for preventing unauthorized access to data during transmission or interaction. However, preventing authorized users from illegally disseminating private data after receiving and decrypting it is also a significant concern. Increasing research recognizes the importance of traitor tracking, with digital fingerprinting technology—embedding user-related fingerprint information into original digital products to trace data origins—playing a crucial role in this field. Embedding user fingerprint information into image, video, and other data carriers is a primary implementation of digital fingerprinting. However, research on digital fingerprinting schemes often neglects the privacy protection of the carrier data, a situation clearly insufficient in today's complex and ever-changing network interaction environment.

[0003] Current research on digital fingerprinting that addresses the privacy of carrier data mostly treats privacy protection and rebellious tracking as two separate studies. Even within the same system model, they are rarely fully integrated. Typically, the approach involves adding a tracking module after implementing data privacy protection functions, such as using data sanitization mechanisms or k-anonymization techniques to protect the privacy of the carrier data, and then using digital fingerprinting technology to track the data or related users.

[0004] Solving multiple corresponding problems sequentially by layering multiple technologies lacks innovation and reduces the efficiency of solution implementation. In today's high-speed, real-time data interaction environment, the approach of implementing different technologies step-by-step and in stages to achieve multiple goals is no longer applicable. Furthermore, although traditional digital fingerprinting technology has been further developed with the development and popularization of big data, text-based digital fingerprinting schemes remain scarce. The fundamental reason is that text-based data contains too little redundant information, which is not conducive to fingerprint embedding. Most text-based data fingerprinting schemes meet basic requirements by annotating attribute or tuple values ​​differently. However, they usually require the layering of complex encryption algorithms, which greatly reduces usability and simplicity. In summary, designing an algorithm based on digital fingerprinting technology that can simultaneously protect the privacy of text-based data carriers and track traitors is a gap in existing research. Summary of the Invention

[0005] Purpose of the invention: To address the problem that existing digital fingerprint algorithms cannot simultaneously protect the privacy of text-based data carriers and track traitors, as well as the lack of text-based digital fingerprint recognition schemes, this invention proposes a carrier data privacy protection method based on traceable privacy-enhanced semi-blind digital fingerprints and a method for detecting the identifying information of relevant users in illegally transmitted data sets. Using large-scale text-based numerical aggregate data as the fingerprint carrier, it achieves privacy protection for large-scale aggregate data while tracking illegally transmitted data objects.

[0006] Technical solution: A data privacy protection method based on traceable, privacy-enhanced semi-blind digital fingerprints, comprising:

[0007] Receive privacy protection requests from data requesters and obtain the data set and user fingerprint information of the data subject to privacy protection;

[0008] Divide the carrier dataset D into appropriate k×k data blocks to obtain the aggregated data block set D. k×k ;

[0009] Encode the user's fingerprint information into a fingerprint coordinate matrix S. * Then, a permutation encryption algorithm is used to obtain the fingerprint coordinate matrix S. * Permutation encryption is performed to obtain a new fingerprint coordinate matrix S. * ;

[0010] The differential privacy sensitivity of the carrier dataset D is calculated, and the differential privacy sensitivity is combined with the differential privacy parameters ε and δ to calculate the range of values ​​for the variance σ of the differential privacy Gaussian mechanism. The range of values ​​for the variance σ of the differential privacy Gaussian mechanism is divided into k equal parts and placed into matrix V according to their magnitude to obtain matrix V. 1×k , matrix V 1×k Each scalar in the matrix is ​​substituted into the differential privacy Gaussian mechanism to generate a random noise block matrix P;

[0011] Extract the new fingerprint coordinate matrix S by column. * Each column has two values ​​used to locate the aggregated data block set D. k×k At the locations where random noise blocks need to be embedded, embed random noise blocks at the corresponding locations to obtain a noisy dataset. The noisy dataset The elements in the data are data stored in a privacy-protected manner.

[0012] Noisy dataset Feedback is sent to the data requester.

[0013] Furthermore, the aforementioned carrier dataset D consists of text-based numerical aggregated data.

[0014] Furthermore, the process of encoding user fingerprint information into a fingerprint coordinate matrix S... * Specifically, it includes:

[0015] Convert the user fingerprint information s in string form into a decimal matrix (S). 10 The form is then used to construct a decimal matrix (S). 10 Convert the form to binary matrix (S)2 form;

[0016] Based on the aggregated data block set D k×k The number of data blocks in the matrix is ​​used to convert the binary matrix (S)² form into a matrix (S) with the corresponding radix k. k form;

[0017] Matrix (S) k Transformed into a 2xm fingerprint coordinate matrix S * .

[0018] Furthermore, the aforementioned permutation encryption algorithm is used to obtain the fingerprint coordinate matrix S. * The process of performing substitution encryption includes:

[0019] The fingerprint coordinate matrix S obtained through the permutation matrix R * Permutation encryption is performed; the permutation matrix R is calculated based on the sequence of each letter in the custom permutation key κ1.

[0020] Furthermore, the differential privacy sensitivity of the computational carrier dataset D includes:

[0021] The differential privacy sensitivity Δ2f of the carrier dataset D is calculated according to equation (1):

[0022] Δ2f=max D,D′ ||f(D)-f(D′)||2 (1)

[0023] In the formula, f is a query function, D represents the carrier dataset, and D′ represents the sibling dataset of the carrier dataset D; D and D′ differ by only one data point.

[0024] The method of combining differential privacy sensitivity with differential privacy parameters ε and δ to calculate the range of values ​​for the variance σ of the differential privacy Gaussian mechanism includes:

[0025] According to equation (2), the range of values ​​for the variance σ of the differential privacy Gaussian mechanism is calculated as follows:

[0026]

[0027] In the formula, Δ2f represents the differential privacy sensitivity, and ε and δ are both differential privacy parameters.

[0028] This invention also discloses a method for detecting the identifying information of relevant users in illegally disseminated datasets, comprising the following steps:

[0029] According to the data privacy protection method, the carrier dataset is protected for privacy to obtain the corresponding noisy dataset; the data privacy protection method is a data privacy protection method based on traceable privacy-enhanced semi-blind digital fingerprint;

[0030] For the aggregated data block set D k×k Each data block in the dataset is hashed to obtain a hash matrix H;

[0031] Obtain the carrier dataset D to be detected * The carrier dataset D to be detected * Divide the dataset into k×k data blocks using the same block division method as the carrier dataset D, resulting in the aggregated data block set {D}. *} k×k ; Calculate the aggregated data block set {D *} k×k The hash value of each data block is used to obtain the corresponding hash matrix H. * ;

[0032] By comparing hash matrix H and hash matrix H * Record the coordinates of the two hash matrix values ​​that are different, and obtain a 2x1m coordinate matrix S. * ;

[0033] Based on coordinate matrix S * Locate and extract the aggregated data block set {D} from the coordinate points in the dataset. *} k×k For the corresponding data blocks, the variance σ of these data blocks is calculated by Gaussian fitting, and the variance σ calculated by Gaussian fitting is stored in matrix M;

[0034] The coordinate matrix S * Merge matrix M row by row to form matrix U, and then rearrange matrix U column by column according to the size of the elements in matrix M to obtain a new matrix U;

[0035] Extract the first 2 rows and the first k columns of the new matrix U as a k-ary noise fingerprint matrix (S). * ) k ; The noise fingerprint matrix (S * ) k Noisy fingerprint data converted to decimal form;

[0036] The identification information string of the relevant user is obtained from the noisy fingerprint data in decimal form by the permutation decryption algorithm. This identification information string is the user fingerprint information in string form.

[0037] Furthermore, by recording the coordinates of the two hash matrix values ​​that are different, a coordinate matrix S with 2 rows and m columns is obtained. * ,include:

[0038] Record the coordinates of the two hash matrix values ​​that are different, and record the x-coordinates of each column into the coordinate matrix S. * The first row, with its ordinate, is recorded in the coordinate matrix S. * The second row ultimately yields a 2x1m coordinate matrix S. * .

[0039] Furthermore, the calculation of the variance σ of these data blocks through Gaussian fitting includes:

[0040] The variance σ of these data blocks is calculated using the following formula:

[0041]

[0042] In the formula, fitGauss represents the Gaussian fitting function. Represents the aggregated data block set {D *} k×k The data block in row r and column c.

[0043] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0044] (1) This invention combines differential privacy technology and digital fingerprint technology, embedding carefully designed noise into specific locations in numerical aggregate datasets, enabling fingerprint embedding and noise interference operations on carrier data in one step, while ensuring the protection of privacy information in the carrier dataset and the tracking function of traitors.

[0045] (2) In terms of protecting the privacy of carrier data, the present invention can flexibly meet different privacy protection needs. In particular, it can achieve high-precision statistical availability for numerical carrier data.

[0046] (3) In terms of traitor tracking, the present invention can meet the basic requirements of digital fingerprints in terms of imperceptibility, robustness and trustworthiness, and can resist conspiracy attacks, as well as new attacks such as buyer framing, identity authentication and man-in-the-middle attacks.

[0047] (4) This invention is applicable to text-based data carriers as objects for fingerprint embedding, and can achieve privacy protection and tracking functions for carrier data on the basis of simplified calculation. Attached Figure Description

[0048] Figure 1 This is a system model diagram of the present invention;

[0049] Figure 2 This is a system security model diagram of the present invention;

[0050] Figure 3 This is a schematic diagram of the digital fingerprint generation and noise embedding process of the present invention;

[0051] Figure 4 This is a flowchart illustrating the simulation implementation of the digital fingerprint generation and noise embedding process of the present invention.

[0052] Figure 5 This is a schematic diagram of the semi-blind fingerprint detection process of the present invention;

[0053] Figure 6 This is a simulation flowchart of the semi-blind fingerprint detection process of the present invention;

[0054] Figure 7 This is a statistical probability comparison chart of the total original dataset and the total noisy dataset in the simulation experiment of this invention;

[0055] Figure 8 This is a statistical probability comparison chart of 8 groups of labeled original data blocks and noisy data blocks in the simulation experiment of this invention;

[0056] Figure 9 This is a box plot comparing the key values ​​of the eight sets of data blocks in the simulation experiment of this invention after noise processing with the corresponding original data blocks.

[0057] Figure 10 This is a graph showing the fingerprint recognition rate (CRR) after randomly deleting different amounts of data in the simulation experiment of this invention.

[0058] Figure 11 This is a comparison chart of key parameters in the simulation experiment of this invention by randomly deleting different amounts of data;

[0059] Figure 12 This is a graph showing the CRR relationship of fingerprint recognition accuracy after randomly adding different amounts of data in the simulation experiment of this invention.

[0060] Figure 13 A comparison chart of key parameters with randomly added data in the simulation experiment of this invention. Detailed Implementation

[0061] The technical solution of the present invention will now be further described in conjunction with the accompanying drawings and embodiments.

[0062] The method of this invention comprises two basic processes: a digital fingerprint generation and noise embedding process, and a semi-blind fingerprint detection process. The digital fingerprint generation and noise embedding process is used to generate fingerprint-encoded data specific to a particular user and to embed fingerprint noise into the original dataset; the semi-blind fingerprint detection process is used to detect user-related identifying information in the captured illegally distributed dataset.

[0063] Example 1:

[0064] This embodiment discloses a privacy-enhanced semi-blind digital fingerprint data privacy protection method, the implementation steps of which include:

[0065] Step 1: Receive a privacy protection request from the data requester and obtain the dataset containing the data to be protected and the user's fingerprint information; encode the user's fingerprint information into a fingerprint coordinate matrix S. * Specifically, it includes:

[0066] Convert the user fingerprint information s in string form into a decimal matrix (S). 10 This form is then further transformed into a binary matrix (S)2, a process represented as s→(S). 10 →(S)2. Divide the original dataset D into appropriate k×k data blocks to obtain the aggregated data block set D. k×k According to the aggregated data block set D k×k The number of data blocks in the matrix is ​​used to transform (S)2 into a matrix (S) with the corresponding base k. k The form can be represented as (S)2→(S). k Finally, the matrix (S) k Transform into a 2xm matrix S * S * This is the required fingerprint coordinate matrix, and this process is represented as (S). k =(S k ) 1n →(S k ) 2m =S * .

[0067] Step 2: To ensure stronger privacy, the fingerprint coordinate matrix S obtained in Step 1 is modified by introducing a permutation matrix R. * Perform a permutation encryption operation to obtain a new fingerprint coordinate matrix S. * The permutation matrix R is calculated based on the sequence of each letter in the permutation key κ1, and can be represented as follows: For example, if the value of κ1 is "MAKE", then R = [4, 1, 3, 2]. Using the permutation matrix R, the fingerprint coordinate matrix S can be... * Replace with a new fingerprint coordinate matrix S * .

[0068] Step 3: Based on the privacy protection requirements of the dataset, calculate the differential privacy sensitivity of the original dataset D. Where f is a query function, and the formula for calculating differential privacy sensitivity is:

[0069] Δ2f=max D,D′ ||f(D)-f(D′)||2 (1)

[0070] Then, combining the known differential privacy parameters ε and δ, the range of values ​​for the variance σ of the differential privacy Gaussian mechanism can be calculated using the following formula:

[0071]

[0072] The range of variance σ is then divided into k equal parts, and the values ​​are arranged in order of magnitude and placed into matrix V, resulting in matrix V. 1×k Finally, matrix V 1×k Each scalar in the equation is substituted sequentially into the differential privacy Gaussian mechanism N(0, σ). 2 A random noise block matrix P is generated in the process.

[0073] Step 4: Based on the fingerprint coordinate matrix S generated in Step 2 * The original dataset D of the positioning carrier needs to be embedded with noise locations. Random noise blocks generated in step 3 are sequentially embedded at the corresponding locations. Specific operations include:

[0074] Extract the fingerprint coordinate matrix S generated in step 2 by column. * Each column has two values ​​used to locate the aggregated data block set D. k×k The coordinates of the random noise blocks to be embedded are required. Random noise blocks are extracted column-wise from the random noise block matrix P and then superimposed as additive noise onto the aggregated data block set D. k×k The corresponding data blocks are then used to obtain the traceable noisy dataset (TND).

[0075] Noisy dataset Feedback is sent to the data requester.

[0076] Example 2:

[0077] This embodiment discloses a method for detecting the identifying information of relevant users in an illegally disseminated dataset, including the following steps:

[0078] To achieve semi-blind fingerprint detection, the aggregated data block set D is calculated before fingerprint noise embedding. k×k The hash value. Extract the aggregated data block set D sequentially. k×k For each data block, the MD5 algorithm is used to calculate a 128-bit hash value for each data block, and these hash values ​​are stored row-by-row in a hash matrix H. The hash matrix H is published to provide a reference for potential semi-blind testing by a third-party arbitration body.

[0079] The vector dataset D to be detected * Divide the data into k×k blocks using the same partitioning method to obtain the aggregated data block set {D}. *} k×k Calculate {D} sequentially *} k×k The hash value of each data block is calculated and these hash values ​​are stored row-by-row in the hash matrix H. * middle.

[0080] Compare the two hash matrices H and H row by row. * Record the coordinates of the two matrix points with different values, and record the x-coordinates of each matrix into the coordinate matrix S. * The first row, with its ordinate, is recorded in the coordinate matrix S. * The second row ultimately yields a 2x1m coordinate matrix S. * .

[0081] Based on the calculated coordinate points, the aggregated data block set {D} is located and extracted sequentially. *} k×k The variance σ of the corresponding data blocks is calculated using Gaussian fitting, and can be specifically expressed as:

[0082]

[0083] The variance σ obtained from the fitting calculation is then stored column-wise in matrix M.

[0084] The coordinate matrix S * The matrix M is merged row-wise into a new matrix U, which can be represented as follows:

[0085] U = [S] * M]′;

[0086] Then rearrange matrix U according to the size of the elements in matrix M by column to obtain a new matrix U.

[0087] Extracting the first two rows and the first k columns of the new matrix U obtained in step 9, which are then used as the k-ary noise fingerprint matrix (S) * ) k Specifically, it can be expressed as:

[0088] S * =U 1:2,1:end S * =(s k ) 2m →(s k ) 1n =(S) k ;

[0089] Then convert it to a decimal number, which can be represented as (S). k→(S)2,(S)2→(S) 10 Finally, the identification information string s of the relevant user is obtained through the substitution decryption algorithm.

[0090] Example 3:

[0091] like Figure 1 As shown, the method in this embodiment consists of two processes: digital fingerprint generation and noise embedding, and semi-blind fingerprint detection. Figure 1 This is a system model diagram of the method. To maximize the effectiveness of the method, placing it in a suitable environment is crucial. The method in this embodiment is applicable to multi-user data interaction modes. In the diagram, the black solid line and circular numerical markers represent the traceable noisy dataset published by this method. The generation process of the corresponding hash matrix H is illustrated by the dashed lines and square numerical markers, which represent the semi-blind fingerprint detection process. The entire method model comprises three types of entities: data requesters, data service providers, and third-party arbitration institutions. Data requesters need to obtain specially customized datasets for querying, analysis, or prediction. In this embodiment, data requesters are specifically required to send identifying information such as User ID (UID), time, and device physical address (MAC address) along with their data request. Furthermore, data requesters may also be malicious users illegally disseminating confidential information. Data service providers generally have two functions: storing and classifying all collected data, and providing data computation services. In this embodiment, the computation service needs to include generating fingerprints and embedding them into the original dataset, calculating hash values, and generating traceable noisy datasets (TNDs). In the event of a dispute, a third-party arbitration institution performs identifier extraction calculations and arbitration operations on the illegitimate data. It holds the hash results of the original dataset and calculates the hash results of the dataset under test, comparing the differences to accurately and securely calculate the corresponding coordinate fingerprint matrix, thus identifying malicious or related users.

[0092] The security model aims to ensure that each subject in the method has the most secure yet absolutely usable access to the object's data. For example... Figure 2As shown in the diagram, the arrows represent the response and constraint relationships between entities. The main objects of this security model consist of the data requester, the data service provider, and the third-party arbitration institution. The object objects include the original dataset, the corresponding noisy fingerprint dataset, the original hash matrix, and the noisy hash matrix. Only the data service provider can possess and access the original dataset. The data requester receives a noisy dataset that has undergone privacy protection processes such as encryption, replacement, and differential privacy, and this noisy dataset contains the data requester's identifiable information. The third-party arbitration institution can only access the hash matrices of the original dataset and the dataset to be inspected. Because it cannot access the original dataset, it essentially prevents vulnerabilities caused by untrusted third parties.

[0093] The core of this embodiment's method is to achieve privacy protection by embedding a carefully designed noise set into specific locations within the original aggregated dataset. This noise set satisfies differential privacy, and the data requester's identifying information is transformed into specific coordinates embedded in the noise within the original dataset. Different data requesters possess different identifying information, which may be a UID, MAC address, or a combination thereof. The original dataset generates different copies by embedding the designed random noise set at different coordinates. These coordinates are generated from fingerprints encoded by the data requester's seed. These copies are unique throughout the data interaction process and are referred to as TNDs in this embodiment.

[0094] To facilitate simulation experiments and explanations, all data involved in this method, including the original aggregated dataset and the noisy dataset, are represented in matrix form. For example... Figure 3 As shown, the original aggregated dataset is divided into a set of data blocks in the shape of a square matrix. The diagram and subsequent detailed process descriptions will use an 8×8 format (i.e., k=8) as an example to illustrate the dataset partitioning. Based on the size of the dataset and the strength of privacy protection requirements, the differential privacy sensitivity of the dataset is calculated. The differential privacy calculation parameters ε and δ are given. Based on the formula... Calculate the range of variance σ required for the differential privacy Gaussian mechanism. Divide the variance σ equally into k sub-variances σ of different sizes. i 2 Based on the differential privacy Gaussian mechanism N(0, σ) i 2 Generate random noise blocks P that satisfy different Gaussian distributions. i Each noise block P i The amount of noise in the dataset should be equal to the size of a block of data in the original dataset. Since the original aggregated dataset D is divided into 8×8 blocks, assuming the size of the original dataset is represented by r, the number of data blocks in each block is r / 64. Therefore, the random noise block P... iThe noise level in the noise block is also r / 64. (The last part, "Noise block P," appears to be a separate, unrelated sentence fragment and is left untranslated.) i TNDs can be obtained by sequentially inserting them into specific positions in the original data block. The coordinates of the noise block embedding are compiled from user identification information. The user identification information is first converted into a k-ary one-dimensional matrix (k=8), and then transformed into a 2-row, m-column two-dimensional matrix S through matrix transformation. * Each column in the matrix identifies a specific coordinate point; the first row represents the x-coordinate of the coordinate point, and the second column represents the y-coordinate. To address potential fingerprint detection, the original dataset D is partitioned into blocks using the MD5 algorithm. k×k The hash value is calculated, and the corresponding 128-bit hash value of each piece of data is stored in matrix H row by row.

[0095] The method described in this embodiment can be implemented using a compilable programming language. Figure 4 The simulation demonstrates the fingerprint generation and noise embedding process. The function `embed_GetBinseed()` retrieves the current time and user identification information and stores it in binary form into a one-dimensional matrix. Subsequently, the function `embed_GetCoordinate()` converts this matrix into a 2xn octal coordinate matrix S. * The function `getData()` extracts five columns of target data from the original dataset and places them into a matrix named `Originaldata`. This matrix is ​​then uniformly divided into an 8×8 square matrix of identity arrays called `OriginalCell`. The `OriginalCell`, differential privacy parameters ε and δ are substituted into the function `embed_ComputeSigmaMat()` to calculate the range of values ​​for the parameter σ of the differential privacy Gaussian mechanism. Finally, the function `embed_NoiseToOD()` sequentially inserts Gaussian noise sets generated according to different σ values ​​into different positions within `OriginalCell`, where these positions use a coordinate matrix generated by the function `embed_GetCoordinate()`. Furthermore, the function `embed_GetTotalODHash()` calculates the hash value of the original dataset and stores it row-wise in a matrix `H` for potential subsequent fingerprint detection.

[0096] The process of semi-blind fingerprint detection is as follows Figure 5As shown in the figure. This embodiment uses a method that compares the hash result of the illegal dataset to be detected with the hash result of the corresponding original dataset to achieve semi-blind detection of digital fingerprints. By comparing these two hash results, coordinate points with different hash values ​​can be marked. These coordinate points are a series of values ​​from 0 to 63 (a total of k×k = 8×8 = 64). Dividing these values ​​by k (k = 8), the quotient and remainder are the horizontal and vertical coordinates of the desired marked points, respectively. These calculated results are stored column-wise in matrix S. * In the middle, the horizontal coordinates are stored in matrix S. * The first row contains the y-coordinates, and the second row contains the y-coordinates. The correct order of the marked coordinate points can be determined by calculating and rearranging the data blocks at specific locations. According to S... * The coordinates of the points are extracted and fitted sequentially to calculate {D}. *} k×k The variance of the Gaussian distribution corresponding to the data block. Store these variance values ​​column-wise in matrix M. Then store S... * Combine M and M row by row to form a new matrix U = [S * M]′, and rearrange U column-wise according to the values ​​of the last row of the new matrix. The first 2 rows and first k columns of matrix U constitute the noise fingerprint matrix (S) of the dataset to be detected. * ) k Finally, in the case of (S) * ) k Matrix transformations and number system conversions are performed to finally calculate the identifying information s.

[0097] The semi-blind fingerprint detection process can be implemented using a compilable programming language. Figure 6 This section specifically demonstrates the simulation implementation flow of the semi-blind digital fingerprint detection process of this method. The function detect_GetNDHash(ND) calculates the dataset to be detected {D}. *} k×k The hash value of each data block is given, where the input parameter ND of the function refers to the dataset to be detected. These are then stored in a 64×32 character array H. * The function `detect_compareHash()` compares the hash matrix H of the original dataset with the matrix H generated by the function `detect_GetNDHash(ND)`. * And record the coordinates of the positions with different contents in matrix S. * In the middle. It is worth noting that the coordinate matrix S at this time... * It is unordered. The function detect_resetCoor() is used to fit {D} *} k×k The variance of the data block corresponding to the coordinate position in the middle, where the coordinate position is based on S. *The matrix S was extracted column-wise and then rearranged according to the magnitude of the different variance values. * At this time, S * The first k columns are the desired fingerprint coordinates. Finally, the function `detect_getSeed()` performs matrix transformations and number system conversions to ultimately calculate S. * The corresponding identifying information s.

[0098] The simulation experiment uses a record with 420,768 numerical data points, each record containing 5 columns of attribute data. A time accurate to the second, in the form of a string like "20221106212056", is chosen as the identifying information. This is because time is dynamic and difficult to control compared to user IDs or MAC addresses. If the time can be accurately embedded and detected, other static seeds are easier to implement. During the simulation, values ​​are assigned to the key data of the method, where k = 8, ε ∈ [0.5, 1], δ ∈

[10] . -5 [1]. Through simulation experiments and analysis, our method can successfully embed and extract correct identifying information. Figure 7 , 8 Points 9 and 9 respectively demonstrate the usability of TND from different perspectives. Figure 7 The overall data differences between the original dataset D and the noisy dataset TND were compared; Figure 8 and 9 Specifically, the differences between the fingerprint coordinates of the original dataset and the noisy dataset were compared among eight sets of data blocks. Figure 8 The differences between data blocks were compared from a statistical probability perspective, and Figure 9 This compares the differences in key values ​​between data blocks, such as the mean and variance.

[0099] Figure 7 The probability distribution of the entire dataset across different numerical ranges is displayed. For the experiment in this embodiment, the entire data range was divided into 30 equal parts, and the percentage of data in each range was calculated. The bar chart shows that the original dataset and the noisy dataset do not differ significantly in probability distribution, and the difference is positively correlated with the amount of data. For example, the difference between the original dataset and the noisy dataset in the range of 0 and 100 is significantly greater than the difference between 500 and 600.

[0100] Figure 8 The comparison focuses on the difference in probability statistics between the data blocks corresponding to fingerprint coordinates in the original and noisy datasets. It can be viewed as... Figure 7 Extraction and refinement. The fingerprint coordinates are ultimately determined by using 8 coordinate points from 64 data block partitions as locations for embedding differential privacy Gaussian noise. Figure 8The probability distributions of the original data blocks and the noisy data blocks at these 8 locations were compared. The comparison results show that the probability distributions of the two datasets are consistent, and the differences are negligible.

[0101] Figure 9 The differences between key values ​​of fingerprint coordinates across the eight datasets were supplemented by a combination of box and bar graphs. Figure 9 The diagram displays the minimum (min), 25th percentile (Q1), median, 75th percentile (Q3), and maximum (max) values ​​for eight pairs of data blocks. Q1, median, and Q3 form a box with compartments, and an extension line between Q3 and the maximum, and between Q1 and the minimum, indicating the dispersion of the data. The comparison in this diagrammatic form shows that the key values ​​between the original data blocks and the TND blocks are not significantly different, and the larger numerical differences between the groups are reasonable.

[0102] Robustness describes the survivability of digital fingerprints after data processing operations. Based on the high randomness of the method in this embodiment and the characteristics of hash verification, the focus is on the impact of randomly deleting, inserting, and modifying data in a skip-like manner on the detection results. Experiments are conducted to examine the relationship between the amount of data deleted or inserted and the accuracy of fingerprint detection.

[0103] Figure 10 The display shows the fingerprint correct recognition rate (CRR) after randomly deleting different amounts of data. Figure 11 The relationship between the mean and variance of the CRR before and after removing different amounts of data is shown.

[0104] Deletion attacks delete data in an incremental manner. We randomly select 2 from TND. n Data, where n = 1, 2, ..., 20. The initial size of TND is r = 420768 × 5. The Gaussian mechanism should produce (r / 64) × 8 random noise embedded at specific fingerprint coordinates in the original dataset.

[0105] Because the noise generated by the Gaussian mechanism is random, the TND obtained each time is also different. And the experiment needs to delete 2... n The data was also randomly selected. (After deleting 2...) n After processing the data, a binary fingerprint (S) is calculated. D )2. Then (S) D Compare 2 and (S)2 bit by bit. Delete 2. n The CRR after the data is (S) D The ratio of the number of digits that have the same value as (S)2 to the total number of digits. Figure 10 The results show that the CRR is inversely proportional to the amount of data deleted, and when the amount of data deleted reaches 2...15 At that time, the CRR decreased significantly.

[0106] For universality, delete 2 for each n The experiment was conducted 50 times for each data point. Figure 11 Compare and calculate deletion 2 n The mean and variance of the CRR for 50 trials were calculated for each data set. To better illustrate the CRR, we calculated the mean and variance of the CRR for each group of 50 data points. As shown in the figure, there are 9 groups of bar charts. Each group consists of three closely connected bar charts, representing the amount of data deleted, the mean of the CRR after 50 trials, and the variance. The graph shows that the mean is inversely proportional to the amount of data deleted; when the amount of data deleted reaches 2... 15 At that point, the average value stabilized. And when the deletion count reached 2... 15 At that time, the average CRR stabilized at around 50%, which is consistent with... Figure 10 Consistent.

[0107] Similar to the deletion experiment, 2 n Data can also be randomly embedded into the TND. Figure 12 and Figure 13 The results for the randomly embedded data are shown. Comparable results are provided by 2... n The consistency between the binary fingerprint (S1)2 generated from the result set of random data and the original fingerprint (S)2. Figure 12 It can be seen that the embedding attack experiment is slightly less effective than the deletion attack experiment. Figure 13 The insertion 2 in each experiment was compared and calculated. n The mean and variance of 50 CRR executions of the data. Figure 1 The results are consistent with those in the previous section. Figure 13 This indicates that when the amount of embedded data reaches 2 10 At that time, the average CRR decreased and stabilized at around 50%.

[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0109] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for detecting the identifying information of relevant users in an illegally disseminated dataset, characterized in that: Includes the following steps: Following data privacy protection methods, the carrier dataset is protected to obtain the corresponding noisy dataset. In the process of implementing data privacy protection methods, the aggregated data block set is processed before embedding random noise blocks. Each data block is hashed to obtain a hash matrix. ; Obtain the vector dataset to be detected The vector dataset to be detected According to the carrier dataset The same block partitioning method is used to divide into From the data blocks, we obtain the aggregated data block set. ; Calculate the aggregated data block set The hash value of each data block is used to obtain the corresponding hash matrix. ; By comparing the hash matrices and hash matrix Record the coordinates of the two hash matrix values ​​that are different, resulting in 2 rows. Column coordinate matrix ; Based on the coordinate matrix Locate and extract the aggregated data block set from the coordinates in the data. The variance of the corresponding data blocks is calculated using Gaussian fitting. The variance obtained by Gaussian fitting Stored in matrix middle; coordinate matrix sum matrix Merge rows into a matrix Then the matrix Based on the matrix The elements in the matrix are rearranged column by column to obtain a new matrix. ; Extracting new matrices The first two lines, the first Column as Noise fingerprint matrix in base 1 ; Noise fingerprint matrix Noisy fingerprint data converted to decimal form; The identification information string of the relevant user is obtained from the noisy fingerprint data in decimal form by the permutation decryption algorithm. This identification information string is the user fingerprint information in string form. The data privacy protection method includes: Receive privacy protection requests from data requesters and obtain the data set and user fingerprint information of the data subject to privacy protection; Carrier dataset Divided into appropriate From the data blocks, we obtain the aggregated data block set. ; Encode the user's fingerprint information into a fingerprint coordinate matrix. Then, a permutation encryption algorithm is used to obtain the fingerprint coordinate matrix. Permutation encryption is performed to obtain a new fingerprint coordinate matrix. ; Computational vector dataset The differential privacy sensitivity is calculated, and the differential privacy sensitivity is compared with the differential privacy parameter. and By combining these methods, the variance of the differential privacy Gaussian mechanism is calculated. The range of values; by using the variance of the differential privacy Gaussian mechanism The range of values ​​is divided into Divide into equal parts and place them into the matrix in order of their values. In the middle, we obtain the matrix , matrix Each scalar in the matrix is ​​substituted sequentially into the differential privacy Gaussian mechanism to generate a random noise block matrix. ; Extract new fingerprint coordinate matrix by column Each column has two values ​​used to locate the aggregated data block set. At the locations where random noise blocks need to be embedded, embed random noise blocks at the corresponding locations to obtain a noisy dataset. This noisy dataset The elements in the data are data stored in a privacy-protected manner. Noisy dataset Feedback is sent to the data requester; The coordinates of the two hash matrix values ​​that are different are recorded, resulting in 2 rows. Column coordinate matrix ,include: Record the coordinates of the points where the two hash matrix values ​​differ, and record the x-coordinates of each point in the coordinate matrix column by column. The first row, the ordinate, is recorded in the coordinate matrix. The second line ultimately results in a 2-line output. Column coordinate matrix ; The variance of these data blocks is calculated using Gaussian fitting. ,include: Calculate the variance of these data blocks using the following formula. : (3) In the formula, This represents the Gaussian fitting function. Represents a collection of aggregated data blocks The data block in row r and column c.

2. The method for detecting the identifying information of relevant users in an illegally disseminated dataset according to claim 1, characterized in that: The aforementioned carrier dataset It consists of aggregated text-based numerical data.

3. The method for detecting the identifying information of relevant users in an illegally disseminated dataset according to claim 1, characterized in that: The process of encoding user fingerprint information into a fingerprint coordinate matrix Specifically, it includes: User fingerprint information in string format Convert to decimal matrix In this form, the decimal matrix is ​​then... Convert the form to a binary matrix form; Based on the aggregated data block set The number of data blocks in the binary matrix The form is converted to have the corresponding base. matrix form; matrix Transform into a 2-line array Fingerprint coordinate matrix of the column .

4. The method for detecting the identifying information of relevant users in an illegally disseminated dataset according to claim 1, characterized in that: The aforementioned permutation encryption algorithm is used to obtain the fingerprint coordinate matrix. The process of performing substitution encryption includes: Through permutation matrix The obtained fingerprint coordinate matrix Perform permutation encryption; the permutation matrix It is based on a custom permutation key It is calculated from the sequence of each letter in the alphabet.

5. The method for detecting the identifying information of relevant users in an illegally disseminated dataset according to claim 1, characterized in that: The aforementioned computing carrier dataset Differential privacy sensitivity includes: The carrier dataset is calculated according to equation (1). Differential privacy sensitivity : (1) In the formula, It is a query function. Represents the carrier dataset, Represents the carrier dataset sibling datasets; and There is only one data point that differs between them; The aforementioned combination of differential privacy sensitivity and differential privacy parameters and By combining these methods, the variance of the differential privacy Gaussian mechanism is calculated. The range of values ​​for includes: According to equation (2), the variance of the differential privacy Gaussian mechanism is calculated. The range of values ​​for: (2) In the formula, To mitigate the sensitivity of differential privacy, and All are differential privacy parameters.