Method for improving data collection quality based on worker task completion recording and privacy protection

Through the reputation management and privacy protection mechanism based on worker history, combined with ABOD anomaly detection and ring signature technology, the problems of workers' privacy leakage and data quality decline in group intelligence perception networks are solved, efficient and reliable data collection is achieved, and the application of smart cities and smart networks is promoted.

CN120523802APending Publication Date: 2025-08-22CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604419.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

In the group intelligence perception network, the existing technology has problems of workers' privacy leakage and degradation of data quality, especially the reputation management mechanism is prone to privacy leakage, and the CRH algorithm fails when the proportion of malicious workers increases, making it difficult to ensure high-quality data collection.

Method used

By introducing a reputation management scheme and privacy protection mechanism based on worker history completion records, elliptic curve encryption and polynomial ring structure are used to generate keys for workers and data demanders, combined with ABOD anomaly detection and ring signature technology, the protection of workers' identity and reputation is achieved, and data quality classification is carried out through trusted indicators and approximate angle variance.

Benefits of technology

It realizes that while protecting workers' privacy, selecting high-quality workers, improving the quality of data collection, reducing the time and iteration of truth discovery, and improving the reliability and accuracy of data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523802A_ABST
    Figure CN120523802A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving data collection quality based on worker task completion recording and privacy protection, and relates to the field of trusted computing and data collection in a crowd-sourcing network. The network structure is composed of a task platform, a reputation management center, a data demander and workers. In a network, a task platform issues a task and recruits a plurality of workers to collect target data, a privacy protection protocol is designed, the reputation and identity information of the workers is subjected to anonymous and encrypted protection, and a novel truth discovery algorithm is utilized to aggregate the data of the plurality of workers, so that a more accurate estimated truth value is obtained. Particularly, different from the prior art that a specific reputation score is given to the worker, the task completion condition of the worker is evaluated into three types of indexes, and the reputation of the worker is fairly evaluated through a statistical method, so that the problems of data quality deficiency and difficulty in accurate measurement of the reputation of the worker in a traditional method are solved, and high-quality data collection work is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of trusted computing and data collection in crowd intelligence networks, and more particularly to a method for improving data collection quality based on worker task completion records and privacy protection. Background Art

[0002] With the emergence of new concepts such as smart cities and intelligent networks, mobile crowd sensing (MCS) has become increasingly prominent due to its powerful data collection and efficient network sensing capabilities. By recruiting a critical number of workers, crowd sensing networks can quickly and accurately collect critical urban information such as weather, temperature, water quality, and traffic flow. However, since workers must submit sensitive information such as location and time when collecting data, this information must be protected. Furthermore, malicious data submitted by dishonest workers can degrade the quality of the collected data, making quality a key issue in crowd sensing networks.

[0003] Currently, in the widely accepted working model of the crowd-intelligence perception network, the central server acts as the task center, responsible for announcing to workers the location, time, and other information of the data to be collected. To ensure the quality of the collected data, the same task is jointly performed by N workers and finally submitted to the data demander. The data demander obtains the target data through a certain truth discovery strategy, thus completing a typical task process of the crowd-intelligence network.

[0004] However, the above task execution process has the following shortcomings: First, to recruit higher-quality workers, each worker is often assigned a reputation weight. This reputation is then updated over time by comparing the quality of their data to distinguish between different workers. Because each worker has a unique reputation value, comparing and labeling the reputations of different workers makes it easy to surveil their identities and private information. Furthermore, the data submitted by workers cannot be easily shared with entities other than the data requester, such as the platform or the worker. Otherwise, data leakage would result in serious privacy breaches. Second, current methods for calculating the true value mostly rely on the reputation-based CRH algorithm. When workers are of high quality, the CRH algorithm can aggregate low-error estimates of the true value. However, as the proportion of malicious workers increases, the aggregated data from the CRH algorithm often drops sharply, and can even fail completely when the number of malicious workers exceeds half. Although the reputation-based CRH algorithm can reduce some errors, it does not fully address the quality issue. Furthermore, the reputation-based CRH algorithm is difficult to integrate with privacy-preserving methods, necessitating the simultaneous resolution of worker privacy issues. Third, in terms of current reputation management, the current reputation management mechanism often favors the quality of recent tasks. That is, if a worker completes several tasks in a row and submits good data, his or her reputation value will increase rapidly, thereby covering up the decline in reputation caused by previous low-quality work and resulting in an incorrect evaluation of the worker's quality.

[0005] For the above reasons, there is a need for a truth discovery algorithm that can solve the privacy leakage caused by workers' reputation while designing a truth discovery algorithm different from CRH, and cooperate with the workers' reputation management mechanism to complete high-quality data collection, so that MCS can further develop. Summary of the Invention

[0006] In light of this, the present invention provides a method for improving data collection quality based on worker task completion records and privacy protection. By innovatively introducing a reputation management scheme and privacy protection mechanism based on workers' historical completion records, this method addresses the shortcomings of traditional reputation management schemes in this context, providing more objective and fair evaluations of workers. Furthermore, using an anomaly detection scheme based on ABOD, it classifies worker data quality, achieving more efficient data collection and protecting worker privacy. Overall, the present invention first assigns workers and data requesters their own keys, enabling data requesters to publish data-demanding tasks on the task platform and select an appropriate number of task executors from among the registered workers to achieve data collection. The present method has three key points: First, a worker task registration and reputation verification mechanism that fully protects worker privacy allows workers to effectively register for tasks and verify their identities, while data requesters can select workers and efficiently obtain the data they collect. Second, workers' historical completion records are used to comprehensively evaluate their reputation. The main idea can be summarized as categorizing each worker's historical task completion record into three categories: excellent, average, and unqualified. Statistical principles are then used to calculate the probability that the worker's next task will be categorized into one of these three categories. Uncertainty is also used to correct for missing samples caused by insufficient number of completed tasks. This approach addresses the challenges of traditional reputation management while ensuring a fair and objective evaluation of a worker's reputation. Finally, an anomaly detection scheme similar to ABOD is used to objectively and impartially classify the data from the worker's current task. The third mechanism is a reputation update check mechanism. After a data demander updates a worker's reputation, the worker can check the status of their own task update, thereby preventing the data demander from irresponsibly updating their reputation.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] First, the system is initialized, and the task platform assigns the worker and the data requester their own initial keys.

[0009] The data demander publishes a data demand task T on the task platform, hoping to recruit N workers to complete the task;

[0010] The task platform broadcasts task T to workers;

[0011] Workers send requests to the task platform to apply for corresponding tasks;

[0012] Workers upload relevant data and credibility certificates of tasks;

[0013] The data demander obtains and decrypts the corresponding worker-submitted data;

[0014] Data demanders verify the worker's credit record and identity;

[0015] Data demanders evaluate the data quality of workers and obtain target data;

[0016] Data demanders update workers’ credit records;

[0017] Workers verify the records updated by data requesters.

[0018] Optionally, during the system initialization phase, the task platform generates identity keys for workers and data demanders. i , generate a private key k that satisfies elliptic curve encryption i and public key K i ; Then, suppose there is a polynomial ring f(y)=y n -1, Represents the set of all polynomials with integer coefficients whose highest degree does not exceed n-1, where n is a prime number. is a power of 2. The worker chooses Polynomial with coefficients 0, 1, -1 And stipulate satisfy is a system parameter equal to 3. Next, calculate the worker’s second private key Public Key Above, the worker's key pair set is represented as ({K i ,k i},{Pk i ,Sk i Similarly, according to the above method, the key pair set generated for the data demander is ({K D ,k D},{Pk D ,Sk D}).

[0019] Optionally, workers apply for tasks using a privacy-preserving framework to protect the worker’s identity and reputation privacy. Workers send task request information to the task platform. <Pk i ,T>. After the data demander receives the task request information from the worker, if the corresponding worker is selected to perform the task, the following operations will be performed: First, randomly select And perform the following calculations: R = bG, R' = b'G, k T =H1(b·Pk i ), K T =k T G, above, data demanders publish information on selecting workers on the task platform <T,K T ,R,R′>. The worker uses his personal private key Ski Can verify whether they are selected to perform task T, that is, if the worker verifies that K T =H1(R·Sk i )G, it means that the worker knows that he has been selected to perform the task, and k T is the corresponding task key.

[0020] Optionally, workers can use data encryption to protect their personal identity and reputation privacy. A worker's reputation is determined by their personal record of completed tasks. Each worker is assigned a set of triples A = (α h ,α u ,α l ), α h Represents the number of times a worker’s performance in completing tasks has been rated as excellent, α u represents the number of times a worker's task completion performance was rated as average in the past, α l The number of times a worker's performance in completing tasks was rated as unsatisfactory. i A i When first recorded in the reputation center, the reputation management center generates an initial reputation certificate Pro (A i )=H1(H1(0),A i ), when a worker completes the task numbered T-1, the corresponding data demander evaluates the worker and then uses the relevant task key k T-1 In the new block update proof is Pro(A i )=H1(H1(k T-1 ),A i ). In task T, when the worker collects data Post-selection And use Pk i Encrypt the data to obtain Among them, k T-1 is the private key of the worker's last completed task, and | is the character concatenation operation. Subsequently, the worker selects Polynomials on satisfy Finally, the worker uses the public key Pk of the data demander D Get encrypted data and calculate In summary, the information submitted by workers on the task platform is

[0021] Optionally, the data demander can decrypt the worker’s task data and reputation information by re-encrypting the key to protect the worker’s privacy. D Decryption get and calculate Among them Sk D satisfy Data demanders will Send to the task center to obtain the re-encryption key Rk iB The mission center receives After that, calculate And Rk iB Return to the data demander. The data demander uses Rk iB and Sk D Data Decryption is performed to obtain data and verify credibility. The decryption method is:

[0022] Optionally, workers can verify the legitimacy of their personal identities through proof of credibility. After decryption, the data demander obtains the data set of all workers on task T. Using the block height information, data demanders can extract the latest credit proof Pro(A) of each worker. i ), verified by Pro(A i )=H1(H1(k T-1 ),A i ) Determine whether the worker’s credibility is genuine.

[0023] Optionally, the worker's data quality can be evaluated by statistics and an anomaly detection scheme based on approximate ABOD. The evaluation method is as follows: using the ternary set A, the probability distribution space of the worker's current task performance as excellent, average, and unqualified is predicted to be X = (X h ,X u ,X l ), where the probability distribution satisfies ∑ β∈{h,u,l} X β =1,0≤X β ≤1. It is known that Γ(·) is a Gamma function, and the weight parameter K is set to (k h ,k u ,k l ) is the factor that adjusts the probability of the distribution space, from which we can calculate It represents the possibility of predicting the three types of evaluations of workers for the current task using the relevant statistical distribution, which satisfies the following equation:

[0024]

[0025] Therefore, using The expected results under the posterior distribution are obtained by summing A, and Laplace smoothing is added to correct the expectation to prevent the problem of zero probability, and the expected results of the last three probability distributions of the workers are obtained. It can be defined by the following equation:

[0026]

[0027] Next, calculate the uncertainty in the expected result Given the probability density function f(·) under the Dirichlet distribution, the uncertainty The expression can be described as:

[0028]

[0029] Combining the above process, we define a trust index Φ of a triple i =(φ h,i ,φ l,i ,φ u,i ) to evaluate the comprehensive reputation of workers, the credibility index Φ i and parameter set A, The following relationship is satisfied between K:

[0030]

[0031] φ u,i =1-φ h,i -φ l,i (5)

[0032] Next, using the credibility index Φ i Perform truth discovery and use the approximate angle variance method to evaluate the quality of worker data. In each truth discovery task T, the worker group Submit perception data collection In the perception task, Can often be represented as a set of one-dimensional vectors In order to detect abnormal classes in sensor data, an abnormal data detection method based on ABOD is adopted to overcome the problem that traditional data detection relies on the distance similarity between data, which leads to frequent parameter changes and time-consuming traversal. The ABOD method can be expressed as: vector Given paradigm The scalar product between two vectors is <·,·>: for represents the difference between two vectors, Then the vector The angular variance can be expressed as:

[0033]

[0034]

[0035] The variance of points inside the data set is often high, while the variance of outliers is very low. Setting reasonable thresholds μ1 and μ2 can Divide into interior point sets Boundary point set and the set of outliers The three correspond to high quality, uncertain quality and low quality evaluation respectively. Then, based on the evaluation index Φ, a set of workers with high reputation evaluation is selected. And in the collection Randomly select δ workers from the set to form a subset The data of high-quality workers are often located at the interior and boundary points, which have the largest calculation weights. It is stipulated that δ has a minimum number of people to ensure the lower bound of the approximate variance, so the vector The approximate angular variance of for:

[0036]

[0037] Using approximate angular variance The workers' data is classified by quality, and a weighted average process is performed using high-quality data sets to finally obtain the data required by the data demander.

[0038] Optionally, the worker's reputation can be updated by embedding the ring signature into the blockchain while protecting the worker's privacy. The data demander updates the worker's reputation record in the following way: After evaluating the data quality of each worker, the data demander obtains a new round of reputation evaluation results A′ i =(α′ h,i ,α′ u,i ,α l ' ,i ), and construct a new reputation proof Pro(A′) for each worker i )=H1(H1(b·Pk i ),A′ i ). In order to effectively embed the reputation proof into the blockchain, the data demander constructs a ring signature for each worker to anonymize the worker's identity while verifying the legitimacy of the worker's new reputation proof. i Constructing a ring signature, the data demander combines the task public key randomly generated by the system with Pro(A′ i ), generate the public key group: L T ={K T,1 ,K T,2 ,...,H1(H1(kT-1 ),A′ π )G,...,K n}, true credibility certificate Pro(A′ i ) is hidden in the πth position. Using the public key group L T and H1(H1(k T ),A′ π ), in the reputation management center, the data demander selects k T,π =H1(H1(k T-1 ),A′ π ),}. In this way, the data demander generates a random message m T,i , using the generation algorithm GEN(·) to generate a new ring signature σ for the worker T,i .

[0039] The data demander will sign the ring T,i Sent to the reputation management center, the reputation management center starts the reputation update contract. First, the blockchain miner on the reputation management center receives σ T,i Then, using the verification algorithm VER(m T,i ,σ T,i ) Verify the legitimacy of the group signature. Finally, L T Stored on the new reputation proof block, each public key of the public key group L is stored at the corresponding height h j , j=1,2,...,π,...,n, but only in h π The public key on the worker N is i Proof of legal credibility.

[0040] Optionally, a detailed reputation update verification method is used to ensure that the data demander reasonably updates the worker's reputation. After the data demander completes the reputation update for the worker, select Calculate D = b′ + b″k D , send information to workers The worker decrypts the information and verifies it. The verification is divided into three aspects: First, verify DG = R' + b"K D Is it true? If so, it proves that the credit update is done by the corresponding data demander. β∈{h,u,l} α′ β,i -α β,i =1 to prove that the data demander has correctly updated the interaction record. Finally, verify H1(H1(k T ),A′ π )=H1(H1(H1(R·Sk π )),A′ π) to prove that the credit certificate is correct and legal. After successful verification, the worker can obtain the corresponding task reward and use the new round of credit certificate to apply for the next round of tasks.

[0041] As can be seen from the above technical solutions, compared to existing technologies, this invention provides a method for improving data collection quality based on worker task completion records and privacy protection. This method protects worker identity, reputation, and data privacy while also selecting workers who provide high-quality data through a reasonable reputation management mechanism, thereby further improving the quality of collected data. Furthermore, this invention offers advantages in terms of time consumption and number of iterations required for truth discovery. This provides a reliable, high-quality data collection solution for crowd-sensing networks that effectively protects worker privacy, and can promote the application and development of crowd-sensing in smart cities and smart networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0043] Figure 1 A framework flow chart of the invention method;

[0044] Figure 2 This is the comparison of the mean square error of the true value data obtained;

[0045] Figure 3 Comparison of the number of iterations of the truth discovery algorithm;

[0046] Figure 4 Comparison of the running time required for truth discovery algorithms. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] Example 1

[0049] An embodiment of the present invention discloses a method for improving data collection quality based on worker task completion records and privacy protection, comprising the following steps:

[0050] Step 1: The system performs initialization operations, and the task platform assigns the workers and data demanders their own initial keys.

[0051] The task platform generates identity keys for workers and data demanders respectively. i , generate a private key k that satisfies elliptic curve encryption i and public key K i ; Then, suppose there is a polynomial ring f(y)=y n -1, Represents the set of all polynomials with integer coefficients whose highest degree does not exceed n-1, where n is a prime number. is a power of 2. The worker chooses Polynomial with coefficients 0, 1, -1 And stipulate satisfy is a system parameter equal to 3. Next, calculate the worker’s second private key Public Key Above, the worker's key pair set is represented as ({K i ,k i},{Pk i ,Sk i Similarly, according to the above method, the key pair set generated for the data demander is expressed as ({K D ,k D},{Pk D ,Sk D}).

[0052] Step 2: The data demander publishes the data demand task T on the data platform, hoping to complete the task by recruiting N workers.

[0053] Step 3: The task platform broadcasts task T to workers.

[0054] Step 4: Workers send task request information to the task platform <Pk i ,T>. After the data demander receives the worker's request information, if he chooses the corresponding worker to perform the task, he will do the following: First, randomly select G is the base point of the elliptic curve, and then the following calculation is performed: R = bG, R' = b'G, k T =H1(b·Pk i ), K T =k T G. Above, data demanders publish information on selecting workers on the task platform <T,K T ,R,R′>. The worker uses his personal private key Sk iCan verify whether they are selected to perform task T, that is, if the worker verifies that K T =H1(R·Sk i )G, it means that the worker knows that he has been selected to perform the task, and k T is the corresponding task key.

[0055] Step 5: Workers upload relevant data and credit certificates for the task.

[0056] The reputation of a worker is determined by the historical record of his or her completed tasks. Each worker is assigned a set of triples A = (α h ,α u ,α l ), α h Represents the number of times a worker’s performance in completing tasks has been rated as excellent, α u represents the number of times a worker’s task completion performance was rated as average in the past, α l The number of times a worker's performance in completing tasks was rated as unsatisfactory. i A i When first recorded in the reputation center, the reputation management center generates an initial reputation certificate Pro (A i )=H1(H1(0),A i ), when a worker completes the task numbered T-1, the corresponding data demander evaluates the worker and then uses the relevant task key k T-1 In the new block update proof is Pro(A i )=H1(H1(k T-1 ),A i ). In task T, when the worker collects data Post-selection And use Pk i Encrypt the data to obtain Among them, k T-1 is the private key of the worker's last completed task, and | is the character concatenation operation. Subsequently, the worker selects Polynomials on satisfy Finally, the worker uses the public key Pk of the data demander D Get encrypted data and calculate In summary, the information submitted by workers on the task platform is

[0057] Step 6: The data demander receives the submitted data and uses Sk D Decryption get and calculate Among them SkD satisfy Data demanders will Send to the task center to obtain the re-encryption key Rk iB The mission center receives After that, calculate And Rk iB Return to the data demander. The data demander uses Rk iB and Sk D Data Decryption is performed to obtain data and verify credibility. The decryption method is:

[0058] Step 7: The data demander verifies the worker’s credit record and identity. Through decryption, the data demander obtains the data set of all workers on task T. Using the block height information h, the data demander extracts the latest credit proof Pro (A i ), verified by Pro(A i )=H1(H1(k T-1 ),A i ) Determine whether the worker’s credibility is genuine.

[0059] Step 8: The data demander evaluates the data quality of the worker and obtains the target data. The evaluation method is as follows: Using the ternary set A, the probability distribution space of the worker's current task performance as excellent, average and unqualified is predicted as X = (X h ,X u ,X l ), where the probability distribution satisfies ∑ β∈{h,u,l} X β =1,0≤X β ≤1. It is known that Γ(·) is a Gamma function, and the weight parameter K is set to (k h ,k u ,k l ) is the factor that adjusts the probability of the distribution space, from which we can calculate It represents the possibility of predicting the three types of evaluations of workers for the current task using the relevant statistical distribution, which satisfies the following equation:

[0060]

[0061] Therefore, using The expected results under the posterior distribution are obtained by summing A, and Laplace smoothing is added to correct the expectation to prevent the problem of zero probability, and the expected results of the last three probability distributions of the workers are obtained. It can be defined by the following equation:

[0062]

[0063] Next, calculate the uncertainty in the expected result Given the probability density function f(·) under the Dirichlet distribution, the uncertainty The expression can be described as:

[0064]

[0065] Combining the above process, we define a trust index Φ of a triple i =(φ h,i ,φ l,i ,φ u,i ) to evaluate the comprehensive reputation of workers, the credibility index Φ i and parameter set A, The following relationship is satisfied between K:

[0066]

[0067] φ u,i =1-φ h,i -φ l,i (5)

[0068] Next, using the credibility index Φ i Perform truth discovery and use the approximate angle variance method to evaluate the quality of worker data. In each truth discovery task T, the worker group Submit perception data collection In the perception task, Can often be represented as a set of one-dimensional vectors In order to detect abnormal classes in sensor data, an abnormal data detection method based on ABOD is adopted to overcome the problem that traditional data detection relies on the distance similarity between data, which leads to frequent parameter changes and time-consuming traversal. The ABOD method can be expressed as: vector Given paradigm The scalar product between two vectors is <·,·>: for represents the difference between two vectors, Then the vector The angular variance can be expressed as:

[0069]

[0070]

[0071] The variance of points inside the data set is often high, while the variance of outliers is very low. Setting reasonable thresholds μ1 and μ2 can Divide into interior point sets Boundary point set and the set of outliers The three correspond to high quality, uncertain quality and low quality evaluation respectively. Then, based on the evaluation index Φ, a set of workers with high reputation evaluation is selected. And in the collection Randomly select δ workers from the set to form a subset The data of high-quality workers are often located at the interior and boundary points, which have the largest calculation weights. It is stipulated that δ has a minimum number of people to ensure the lower bound of the approximate variance, so the vector The approximate angular variance of for:

[0072]

[0073] Using approximate angular variance The workers' data is classified by quality, and a weighted average process is performed using high-quality data sets to finally obtain the data required by the data demander.

[0074] Step 9: The data demander updates the worker's credit record. After evaluating the data quality of each worker, the data demander obtains a new round of credit evaluation results A′ i =(α′ h,i ,α′ u,i ,α l ' ,i ), and construct a new reputation proof Pro(A′) for each worker i )=H1(H1(b·Pk i ),A′ i ). In order to effectively embed the reputation proof into the blockchain, the data demander constructs a ring signature for each worker to anonymize the worker's identity while verifying the legitimacy of the worker's new reputation proof. i Constructing a ring signature, the data demander combines the task public key randomly generated by the system with Pro(A′ i ), generate the public key group: L T ={K T,1 ,K T,2 ,...,H1(H1(k T-1 ),A′ π )G,...,K n}, true credibility certificate Pro(A′ i ) is hidden in the πth position. Using the public key group L T and H1(H1(kT ),A′ π ), in the reputation management center, the data demander selects k T,π =H1(H1(k T-1 ),A′ π ),}. In this way, the data demander generates a random message m T,i , using the generation algorithm GEN(·) to generate a new ring signature σ for the worker T,i :

[0075] The data demander will sign the ring T,i Sent to the reputation management center, the reputation management center starts the reputation update contract. First, the blockchain miner on the reputation management center receives σ T,i Then, using the verification algorithm VER(m T,i ,σ T,i ) Verify the legitimacy of the group signature. Finally, L T Stored on the new reputation proof block, each public key of the public key group L is stored at the corresponding height h j , j=1,2,...,π,...,n, but only in h π The public key on the worker N is i Proof of legal credibility.

[0076] Step 10: The worker verifies the records updated by the data demander. After the data demander completes the credit update for the worker, select Calculate D = b′ + b″k D , send information to workers The worker decrypts the information and verifies it. The verification is divided into three aspects: First, verify DG = R' + b"K D Is it true? If so, it proves that the credit update is done by the corresponding data demander. β∈{h,u,l} α′ β,i -α β,i =1 to prove that the data demander has correctly updated the interaction record. Finally, verify H1(H1(k T ),A′ π )=H1(H1(H1(R·Sk π )),A′ π ) to prove that the credit certificate is correct and legal. After successful verification, the worker can obtain the corresponding task reward and use the new round of credit certificate to apply for the next round of tasks.

[0077] Example 2

[0078] The difference between this embodiment and embodiment 1 is that:

[0079] In a smart city's perception network, it's necessary to capture key data from a specific location in the city, ensuring high-quality data standards. For example, information such as traffic flow, weather conditions, and real-time humidity must be maintained. Furthermore, all data collected by mobile workers must be transmitted to the platform via the network for processing. After data collectors obtain the data collected by workers, they can use the method of this invention to improve data aggregation quality without revealing the workers' specific identities, thereby obtaining more accurate data.

[0080] The experimental results of the inventive method are given below.

[0081] Figure 1 The framework flowchart of the present invention's method is provided. As can be seen from the figure, this system has four main entities. This invention assumes that all four entities conform to the semi-honest model: they will honestly fulfill their respective obligations but may curiously infer sensitive content from existing information. The framework flowchart is clearly described in the claims of this invention.

[0082] Figure 2 The error between the estimated true value and the actual value obtained by the algorithm of the present invention and the comparison algorithm is given, and we use the root mean square error (RMSE) to describe it. Benchmark algorithm one PRTD is a truth discovery method that incorporates reputation as a part of the weight into the CRH process; benchmark algorithm two CRH-SVM is a strategy that first eliminates a small amount of abnormal data through the SVM algorithm, and then uses the CRH method to discover the truth; benchmark algorithm three CRH is a typical truth discovery algorithm, which obtains the truth value by aggregating data from different workers through a certain strategy; benchmark method four RTD uses reputation as the initial weight to perform the CRH process; benchmark algorithm five QE uses the traversal method to measure the distance between each two points, and finally obtains the point at the center as the true value. Through Figure 2 It can be concluded that the method of the present invention has the smallest error of 0.199, which is 31.8% lower than that of the CRH algorithm.

[0083] Figure 3 A comparison of the runtime of our algorithm and other benchmark algorithms is presented. This comparison shows that our algorithm has the lowest runtime. After 100 iterations, the runtimes of the benchmark algorithm and our algorithm are 44.33ms, 54.68ms, 90.8822ms, 43.94ms, 1226.24ms, and 14.45ms, respectively. This demonstrates our algorithm's superior runtime performance.

[0084] Figure 4A memory cost comparison is presented between the proposed algorithm and other benchmark algorithms. After 100 iterations, the proposed algorithm's cumulative memory cost is 322.35KB, while the benchmark algorithms one through five use 395.21KB, 358.20KB, 426.76KB, 448.05KB, and 422.07KB, respectively. This demonstrates that the proposed algorithm can accurately determine the true value on aggregated data and also offers advantages in terms of runtime and memory consumption.

[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0086] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for improving data collection quality based on worker task completion records and privacy protection, characterized in that: The following steps are involved: The system performs initialization operations, and the task platform assigns the worker and the data requester their own initial keys; The data demander publishes a data demand task T on the task platform, hoping to recruit N workers to complete the task; The task platform broadcasts task T to workers; Workers send requests to the task platform to apply for corresponding tasks; Workers upload relevant data and credibility certificates of tasks; The data demander obtains and decrypts the data submitted by the corresponding worker; Data demanders verify the worker's credit record and identity; Data demanders evaluate the data quality of workers and obtain target data; Data demanders update workers’ credit records; Workers verify the records updated by data requesters.

2. A method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: In the initialization phase, the task platform generates identity keys for workers and data demanders respectively. i , generate a private key k that satisfies elliptic curve encryption i and public key K i ; Then, suppose there is a polynomial ring f(y)=y n -1, Represents the set of all polynomials with integer coefficients whose highest degree does not exceed n-1, where n is a prime number. is a power of 2. The worker chooses Polynomial with coefficients 0, 1, -1 And stipulate satisfy is a system parameter equal to 3. Next, calculate the worker’s second private key Public Key Above, the worker's key pair set is represented as ({K i ,k i },{Pk i ,Sk i }). Similarly, according to the above method, the key pair set generated for the data demander is expressed as ({K D ,k D },{Pk D ,Sk D }).

3. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for workers to apply for tasks is to send task request information to the task platform <Pk i ,T>. After the data demander receives the worker's request information, if he chooses the corresponding worker to perform the task, he will do the following: First, randomly select G is the base point of the elliptic curve, and then the following calculation is performed: R = bG, R' = b'G, k T =H1(b·Pk i ), K T =k T G. Above, data demanders publish information on selecting workers on the task platform <T,K T ,R,R′>. The worker uses his personal private key Sk i Can verify whether they are selected to perform task T, that is, if the worker verifies that K T =H1(R·Sk i )G, it means that the worker knows that he has been selected to perform the task, and k T is the corresponding task key.

4. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The method and rules for workers to upload data and prove their reputation are as follows: the reputation of a worker is determined by the historical records of tasks completed by the worker. Each worker is assigned a set of triples A = (α h ,α u ,α l ), α h Represents the number of times a worker’s performance in completing tasks has been rated as excellent, α u represents the number of times a worker’s task completion performance was rated as average in the past, α l The number of times a worker's performance in completing tasks was rated as unsatisfactory. i A i When first recorded in the reputation center, the reputation management center generates an initial reputation certificate Pro (A i )=H1(H1(0),A i ), when a worker completes the task numbered T-1, the corresponding data demander evaluates the worker and then uses the relevant task key k T-1 In the new block update proof is Pro(A i )=H1(H1(k T-1 ),A i ). In task T, when the worker collects data Post-selection And use Pk i Encrypt the data to obtain Among them, k T-1 is the private key of the worker's last completed task, and | is the character concatenation operation. Subsequently, the worker selects Polynomials on satisfy Finally, the worker uses the public key Pk of the data demander D Get encrypted data and calculate In summary, the information submitted by workers on the task platform is 5. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for data demanders to decrypt data submitted by workers is as follows: the data demanders receive the submitted data and use Sk D Decryption get and calculate Among them Sk D satisfy Data demanders will Send to the task center to obtain the re-encryption key Rk iB The mission center receives After that, calculate And Rk iB Return to the data demander. The data demander uses Rk iB and Sk D Data Decryption is performed to obtain data and verify credibility. The decryption method is:

6. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for data demanders to verify the reputation of workers is: through decryption, data demanders obtain the data set of all workers on task T Using the block height information h, the data demander extracts the latest credit proof Pro (A i ), verified by Pro(A i )=H1(H1(k T-1 ),A i ) Determine whether the worker’s credibility is genuine.

7. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for data demanders to evaluate the quality of workers' data is as follows: using the ternary set A, the probability distribution space of workers' current task performance being excellent, average, and unqualified is predicted to be X = (X h ,X u ,X l ), where the probability distribution satisfies ∑ β∈{h,u,l} X β =1,0≤X β ≤1. It is known that Γ(·) is a Gamma function, and the weight parameter K is set to (k h ,k u ,k l ) is the factor that adjusts the probability of the distribution space, from which we can calculate It represents the possibility of predicting the three types of evaluations of workers for the current task using the relevant statistical distribution, which satisfies the following equation: Therefore, using The expected results under the posterior distribution are obtained by summing A, and Laplace smoothing is added to correct the expectation to prevent the problem of zero probability, and the expected results of the last three probability distributions of the workers are obtained. β∈{h,u,l}, which can be defined by the following equation: Next, calculate the uncertainty in the expected result Given the probability density function f(·) under the Dirichlet distribution, the uncertainty The expression can be described as: Combining the above process, we define a trust index Φ of a triple i =(φ h,i ,φ l,i ,φ u,i ) to evaluate the comprehensive reputation of workers, the credibility index Φ i and parameter set A, The following relationship is satisfied between K: f u,i =1-φ h,i -f l,i (5) Next, using the credibility index Φ i Perform truth discovery and use the approximate angle variance method to evaluate the quality of worker data. In each truth discovery task T, the worker group Submit perception data collection In the perception task, Can often be represented as a set of one-dimensional vectors In order to detect abnormal classes in sensor data, an abnormal data detection method based on ABOD is adopted to overcome the problem that traditional data detection relies on the distance similarity between data, which leads to frequent parameter changes and time-consuming traversal. The ABOD method can be expressed as: vector Given the paradigm ‖·‖: The scalar product between two vectors is <·,·>: for represents the difference between two vectors, Then the vector The angular variance can be expressed as: The variance of points inside the data set is often high, while the variance of outliers is very low. Setting reasonable thresholds μ1 and μ2 can Divide into interior point sets Boundary point set and the set of outliers The three correspond to high quality, uncertain quality and low quality evaluation respectively. Then, based on the evaluation index Φ, a set of workers with high reputation evaluation is selected. And in the collection Randomly select δ workers from the set to form a subset The data of high-quality workers are often located at the interior and boundary points, which have the largest calculation weights. It is stipulated that δ has a minimum number of people to ensure the lower bound of the approximate variance, so the vector The approximate angular variance of for: Using approximate angular variance The workers' data is classified by quality, and a weighted average process is performed using high-quality data sets to finally obtain the data required by the data demander.

8. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for data demanders to update worker credit records is as follows: after evaluating the data quality of each worker, data demanders obtain a new round of credit evaluation results A′ i =(α′ h,i ,α′ u,i ,α l ' ,i ), and construct a new reputation proof Pro(A′) for each worker i )=H1(H1(b·Pk i ),A′ i ). In order to effectively embed the reputation proof into the blockchain, the data demander constructs a ring signature for each worker to anonymize the worker's identity while verifying the legitimacy of the worker's new reputation proof. i Constructing a ring signature, the data demander combines the task public key randomly generated by the system with Pro(A′ i ), generate the public key group: L T ={K T,1 ,K T,2 ,...,H1(H1(k T-1 ),A′ π )G,...,K n }, true credibility certificate Pro(A′ i ) is hidden in the πth position. Using the public key group L T and H1(H1(k T ),A′ π ), in the reputation management center, the data demander selects k T,π =H1(H1(k T-1 ),A′ π ),}. In this way, the data demander generates a random message m T,i , using the generation algorithm GEN(·) to generate a new ring signature σ for the worker T,i . The data demander will sign the ring T,i Sent to the reputation management center, the reputation management center starts the reputation update contract. First, the blockchain miner on the reputation management center receives σ T,i Then, using the verification algorithm VER(m T,i ,σ T,i ) Verify the legitimacy of the group signature. Finally, L T Stored on the new reputation proof block, each public key of the public key group L is stored at the corresponding height h j , j=1,2,...,π,...,n, but only in h π The public key on the worker N is i Proof of legal credibility.

9. The method for improving data collection quality based on worker task completion records and privacy protection according to claim 1, characterized in that: The way for workers to verify the records updated by data demanders is as follows: after the data demander completes the credit update for the worker, select Calculate D = b′ + b″k D , send information to workers The worker decrypts the information and verifies it. The verification is divided into three aspects: First, verify DG = R' + b"K D Is it true? If so, it proves that the credit update is done by the corresponding data demander. β∈{h,u,l} α′ β,i -α β,i =1 to prove that the data demander has correctly updated the interaction record. Finally, verify H1(H1(k T ),A′ π )=H1(H1(H1(R·Sk π )),A′ π ) to prove that the credit certificate is correct and legal. After successful verification, the worker can obtain the corresponding task reward and use the new round of credit certificate to apply for the next round of tasks.