Trust-based data verification method and system in crowd sensing

By introducing social exchange theory and localized differential privacy into mobile crowd sensing, and combining the Mann-Whitney U test and locality-sensitive hashing algorithm, the problems of new user trust assessment and malicious data filtering are solved, achieving efficient and reliable data verification and privacy protection in cold start scenarios.

CN121580445BActive Publication Date: 2026-04-28GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-01-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In mobile crowd sensing, existing technologies struggle to effectively assess the reliability of new users and protect user privacy in cold start scenarios. They also cannot effectively defend against collusion attacks by malicious users, resulting in inconsistent data quality and misleading system judgments.

Method used

The trust assessment process is modeled as a social exchange process with three dimensions: historical contribution, privacy utility, and data similarity. Localized differential privacy processing of user data is adopted, and the Mann-Whitney U test and locality-sensitive hashing algorithm are combined to filter malicious data, thereby achieving dynamic assessment of trust values ​​and efficient filtering of false data.

Benefits of technology

While protecting user privacy, it effectively filters out false data uploaded by malicious users in collusion, improves data quality and system reliability, reduces the impact of collusion attacks, and adapts to trust assessment and data verification in large-scale and complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580445B_ABST
    Figure CN121580445B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data inspection, and discloses a trust-based data inspection method and system in crowd wisdom perception, wherein the trust evaluation process is modeled as a social exchange process in three dimensions of historical contribution, privacy utility and data similarity, the exchange parties are a platform and a user, and the cold start problem is solved to a certain extent; the user data is processed by using localized differential privacy, and the similarity of the user data is quantified, so that a trade-off is made between privacy protection and trust evaluation; the consistency of the uploaded data of the user is combined with Mann-Whitney U inspection, local sensitive hashing (LSH) is used to improve the inspection efficiency, false user data is filtered, and the accuracy of the task result is improved. Finally, through theoretical analysis and experimental verification, it is proved that the TDA scheme is feasible and effective when completing the perception task, and can effectively filter the false data uploaded by malicious users in collusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data verification technology, and in particular relates to a trust-based data verification method and system in crowd intelligence perception. Background Technology

[0002] In mobile crowdsourcing sensing, ordinary users can participate in data collection as sensing nodes, leveraging the widespread availability of mobile smart devices and internet connectivity. This creates a low-cost, large-scale, and diverse data source with broad applications in fields such as environment, transportation, and healthcare. However, due to the difficulty in standardizing user device status and motivation, the quality of submitted data often varies. To ensure data reliability, sensing platforms typically employ trust assessment models to screen trustworthy users.

[0003] Existing research largely relies on user history records for trust assessment. For example, in vehicular networks, this involves integrating entity and data trust to improve message propagation efficiency, or using trust values ​​to select reliable lead vehicles. Additionally, some methods combine privacy-preserving technologies to reduce communication overhead while maintaining user privacy. However, for newly joined users (cold start scenarios), the lack of historical data makes it difficult to directly assess their reliability. While some studies have introduced drones to collect baseline data or utilized recommended trust, these methods are costly and susceptible to collusion attacks. Another key issue is that existing trust assessment models often do not adequately consider user privacy preferences. In reality, assessment accuracy often depends on users providing more information (such as location and time), but most solutions do not allow users to choose their level of privacy exposure, and related research remains scarce. Furthermore, platforms often employ redundancy strategies to expand data collection, but the increased number of users also introduces the risk of collusion attacks. Malicious users may collude to submit false data, interfering with system judgment, such as falsifying epidemic data in health monitoring, misleading decision-making, and wasting resources. To address this, existing research has proposed methods based on incentive mechanisms, truth-finding algorithms, and privacy-preserving aggregation to identify and resist collusion behavior, or to reduce the likelihood of collusion from the task allocation stage.

[0004] While numerous studies have been conducted on trust assessment and collusion attacks, most focus on a single perspective and lack in-depth exploration of their combined impact. In real-world systems, user reliability and collusion attacks often occur in tandem. Therefore, it is necessary to balance privacy protection and reliability assessment under cold start conditions, and to effectively suppress the impact of collusion on data quality after the assessment.

[0005] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0006] (1) In practice, user reliability and collusion attacks are often intertwined, so it is necessary to consider the problem from two dimensions. How to make a reasonable judgment on user reliability in the cold start scenario and protect user privacy; after assessing user reliability, how to reduce the impact of collusion attacks on data quality. (2) In crowd intelligence perception, the rationality of traditional trust assessment methods is difficult to guarantee when there is a lack of user history records, and trust assessment can only indirectly infer user reliability and cannot avoid collusion deception by malicious users. Summary of the Invention

[0007] To address the problems existing in the prior art, this invention provides a trust-based data verification method for crowd intelligence sensing.

[0008] This invention is implemented as follows: a trust-based data verification method in crowd intelligence sensing includes:

[0009] Step 1: Model the trust assessment process as a social exchange process with three dimensions: historical contribution, privacy utility, and data similarity. The two parties in the exchange are the platform and the user, thus solving the cold start problem.

[0010] Step 2: Use localized differential privacy processing to process user data and quantify the similarity of user data, making a trade-off between privacy protection and trust assessment;

[0011] Step 3: A user data consistency determination method based on the Mann-Whitney U test is adopted, and the locality-sensitive hashing algorithm is combined to improve the test efficiency, thereby filtering out false or collusive data uploaded by malicious users;

[0012] Step 4: Through theoretical analysis and experimental verification, the feasibility and effectiveness of the TDA solution in completing the perception task are demonstrated, effectively filtering out false data uploaded by malicious users in collusion.

[0013] Furthermore, the trust assessment includes: S1, historical contribution; S2, privacy utility; S3, data similarity; and S4, trust fusion.

[0014] Furthermore, the aforementioned historical contributions:

[0015] Historical contribution value According to user The historical data is used to calculate a user's historical contribution value, which quantifies their long-term exchange behavior. This historical contribution value is then compared with the user's integrity value. and the user's historical activity value Closely related, integrity score and activity score reflect a user's reliability and initiative in completing historical tasks; a user's integrity score It is the user's historical contribution value A core component of this system reflects the reliability of a user's fulfillment of their "data quality commitment." Generally, more trustworthy users will make a greater contribution to the task. Therefore, a user's trustworthiness score is derived from the proportion of reliable data submitted in historical tasks, calculated using the following formula:

[0016] ,

[0017] in, The total number of times tasks were completed for the user; The number of times reliable data is submitted to users; For growth control parameters;

[0018] User's integrity score The calculation is divided into integrity growth and penalties for dishonesty; integrity growth It is the positive cumulative effect of users submitting reliable data, which controls the growth of the integrity score; This represents the cumulative effect of reliable data submitted by users. This represents the total number of tasks a user participates in, serving as a normalization factor; if a malicious user wants to improve their integrity score, they need to complete a large number of tasks. ),but It will inhibit the growth of its integrity value; in terms of integrity growth, the initial growth is fast and the later growth is slow, which is in line with the security requirement that trust needs to be accumulated over a long period of time. As a form of penalty for breach of trust, no penalty is imposed when a user submits reliable data for every task; however, even if a malicious user submits a large amount of data rapidly in a short period, penalties will be imposed if the failure rate is low. If the rating is too high and violates the exchange rules of trust, the platform will significantly reduce its trust score.

[0019] Historical activity value As one of the reference indicators for historical contributions, user activity value represents whether users can respond promptly to task requests broadcast by the platform and actively complete perception tasks. The calculation of activity value adopts a two-level evaluation system, with single task activity value... Compared with historical activity value Single task activity value Evaluate users In specific tasks Timeliness of data; historical activity value This is a comprehensive assessment of recent task performance, serving as the final result of user activity value; [The user's activity value is then used to evaluate their performance across multiple tasks.] Upload task The time of the data is recorded by the user. For the task The platform will record the completion time of each user's task. The task can be derived. The median completion time of all users is denoted as the task. base time By comparing the difference between two time periods, we can determine the user's... Single activity value The calculation formula is:

[0020] ,

[0021] in, This refers to the time dispersion, which is the median absolute deviation of the task completion time. It is a positive part function;

[0022] In single activity value In the calculation process, the platform only penalizes timeouts. ;when hour, ,when hour, It decays exponentially, and each time it exceeds time, decline ;

[0023] Historical activity value Based on recent user activity The performance of this task is calculated using the following formula:

[0024] ,

[0025] in, , As time decays, the weight of more recent tasks increases. The sliding window mechanism simulates the recent preference effect to prevent users' historical trustworthy behavior from masking recent malicious behavior.

[0026] Through triple verification of user reliability, timeliness, and stability, fine-grained evaluation of user historical behavior is achieved, ultimately resulting in a historical contribution value. The calculation formula is:

[0027] ,

[0028] in, These are the weighting coefficients. It is used to balance user reliability and response speed.

[0029] Furthermore, the privacy benefits mentioned:

[0030] Localized differential privacy technology is used to protect user data privacy and security, assuming that the user... The submitted data is The platform wants to query a certain function. (e.g., user's location coordinates), towards Add Laplace noise to satisfy , It is noise sampled following a Laplace distribution; malicious users may frequently increase their privacy budget in the short term to accumulate privacy utility values, so this invention divides privacy utility into short-term utility and long-term utility; short-term utility value By user Privacy budget per task The calculation shows that:

[0031] ,

[0032] Long-term utility value Exponential decay is used to suppress short-term speculative behavior and ensure the fairness of long-term trust exchange. The calculation formula is as follows:

[0033] ,

[0034] in, It is an adjustment parameter that controls the rate of increase in short-term utility value; It is the attenuation factor. ;

[0035] If a user has no history, the user's privacy utility value is the short-term utility value. If a user has a history, then the user's privacy utility value is the long-term utility value. ;user The formula for calculating privacy utility value is:

[0036] .

[0037] Furthermore, the data similarity:

[0038] Group consensus is generated from user data with historical contributions, and long-term contributors should receive higher weight. Specifically, users are assigned different weights based on their historical trust values, so that users with high trust values ​​have a greater influence on group consensus, that is, regulating individual behavior through normative pressure. Assuming that in the user set... In the middle, there is If a user has historical contributions, then the weighted group consensus value is... The calculation formula is:

[0039] ,

[0040] in, This indicates the user's final trust value, which they can choose to follow group exchange rules to obtain.

[0041] When new users with high privacy sensitivity join, they can choose to comply with group exchange norms to gain trust; data similarity. The calculation formula is:

[0042] ,

[0043] in, yes and The Euclidean distance between them It is the weighted group standard deviation, calculated using the following formula:

[0044] .

[0045] Furthermore, the trust fusion:

[0046] Through trust exchange across three dimensions—historical contribution, privacy utility, and data similarity—a weighted fusion is used to derive a linear trust value for the user. The calculation formula is:

[0047] ,

[0048] in, , , Let be the weighting coefficient, satisfying ,

[0049] This represents the user's privacy utility value;

[0050] Directly use linearly weighted trust values Using a linearly weighted trust score as a user's trust value presents several problems: it struggles to capture the marginal effect of trust and cannot suppress the influence of extreme values; furthermore, malicious users may fabricate high scores in a single dimension, exhibiting short-term speculative behavior. Therefore, a sigmoid function is used to non-linearly map the linearly weighted trust value before outputting the final user trust score. The calculation formula is:

[0051] .

[0052] Another object of the present invention is to provide a trust-based data verification system for crowd intelligence perception, comprising:

[0053] The exchange module is used to model the trust assessment process as a social exchange process with three dimensions: historical contribution, privacy utility, and data similarity. The two parties in the exchange are the platform and the user, which solves the cold start problem.

[0054] The filtering module is used to process user data with localized differential privacy and quantify the similarity of user data, making a trade-off between privacy protection and trust assessment; it combines the Mann-Whitney U test to verify the consistency of user-uploaded data and uses Locality Sensitive Hash (LSH) to improve the test efficiency and filter out fake user data.

[0055] The analysis and verification module is used to prove the feasibility and effectiveness of the TDA solution in completing perception tasks through theoretical analysis and experimental verification, and to effectively filter out false data uploaded by malicious users in collusion.

[0056] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the trust-based data verification method in the crowd perception.

[0057] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the trust-based data verification method in the crowd sensing.

[0058] Another objective of this invention is to provide an information data processing terminal for implementing a trust-based data verification system in the collective intelligence perception.

[0059] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0060] This invention proposes a Trust-Based Data Quality Assessment Mechanism in Crowdsensing Services (TDA), which consists of two parts: trust assessment and data verification. The invention models the trust assessment process as a social exchange process encompassing three dimensions: historical contribution, privacy utility, and data similarity. The exchange partners are the platform and the user, thus addressing the cold start problem to some extent. Localized differential privacy processing of user data and quantification of data similarity strike a balance between privacy protection and trust assessment. The Mann-Whitney U test is used to verify the consistency of user-uploaded data, and Locality Sensitive Hashing (LSH) is employed to improve verification efficiency, filter out fraudulent user data, and enhance the accuracy of task results. Finally, theoretical analysis and experimental verification demonstrate the feasibility and effectiveness of the TDA scheme in completing perception tasks, effectively filtering out fraudulent data uploaded by malicious users.

[0061] This invention proposes a trust-based data quality assessment (TDA) mechanism. Compared with existing work, the contributions of this invention are as follows:

[0062] (1) This invention, from the perspective of social exchange theory, regards trust assessment in collective intelligence perception as a social exchange process and models it as an exchange relationship in three dimensions: historical contribution, privacy utility, and data similarity. This alleviates the limitations of traditional methods that rely on user history records and inhibits the speculative behavior of malicious users to quickly gain platform trust.

[0063] (2) In the process of trust exchange, localized differential privacy (LDP) is combined with trust assessment, and the concept of data similarity is introduced to provide users in the collective intelligence perception with a trust exchange channel that meets their needs.

[0064] (3) Design a method to deal with malicious user collusion attacks in the group intelligence perception system, use trust value to eliminate users with low reliability, combine Mann-Whitney U test to verify the consistency of user data, use local sensitive hashing to improve the efficiency of data verification, and filter false data uploaded by malicious users in collusion.

[0065] This invention proposes a trust-based data verification scheme. This novel scheme comprehensively considers the impact of multiple factors on trust assessment, reducing the influence of the cold start problem on the accuracy of trust assessment to a certain extent. Simultaneously, it incorporates localized differential privacy protection technology to effectively protect user privacy, using trust values ​​as the basis for filtering high-quality data and combining consistency checks to filter out fraudulent data uploaded by malicious users in collusion. To evaluate TDA, experiments were conducted on a simulated dataset and compared with the DTI and RPS methods. Experimental results show that TDA can defend against collusion attacks by malicious users in cold start scenarios. The dataset used in this invention was generated under relatively ideal conditions. Due to the difficulty in obtaining real experimental data, the experiments in this invention have not yet been further validated using real experimental datasets. Future research will seek suitable datasets to verify the effectiveness of TDA.

[0066] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:

[0067] With the widespread adoption of mobile internet and smart terminals, crowdsourcing sensing technology is gradually moving towards commercial applications. This solution overcomes key technical bottlenecks in data quality assurance, user privacy protection, and malicious behavior defense, laying the foundation for large-scale commercial applications and effectively solving the problems of trust establishment and data reliability in open environments. The solution demonstrates excellent performance in data reliability and system security. Theoretical analysis and experimental verification show that even in complex scenarios involving a large number of new and malicious users, it can maintain reasonable computational complexity and response time. The system exhibits good robustness; even under attacks with up to 50% malicious user collusion, the data verification accuracy and fake data filtering rate remain at a high level, ensuring practicality and reliability in real-world business environments. The commercial value of this solution extends beyond optimizing the data quality and efficiency of existing crowdsourcing sensing applications; it also has the potential to foster new business models and service formats. For platform operators, based on three-dimensional trust assessment and efficient data verification capabilities, a more trustworthy and efficient sensing ecosystem can be built, improving data quality and enhancing market competitiveness. For SMEs and developers, this technology can lower development barriers and costs, helping to quickly build data applications with trust guarantees. For end users, localized differential privacy technology can enhance the user experience while protecting personal privacy. Furthermore, user behavior data and trust profiles generated based on multi-dimensional trust assessments also have derivative value, which can be used in areas such as credit evaluation, personalized services, or risk control, driving the development of data value-added services.

[0068] The widespread application of this technology is expected to accelerate progress in fields such as smart cities, the Internet of Things, mobile health, and environmental monitoring, creating new economic growth points and industrial ecosystems. By establishing a reliable data exchange mechanism, it will promote the efficient flow and value release of data elements, providing crucial support for the development of the digital economy.

[0069] (2) This invention systematically solves the core problems that have long restricted the development of the industry in the field of crowd intelligence perception, such as trust assessment in cold start scenarios, balance between privacy protection and trust assessment, and defense against malicious user collusion attacks.

[0070] Addressing the pain points of traditional trust assessment schemes' over-reliance on user history and reliability failure in cold start scenarios, this invention breaks through the single-dimensional assessment framework and creatively constructs a three-dimensional social exchange model of "historical contribution - privacy utility - data similarity": It quantifies historical contributions through integrity value (integrating positive cumulative effects and a penalty mechanism for dishonesty, with rapid initial growth followed by slower growth, conforming to the long-term accumulation pattern of trust) and activity value (a two-level assessment system + sliding window recent preference effect, suppressing "historical credibility masking recent malice"), providing precise trust anchors for users with records. Simultaneously, it designs a dual-path trust acquisition channel for new users—users with low privacy sensitivity can increase their privacy budget to obtain short-term utility value (Laplace noise perturbation ensures privacy security), while users with high privacy sensitivity can establish initial trust based on data similarity (based on weighted group consensus of high-trust users, Euclidean distance + weighted standard deviation normalization). This completely solves the industry-wide common problem of "new users without historical records being unable to be assessed," filling the technical gap in multi-dimensional trust initialization in cold start scenarios.

[0071] Addressing the technical shortcomings of existing solutions that neglect user privacy preferences and fail to balance privacy protection with the accuracy of trust assessment, this invention innovatively integrates Localized Differential Privacy (LDP) with trust assessment. It achieves decentralized privacy protection for user data through Laplace noise injection (meeting the ε-localized differential privacy definition, satisfying privacy and security requirements at both the individual user and system levels). Simultaneously, it designs a dynamic calculation mechanism for short-term utility (driven by privacy budget) and long-term utility (exponential decay suppressing short-term speculation), coupled with a data similarity quantification method, establishing a dynamic balance between the degree of privacy exposure and the accuracy of trust assessment. Users can autonomously choose the intensity of their privacy investment, and the platform can achieve effective trust determination based on the perturbed data. This fills the technical gap in "coordinated optimization of privacy protection and trust assessment," breaking the industry perception that "privacy security and assessment accuracy are mutually exclusive."

[0072] To address the shortcomings of traditional solutions that rely solely on defense against collusion attacks and struggle to handle large-scale fake data clusters, this invention constructs a three-layer defense system: "Trust Pruning - LSH Acceleration - U-Check Verification." First, preliminary pruning is achieved through descending trust value sorting and intra-group threshold filtering (eliminating low-trust data and reducing invalid computation). Then, Locality Sensitive Hash (LSH) is used to construct hash buckets, compressing the data consistency check scope from the entire dataset to data within the same bucket, resulting in an exponential improvement in computational efficiency. Finally, Mann-Whitney... The U-Nonparametric test (which requires no assumptions about data distribution and is robust to identifying fake data clusters) verifies data consistency. Combined with the greedy solution of the maximum consistent subset across groups (an efficient approximation of NP-hard problems, ensuring that the proportion of reliable data exceeds 50%), the fake data filtering rate (CDR) is significantly improved compared with traditional DTI and RPS methods in scenarios where the proportion of malicious users does not exceed 50%. At the same time, the running time for city-level user scales (more than 1,000 people) is controlled within a reasonable range, filling the technical gap of "efficient anti-collusion attack in large-scale collective intelligence perception scenarios" and solving the industry dilemma of "the simultaneous increase in redundant data and the risk of collusion attacks".

[0073] Furthermore, at the technological paradigm level, this invention establishes for the first time a three-in-one data verification framework for crowd-based intelligent sensing, integrating social exchange theory, statistical testing, and privacy computation. It uses social exchange theory as the foundation for trust modeling (satisfying long-term reciprocity, rational choice, and group normativity), nonparametric statistical testing as the core of data verification, and localized differential privacy as the support for privacy protection, forming a closed-loop management system from trust assessment to data verification. This framework not only enables breakthroughs in three key indicators—cold-start adaptability, privacy security, and attack robustness—for crowd-based intelligent sensing systems, but also provides reusable trust-privacy-security collaborative solutions for distributed collaborative scenarios such as the Internet of Things and edge computing. This marks a significant turning point for crowd-based intelligent sensing technology, moving from "single-function optimization" to "system-level trust assurance," further filling the theoretical and practical gaps in "multi-technology integration driving the trustworthiness of crowd-based intelligent sensing" both domestically and internationally.

[0074] (3) The technical solution of the present invention has precisely overcome the three core technical problems that people have long wanted to solve in the field of crowd intelligence perception, but have never been able to achieve success. It has achieved a key leap from breaking through technical bottlenecks to practical application and provided a disruptive solution for industry development.

[0075] First, it successfully solves the long-standing problem of "failure of new user trust assessment in cold start scenarios." In crowdsourced sensing systems, new users lack historical behavior records. Traditional solutions either rely on indirect inference based on recommendation trust (vulnerable to malicious collusion attacks) or require additional deployment of drones to collect baseline data (significantly increasing costs). A balance between lack of historical records, low cost, and high reliability has always been difficult to achieve, resulting in high barriers to new user participation and limited system expansion. This invention innovatively constructs a three-dimensional social exchange model of "historical contribution - privacy utility - data similarity," designing a dual-path trust acquisition channel for new users: users with low privacy sensitivity can increase their privacy budget (combined with localized differential privacy technology to ensure security) to obtain short-term utility values, while users with high privacy sensitivity can establish initial trust based on the weighted group consensus with highly trusted users. This solution can achieve reasonable trust assessment of new users without relying on historical records or additional hardware. Second, it overcomes the industry-wide problem of "difficulty in coordinating privacy protection and trust assessment accuracy." For a long time, the field of crowdsourced sensing has faced a contradiction: "the more privacy is exposed, the more accurate the trust assessment." Traditional solutions either ignore user privacy preferences and force data collection (leading to user resistance) or overprotect privacy, resulting in decreased data usability (distorted trust assessment). They have consistently failed to provide users with a privacy-trust balance mechanism that allows for autonomous choice. This invention deeply integrates Localized Differential Privacy (LDP) with trust assessment. It achieves decentralized protection of user data through Laplace noise injection (satisfying the ε-localized differential privacy definition). Simultaneously, it designs a computational mechanism for short-term utility (dynamically adjusted with the privacy budget) and long-term utility (exponential decay suppressing speculative behavior). Combined with a data similarity metric, this allows users to autonomously choose the degree of exposure based on their privacy preferences, while the platform can achieve accurate trust assessment based on the perturbed data. Thirdly, it solves the technical challenge of "low efficiency in defending against malicious user collusion attacks in large-scale scenarios." As the user base of crowdsourced sensing expands, the risk of malicious users colluding to submit false data (such as exaggerating epidemic data or falsifying environmental monitoring information) increases dramatically. Traditional solutions either rely on historical data verification (ineffective against new colluding users) or employ full data consistency checks (high computational complexity and unacceptable latency), consistently failing to achieve a balance between defense effectiveness and system efficiency. This invention constructs a three-layer defense system: "trust pruning - LSH acceleration - U-test verification." First, low-trust data is filtered and eliminated using trust values ​​(reducing computational load). Then, Locality Sensitive Hash (LSH) is used to compress the verification scope to data within the same bucket (improving efficiency). Finally, the Mann-Whitney U nonparametric test verifies data consistency (no assumptions about data distribution are required, resulting in strong anti-interference capabilities).

[0076] In summary, this invention addresses the three major technical challenges that have long existed in the field of crowd intelligence sensing: cold start trust assessment, privacy-trust collaboration, and large-scale collusion defense. It proposes a systematic solution, breaks through the technical bottlenecks in the industry's development, and marks a key turning point for crowd intelligence sensing technology from theoretical exploration to large-scale practical application, meeting the industry's long-term needs for trustworthiness, efficiency, and privacy protection.

[0077] (4) The technical solution of the present invention has significantly overcome the three major technical problems that have long existed in the field of crowd intelligence sensing technology, and promoted the industry to move from a single and static technical understanding to a multi-dimensional and dynamic system design paradigm.

[0078] First, it overcomes the technical bias that "trust assessment must rely on historical records." Traditional solutions generally believe that user trust values ​​must be calculated based on long-term historical behavioral data. This leads to new users (cold start scenarios) being unable to be reasonably assessed due to a lack of records, either being directly excluded from the system or relying on unreliable indirect recommendations, creating the cognitive limitation of "no historical record equals no trust." This invention breaks through this bias by innovatively constructing a three-dimensional trust assessment model of "historical contribution - privacy utility - data similarity": for new users with no historical records, those with low privacy sensitivity can increase their privacy budget (combined with localized differential privacy technology) to obtain short-term utility values, while those with high privacy sensitivity can establish initial trust based on the similarity between their data and the consensus of a high-trust user group. Second, it breaks the technical bias that "privacy protection and trust assessment accuracy are mutually exclusive." For a long time, the industry has generally believed that "the more user privacy is exposed, the higher the data availability, and the more accurate the trust assessment," leading to solutions that either sacrifice privacy for assessment accuracy or overprotect privacy, resulting in distorted trust judgments, creating a cognitive misconception that the two are contradictory. This invention overturns this prejudice by deeply integrating Localized Differential Privacy (LDP) with trust assessment: It achieves decentralized privacy protection for user data through Laplace noise injection (satisfying the ε-localized differential privacy definition), while designing a computational mechanism for short-term utility (dynamically adjusted with the privacy budget) and long-term utility (exponential decay to suppress speculative behavior). Combined with a data similarity quantification method, it allows users to autonomously choose their level of privacy exposure, and the platform can still achieve accurate trust assessment based on the perturbed data. Finally, it overcomes the technical prejudice that "anti-collusion attacks must come at the cost of system efficiency." Traditional anti-collusion solutions either employ full data consistency checks (high computational complexity and unacceptable latency) or rely on historical data verification (ineffective against new colluding users), leading to the cognitive limitation that "defending against attacks requires sacrificing efficiency." This invention overcomes this bias by constructing a three-layered, highly efficient defense system: "trust pruning - LSH acceleration - U-test verification." First, low-trust data is filtered and eliminated using trust values ​​(reducing invalid computation). Then, Locality Sensitive Hash (LSH) is used to compress the verification scope to data within the same bucket (improving matching efficiency). Finally, Mann-Whitney U nonparametric verification is used to verify data consistency (strong anti-interference capability). This achieves a win-win situation of effectiveness against collusion attacks and system efficiency, breaking the technical bias that "defense against attacks inevitably sacrifices efficiency." These innovations not only realize a paradigm shift in crowdsourced sensing technology from "single historical dependence" to "multi-dimensional trust assessment," from "privacy versus accuracy" to "collaborative optimization," and from "defense versus efficiency" to "win-win balance," but also provide a complete technical solution for building a trustworthy, efficient, and privacy-protected crowdsourced sensing system adapted to real-world open environments, propelling industry technology awareness and practical application to new heights. Attached Figure Description

[0079] Figure 1 This is a flowchart of a trust-based data verification method in crowd intelligence perception provided in an embodiment of the present invention.

[0080] Figure 2 This is a block diagram of a trust-based data verification system in crowd intelligence perception provided in an embodiment of the present invention.

[0081] Figure 3 This is an accuracy graph provided by an embodiment of the present invention for different proportions of new users.

[0082] Figure 4 This is a CDR diagram under different proportions of colluding users provided in the embodiments of the present invention.

[0083] Figure 5 This is a MAE diagram under different proportions of colluding users provided in the embodiments of the present invention.

[0084] Figure 6 This is a runtime diagram for different user scales provided in the embodiments of the present invention.

[0085] Figure 7 This is a graph showing the impact of privacy effectiveness and data similarity on TDA under different new user ratios, as provided in the embodiments of the present invention.

[0086] Figure 8 This is a diagram showing the impact of LSH on TDA under different proportions of colluding users, as provided in an embodiment of the present invention. Detailed Implementation

[0087] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0088] The core challenge faced by crowdsourced sensing systems in typical application scenarios such as smart cities, intelligent transportation, and environmental monitoring lies in the inconsistent quality of raw data received by the platform due to the surge in mobile terminals and the increasing complexity of sensing tasks. This data includes both reliable, real observations and noisy, maliciously forged information. Traditional reputation scoring models are mostly based on single-dimensional statistics or linear accumulation of historical behavior, lacking comprehensive capabilities to address dynamic trust evolution, privacy protection, and collusion attacks. This problem is particularly prominent in large-scale distributed crowdsourced networks; the inability to effectively identify false data directly impacts the reliability of decision-making models and the accuracy of subsequent data-driven tasks. Therefore, a technical approach is needed that can achieve reliable data filtering and dynamic trust assessment while protecting user privacy, thereby improving the overall effectiveness of crowdsourced sensing tasks.

[0089] In practical system deployments, the heterogeneity of data sources and the randomness of tasks prevent platforms from relying solely on global models or static parameters for trust determination. Therefore, this method abstracts the trust assessment process into a social exchange model. The interaction between the platform and users no longer depends solely on a single contribution record, but simultaneously considers three dimensions: long-term historical contributions, privacy utility, and current data similarity. This modeling approach transforms trust from a function with fixed weights into a dynamically adjusted quantitative result that varies with time, environment, and individual behavior. In this way, the system can assign initial trust values ​​to new users during the cold start phase and gradually adjust them during ongoing tasks, achieving the coupling of trust accumulation and malicious detection.

[0090] At the data privacy level, this method introduces a localized differential privacy mechanism, utilizing Laplace distribution noise perturbation to achieve obfuscated publication of individual user data. This process mathematically guarantees the irreversibility of sensitive individual information while allowing the platform to recover group distribution characteristics through statistical methods. By dynamically adjusting the control parameters of the privacy budget, a balance is achieved between the strength of privacy protection and data availability. When the privacy budget is too low, data authenticity decreases; when the budget is too high, the risk of leakage increases. The system uses an exponential decay function combining short-term and long-term utility to penalize behaviors that frequently increase the privacy budget, maintaining a long-term stable trust exchange order.

[0091] In the data consistency assessment stage, this method employs the Mann-Whitney U nonparametric test to compare the rank differences between user-uploaded data and group consensus data, thus avoiding misjudgments caused by distributional hypothesis bias. This type of rank statistical test is robust in environments with limited sample sizes and unknown distributions, effectively identifying data deviations from trends. Simultaneously, by using a locality-sensitive hashing algorithm to bucket the high-dimensional data space, the testing efficiency is significantly improved, reducing the computational complexity of large-scale data comparisons. Under this mechanism, the platform can quickly locate similar sample clusters and remove outliers, thereby achieving real-time filtering of fraudulent or collusive uploaded data.

[0092] The trust fusion process theoretically integrates multi-dimensional trust values ​​through a combination of linear weighting and nonlinear mapping. The linear weighting determines weights based on the importance of each dimension to ensure model interpretability; while the introduction of the nonlinear sigmoid function guarantees the controllability of the marginal trust effect—that is, trust growth slows in the high trust range and the penalty effect intensifies in the low trust range. This mechanism prevents malicious users from masking overall risk with high scores in a single dimension, maintaining the system's sensitive response to abnormal behavior. The final trust value serves as a key parameter in task scheduling and reward allocation, enabling bidirectional regulation of incentives and constraints.

[0093] When this method is run on a crowdsourced sensing platform, it forms a closed-loop trust evolution and data verification mechanism. The system constructs a feedback loop between data collection, privacy perturbation, trust updates, and anomaly detection, ensuring dynamic coordination among user behavior, task data, and platform decisions. Theoretical derivation and experimental results show that this method can significantly improve the false data identification rate and system robustness. Especially in environments with a high proportion of colluding nodes, it can still maintain high task completion accuracy and trust distribution stability, providing technical support for the reliable operation of crowdsourced sensing in smart cities and mobile internet applications.

[0094] like Figure 1 As shown in the figure, a trust-based data verification method in crowd intelligence perception provided by an embodiment of the present invention includes the following steps:

[0095] S101 models the trust assessment process as a social exchange process with three dimensions: historical contribution, privacy utility, and data similarity. The two parties in the exchange are the platform and the user, thus solving the cold start problem.

[0096] S102 utilizes localized differential privacy processing of user data and quantifies the similarity of user data, striking a balance between privacy protection and trust assessment; it combines the Mann-Whitney U test to verify the consistency of user-uploaded data and uses Locality Sensitive Hash (LSH) to improve the efficiency of the test and filter out fake user data.

[0097] S103, through theoretical analysis and experimental verification, proves the feasibility and effectiveness of the TDA solution in completing perception tasks, effectively filtering false data uploaded by malicious users in collusion.

[0098] like Figure 2 As shown in the figure, an embodiment of the present invention provides a trust-based data verification system for crowd intelligence perception, comprising:

[0099] The exchange module is used to model the trust assessment process as a social exchange process with three dimensions: historical contribution, privacy utility, and data similarity. The two parties in the exchange are the platform and the user, which solves the cold start problem.

[0100] The filtering module is used to process user data with localized differential privacy and quantify the similarity of user data, making a trade-off between privacy protection and trust assessment; it combines the Mann-Whitney U test to verify the consistency of user-uploaded data and uses Locality Sensitive Hash (LSH) to improve the test efficiency and filter out fake user data.

[0101] The analysis and verification module is used to prove the feasibility and effectiveness of the TDA solution in completing perception tasks through theoretical analysis and experimental verification, and to effectively filter out false data uploaded by malicious users in collusion.

[0102] Specific implementation of the present invention:

[0103] 1. This invention proposes a trust-based data quality assessment (TDA) mechanism. Compared with existing work, the contributions of this invention are as follows:

[0104] This invention, from the perspective of social exchange theory, regards trust assessment in collective intelligence perception as a social exchange process, and models it as an exchange relationship in three dimensions: historical contribution, privacy utility, and data similarity. This alleviates the limitations of traditional methods that rely on users' historical records and inhibits speculative behavior by malicious users to quickly gain the trust of the platform.

[0105] In the trust exchange process, Localized Differential Privacy (LDP) is combined with trust assessment, and the concept of data similarity is introduced to provide users in the crowd intelligence perception with a trust exchange channel that is adapted to their needs.

[0106] In a crowd-sensing system, a method is designed to counter malicious user collusion attacks. This method utilizes trust values ​​to eliminate users with low reliability, combines Mann-Whitney U checks to verify the consistency of user data, employs locality-sensitive hashing to improve the efficiency of data verification, and filters out false data uploaded by malicious users in collusion.

[0107] 2. Prerequisites

[0108] 2.1 Social Exchange Theory

[0109] Social Exchange Theory (SET) posits that social behavior is essentially a process of resource exchange, where individuals expect to gain social benefits through interaction. Unlike economic exchange, SET focuses on intangible resources (such as trust and emotional support). SET satisfies the following principles:

[0110] Definition 1. Long-term reciprocity

[0111] Interactive behaviors in social exchange are based on the expectation of reciprocal returns. That is, after one party provides resources (such as trust), the other party needs to reciprocate through corresponding behaviors (such as cooperation) in the long-term relationship; otherwise, the relationship may become unbalanced or terminate.

[0112] Definition 2. Rational Choice

[0113] Individuals or organizations make decisions in exchanges through cost-benefit analysis to maximize their own interests (economic or non-economic), and trust may also drive choices.

[0114] 2.2 Localized Differential Privacy

[0115] Localized differential privacy (LDP) is commonly used in privacy protection research for crowdsourcing awareness, allowing users to process data according to their own privacy budget. The definition of localized differential privacy is as follows:

[0116] Definition 3. Localized Differential Privacy

[0117] If it exists There are 10 users, and each user corresponds to one data point. For a random algorithm... The domain is The range is ,like In any two different data and ( The same output result is obtained on ), and the following conditions are met:

[0118]

[0119] At this time, the random algorithm satisfy -Localized differential privacy, It's a privacy budget.

[0120] Definition 4. Parallel Combinability

[0121] There is a dataset Divide it into A set of mutually disjoint subsets, namely If the random algorithm To satisfy localized differential privacy, then exist Satisfy differential privacy.

[0122] 2.3 Euclidean distance

[0123] Euclidean distance is The natural length of a vector between two points in a dimensional space, or the true distance between two points

[27] , is sensitive to outliers.

[0124] The formula for calculating Euclidean distance in two-dimensional space is as follows:

[0125] (1),

[0126] (2),

[0127] in, Let be the Euclidean distance between the two points. It is a point Distance to the center point.

[0128] The formula for calculating the Euclidean distance in 3D space is as follows:

[0129] (3),

[0130] 2.4 Mann-Whitney U Test

[0131] The Mann-Whitney U test (also known as the U test) is a nonparametric statistical method used to determine whether two independent samples come from the same distribution, without assuming that the data follow a normal distribution or other specific distribution. The test steps are as follows: First, set the null hypothesis... and alternative hypothesis The sample data are merged and arranged in order, and a rank is assigned to each data point; assuming the sum of the ranks of the two groups of samples are respectively and The sample sizes are respectively and Then we have:

[0132] (4),

[0133] Similarly, we can obtain ,Will and Minimum and critical values Comparison. If Then reject the null hypothesis. ,accept Conversely, accept the original hypothesis. .

[0134] 3. TDA Solution, 3.1 System Architecture

[0135] To reasonably assess user reliability in cold start scenarios, protect user data privacy, and defend against malicious collusion attacks, TDA models the trust assessment process as an exchange of trust between the perception platform and users. Users can exchange trust from the perception platform based on three dimensions: historical contributions, privacy utility, and data similarity. Subsequently, a U-test is used to verify the consistency of user data and reduce the impact of false data on task results. The entities included in TDA are the task requester, the perception platform, and the perception user.

[0136] The perception platform is responsible for receiving tasks initiated by task requesters, broadcasting task information and related requirements to users, assessing user reliability and data quality, and delivering task results to task requesters upon completion. Perception users are responsible for data collection and uploading; they use smart devices to complete data collection tasks, process the data according to privacy protection policies, and then upload it to the platform. Task requesters submit perception tasks to the platform, receive task results from the platform, and apply them to relevant fields after training.

[0137] 3.2 Trust Assessment

[0138] From the perspective of social exchange theory, the trust relationship in a crowd-sensing system can be viewed as a social exchange process between the platform and users. Users provide high-quality data (contribution) in exchange for the platform's trust (reward), while the platform selects reliable users through trust assessment (selection). Based on social exchange theory, this invention models the trust assessment process of TDA as a three-dimensional exchange relationship: historical contribution dimension, privacy utility dimension, and data similarity dimension, and comprehensively calculates the user's trust value. The design details of the trust assessment mechanism of this invention will be described in detail below.

[0139] 3.2.1 Historical Contributions

[0140] In this invention, historical contribution value According to user The historical data is used to calculate a user's historical contribution value, quantifying their long-term exchange behavior. This historical contribution value is then compared to the user's integrity score. and user activity value Closely related, integrity score and activity score reflect a user's reliability and initiative in completing historical tasks.

[0141] User's integrity score It is the user's historical contribution value A core component of this system reflects the reliability of a user's "data quality commitment." Generally, more trustworthy users contribute more to the task. Therefore, a user's trustworthiness score is derived from the proportion of reliable data submitted in historical tasks, calculated using the following formula:

[0142] (5),

[0143] in, The total number of times the user completes a task. The number of times reliable data is submitted to users. These are growth control parameters.

[0144] Integrity Value The calculation is divided into integrity growth and dishonesty punishment. It is the positive cumulative effect of users submitting reliable data that controls the growth of integrity scores. This represents the cumulative effect of reliable data submitted by users. This represents the total number of tasks a user participates in, serving as a normalization factor. If a malicious user wants to increase their integrity score, they need to complete a large number of tasks (…). ),but This will inhibit the growth of its integrity score. In terms of integrity growth, the initial growth is rapid, while the later growth is slow, which aligns with the security requirement that trust needs to be accumulated over a long period. As a form of penalty for breach of trust, no penalty is imposed when a user submits reliable data for every task; however, even if a malicious user submits a large amount of data rapidly in a short period, penalties will be imposed if the failure rate is low. If the exchange is deemed too high and violates trust-based exchange norms, the platform will significantly lower the user's trust score.

[0145] Response time is one of the important factors affecting the data quality of perception tasks, so this invention will use the activity value. As one of the reference indicators for historical contributions, user activity value represents whether a user can respond promptly to task requests broadcast by the platform and actively complete sensory tasks. The calculation of activity value adopts a two-level evaluation system: single-task activity value... Compared with historical activity value Single task activity value assessment of users In specific tasks The timeliness of the data is considered, while the historical activity value is a comprehensive result of the performance of multiple recent tasks, serving as the final result of the user's activity value.

[0146] User Upload task The time of the data is recorded by the user. For the task The platform will record the completion time of each user's task. The task can be derived. The median completion time of all users is denoted as the task. base time By comparing the difference between two time periods, we can determine the user's... Single activity value The calculation formula is:

[0147] (6),

[0148] in, This refers to the time dispersion (the median absolute deviation of task completion time). It is a positive part function.

[0149] In single activity value In the calculation process, the platform only penalizes timeouts. .when hour, ,when hour, It decays exponentially, and each time it exceeds , decline .

[0150] Historical activity value Based on recent user activity The performance of this task is calculated using the following formula:

[0151] (7),

[0152] in, User's historical activity value The coefficient of variation (standard deviation / mean) is a standardized measure of the magnitude of fluctuation. The maximum allowable fluctuation threshold for the system, together with the weighting, constitutes a quantitative evaluation mechanism for user behavior stability. This mechanism detects abnormal fluctuations in user performance and prevents malicious users from employing alternating "normal-malicious" behaviors to evade system detection. and As time decays, the weight of more recent tasks increases. The sliding window mechanism simulates the recent preference effect to prevent users' historical trustworthy behavior from masking recent malicious behavior.

[0153] Through triple verification of user reliability, timeliness, and stability, fine-grained evaluation of user historical behavior is achieved, ultimately resulting in a historical contribution value. The calculation formula is:

[0154] (8),

[0155] in, These are the weighting coefficients. It is used to balance user reliability and response speed.

[0156] 3.2.2 Privacy Utility

[0157] In a crowdsourcing sensing system, new users lack historical records, making it impossible to assess their reliability based on past contributions. Considering the personalized privacy needs of users in crowdsourcing sensing, some have low sensitivity to privacy exposure. Therefore, this invention introduces privacy utility to provide new users with an initial trust exchange channel. Privacy utility assessment allows new users with low privacy exposure sensitivity to actively participate in trust exchange by increasing their privacy budget. Users will rationally weigh the privacy investment against the trust return, deciding whether to expose more data in exchange for a privacy utility value. .

[0158] This invention uses localized differential privacy technology to protect user data privacy and security, assuming that the user... The submitted data is The platform wants to query a certain function. (e.g., user's location coordinates), towards Add Laplace noise to satisfy , It is noise sampled according to a Laplace distribution.

[0159] Malicious users may frequently increase their privacy budget in a short period to accumulate privacy utility value. Therefore, this invention divides privacy utility into short-term utility and long-term utility. Short-term utility value. By user Privacy budget per task The calculation shows that:

[0160] (9),

[0161] Long-term utility value Exponential decay is used to suppress short-term speculative behavior and ensure the fairness of long-term trust exchange. The calculation formula is as follows:

[0162] (10),

[0163] in, It is an adjustment parameter that controls the rate of increase in short-term utility value; It is the attenuation factor. .

[0164] If a user has no history, the user's privacy utility value is the short-term utility value. If a user has a history, the user's privacy utility value is the long-term utility value. .user The formula for calculating privacy utility value is:

[0165] (11),

[0166] 3.2.3 Data Similarity

[0167] In the absence of user history, privacy utility can be used to assess user reliability; however, some users are highly sensitive to privacy exposure and are unwilling to disclose their privacy in exchange for a privacy utility value. Therefore, this invention uses data similarity... Incorporating data similarity into a trust assessment mechanism refers to user... Submitted data Weighted consensus value The degree of similarity between them.

[0168] Group consensus is generated from user data with historical contributions, and long-term contributors should receive higher weight. Specifically, users are assigned different weights based on their historical trust values, so that users with high trust values ​​have a greater influence on group consensus, that is, regulating individual behavior through normative pressure. Assume a user set... There is If a user has historical contributions, then the weighted group consensus value is... The calculation formula is:

[0169] (12),

[0170] When new users with high privacy sensitivity join, they can choose to comply with group exchange guidelines to gain trust. (Data similarity) The calculation formula is:

[0171] (13),

[0172] in, yes and The Euclidean distance between them It is the weighted group standard deviation, calculated using the following formula:

[0173] (14),

[0174] 3.2.4 Trust Integration

[0175] Through trust exchange across three dimensions—historical contribution, privacy utility, and data similarity—a weighted fusion is used to derive a linear trust value for the user. The calculation formula is:

[0176] (15),

[0177] in, , , Let be the weighting coefficient, satisfying .

[0178] Directly use linearly weighted trust values Using a linearly weighted trust score as a user's trust value presents several problems: it struggles to capture the marginal effect of trust and cannot suppress the influence of extreme values; furthermore, malicious users may fabricate high scores in a single dimension, exhibiting short-term speculative behavior. Therefore, a sigmoid function is used to non-linearly map the linearly weighted trust value before outputting the final user trust score. The calculation formula is:

[0179] (16),

[0180] The trust evaluation algorithm is shown in Algorithm 1:

[0181] Algorithm 1. Trust Evaluation Algorithm

[0182] Input: User set Historical records Privacy Budget Collection Task completion time set User Data Set

[0183] Output: Set of user trust values

[0184] 1. Initialize parameters

[0185] 2. for each do

[0186] 3.if

[0187] 4. Calculate according to formula (5) Integrity value

[0188] 5. Calculate according to formula (6) Single activity value

[0189] 6. According to Calculated with formula (7) Historical activity value

[0190] 7. According to , Calculated with formula (8) Historical contribution value

[0191] 8.if

[0192] 9. Calculate according to formula (9) short-term utility value

[0193] 10. According to Calculated with formula (10) long-term utility value

[0194] 11. Determine if the user has task records.

[0195] 12. Output according to formula (11) Privacy utility value

[0196] 13.end if

[0197] 14. Calculate the group-weighted consensus value according to formula (12).

[0198] 15. According to Calculated with formula (13) Data similarity

[0199] 16. From formula (15), we can derive... Linear weighted trust value

[0200] 17. According to formula (16), nonlinear mapping The conclusion is Final trust value

[0201] 18. return

[0202] 19. end for

[0203] 3.3 Data Validation

[0204] In crowdsourced sensing systems, malicious users may collude to upload false data, thereby affecting data quality and the accuracy of task results. To address this issue, this invention proposes a data verification mechanism that integrates trust assessment and hypothesis testing. This mechanism makes trust decisions based on data consistency checks, filtering out false data uploaded by malicious users and reducing the impact of their false data on task results. The design details of the data verification mechanism are described below.

[0205] 3.3.1 Data Preprocessing

[0206] To reduce the computational involvement of low-trust user data, high-trust user data is processed first, and user-uploaded data sets are processed accordingly. Based on final trust value Sort in descending order to obtain a descending data set. ,

[0207] Indicates the first The trust values ​​of a user, i=1,2,3..., satisfy... Then the data set The data in the table is grouped, and the number of groups is [number]. The size of each group is approximately , obtained the Group data .

[0208] 3.3.2 Trust Pruning

[0209] After grouping the data, removing obviously unreliable data can reduce subsequent computation. Therefore, low-trust-value data within a group is discarded based on the trust value. The trust threshold within each group is calculated as follows:

[0210] (17),

[0211] Based on a trust threshold, data with low trust values ​​is discarded, and data is retained. ,like Then retain the data with the highest trust value within the group. .

[0212] 3.3.3 Accelerating Construction via Hash Methods

[0213] To reduce the number of comparisons, a Locality Sensitive Hash (LSH) table is created, which performs U-tests only on data within the same hash bucket, skipping tests on obviously inconsistent data.

[0214] For data Extract its feature vector :

[0215] in express Quantiles Let Variance be the variance.

[0216] Then, construct the LSH function:

[0217] (18),

[0218] in, The width of the bucket.

[0219] 3.3.4 Majority Election

[0220] First, initialize the candidate data within the group. The data of the user with the highest trust value within the group is used as candidate data. Iteratively eliminate candidate data. traversal ,like and If they are in the same hash bucket, a consistency check is performed.

[0221] Will and The feature vectors are merged, and ranks are assigned to the merged data. For data with the same value, the average rank is taken. The ranks of the candidate data and the data to be tested are calculated separately. and At this time, we have:

[0222] (19),

[0223] in, and Similarly, for data dimensions, we can obtain Ultimately, the U statistic is... .

[0224] Let the null hypothesis be... For: The two sets of data come from the same distribution (consistent); Alternative hypothesis The two sets of data have different distributions (inconsistent). Given a significance level... Critical value If consistent ( ),reserve If they are inconsistent ( ),disuse and The candidate data is then reset to the next highest-trust-value data. The final output is the candidate data within the group, i.e., the data that was not eliminated. .

[0225] 3.3.5 Cross-group aggregation and validation

[0226] Through majority voting within the group, this invention has obtained the candidate sets output by each group. The consistency test is used again to determine the consistency between candidate data, using the consistency matrix. The consistency among all candidate data is quantified by the following formula:

[0227] (20),

[0228] Then, this invention requires finding the largest consistent subset (Clique problem). The candidate data is considered as graph nodes, if... If the nodes are consistent, then there are edges between them. To find the largest complete subgraph (where every node is connected to the next), we need to find the largest complete subgraph. The clique problem is NP-hard. This invention uses a greedy approximation algorithm to solve it, first initializing an empty set. Select with current Completely consistent candidate data Prioritize adding the one with the highest trust value, iterating repeatedly until it cannot be expanded further. Ultimately, if Output data set ,like If the result is inconsistent, the above process will be repeated after increasing the number of groups.

[0229] Algorithm 2. Data Validation Algorithm

[0230] Input: User dataset User trust value set

[0231] Output: A high-quality dataset after verification.

[0232] 1. Sort the dataset in descending order of trust value. To obtain an ordered set

[0233] 2. Divided into Groups, each group being of size

[0234] 3. for each do

[0235] 4. Calculate the intra-group trust threshold according to formula (17).

[0236] 5. Eliminate all members of the group. Data

[0237] 6. If Then retain the data with the highest trust value.

[0238] 7. Return the candidate dataset

[0239] 8. end for

[0240] 9. for each do

[0241] 10. Regarding the data Extracting feature vectors

[0242] 11. Distribute the data into different buckets based on their hash values.

[0243] 12. Initialize candidate data

[0244] 13. for each do

[0245] 14.if and In the same hash bucket

[0246] 15. Perform a consistency check; if consistent, retain the data with the higher trust value.

[0247] 16.end if

[0248] 17. return

[0249] 18. end for

[0250] 19. end for

[0251] 20. Construct a consistency matrix according to formula (20).

[0252] 20. Greedy algorithm for finding the largest consistent subset

[0253] 21.if

[0254] 22. Increase the number of groups and skip to step 1 to repeat the above process.

[0255] 23.else

[0256] 24. return

[0257] 25.end if

[0258] 4. Performance Analysis

[0259] 4.1 Perceived Data Security

[0260] Theorem 1: Data processing in TDA satisfies differential privacy.

[0261] Proof: For any user In this context, the noise added to the data is noise sampled according to a Laplace distribution. Based on the fundamental properties of the Laplace distribution, for the function... In other words, if noise This satisfies differential privacy. In TDA, data processing for each user is performed independently and is independent of the user's privacy settings. Using Laplace noise, therefore for a single user The data processing satisfies differential privacy. For the entire system, since the data processing of each user is independent, based on the parallel combination property of differential privacy, the system processes the user set... The level of data protection still meets the requirements of differential privacy.

[0262] Theorem 1 is proved.

[0263] 4.2 Trust Assessment Mechanism

[0264] Theorem 2: TDA satisfies social exchange theory compatibility.

[0265] Proof: Long-term reciprocity. TDA achieves long-term reciprocal exchange between the platform and users through the accumulation of historical contributions and a penalty mechanism. In TDA, the user's trust value calculation includes the historical contribution dimension (Formula (15)), and its core is the integrity value and the activity value. As can be seen from Formula (5), users need to continuously submit reliable data to accumulate integrity value, and the platform increases their trust value as a reward. Dishonest behavior will trigger penalties and destroy the reciprocal relationship. As can be seen from Formula (6), users can obtain high activity value by responding to tasks in a timely manner. The sliding window mechanism ensures that recent behavior has a higher weight, which strengthens the dynamic balance of long-term reciprocity.

[0266] With a rational approach, TDA offers new users two trust-building pathways, supporting them in optimizing their strategies based on rational choices. If a user has low sensitivity to privacy exposure, they can choose to sacrifice privacy to gain trust (benefits); conversely, they can gain trust through data similarity, adjusting the similarity between their data and group consensus to minimize privacy exposure (costs) while acquiring trust.

[0267] In TDA, users need to adhere to the group consensus to meet the reliability specifications defined by the group. As shown in formula (12), users with high trust values ​​have a greater impact on the group consensus, creating pressure on the group's norms. New users need to upload real data that is close to the group consensus; otherwise, the platform will lower their trust value.

[0268] In conclusion, Theorem 2 is proved.

[0269] 4.3 Data Validation Mechanism

[0270] Theorem 3: When the number of malicious users does not exceed 50%, TDA can effectively identify and filter data uploaded by malicious users in collusion.

[0271] Proof: As shown in formula (16), malicious users, due to their lack of historical contributions and low data similarity, have significantly lower trust values ​​than normal users and are eliminated during the trust pruning stage. When the proportion of malicious users... The percentage of false data in the retained data TDA performs consistency checks on data within the same hash bucket; if the data is inconsistent, it is rejected. The error probability is That is, the probability of correctly identifying fake data is In most election phases, the proportion of false data is high. Reliable data prevails in most elections, which can further eliminate false data.

[0272] A consistency matrix for all candidate data is constructed using formula (20), and a greedy algorithm is used to find the maximum consistent subset. If Then the largest consistent subset must contain Reliable data.

[0273] For any set of data, if there is false data, the consistency check will be performed. probability of rejection For all data, when the percentage of malicious users... Majority elections and cross-group aggregation ensure at least The data is reliable and consistent with the data. Based on accuracy probability, TDA can effectively identify and filter data uploaded by malicious users in collusion.

[0274] In conclusion, Theorem 3 is proved.

[0275] 4.4 Algorithm Time Complexity

[0276] Theorem 4: The time complexity of TDA is O(n log n). .

[0277] Proof: Assume that For a given number of users, the time complexity of Algorithm 1 primarily depends on the dimensionality of the data and the number of tasks selected in the sliding window; the time complexity of the remaining trust calculations is also [missing information]. Since the data dimension and the number of tasks are constant, the time complexity of Algorithm 1 is O(n log n). .

[0278] In Algorithm 2, the time complexity of trust sorting is O(n log n). Assuming it is divided into The time complexity of grouping and pruning for set data is O(n log n). If the amount of data in the hash bucket is constant The time complexity of constructing the LSH and performing the consistency check is then... When greedily solving for the maximum clique, the worst-case time complexity is O(n log n). Number of groups Much smaller than the number of users ,at this time Therefore, the time complexity of Algorithm 2 is O(n log n). Known Since it is a constant, the main time complexity of Algorithm 2 is O(n). .

[0279] In summary, the time complexity of TDA is O(n log n). Therefore, Theorem 4 is proved.

[0280] 5. Experimental Verification

[0281] 5.1 Experimental Setup

[0282] To evaluate the performance of TDA, this invention conducted extensive simulations on a PC running Windows 10. The PC's processor was an AMD Ryzen 7 4800H with Radeon Graphics 2.90 GHz, and it had 8GB of memory. The programming platform used was JetBrains PyCharm 2019.1. This invention uses the literature (DTI) and literature (RPS) as benchmarks to compare the algorithms, comparing their performance in cold start scenarios, defense against collusion attacks, and data quality. It analyzes the contributions of historical contributions, privacy utility, and data similarity to system performance in TDA.

[0283] The simulation experiment was conducted according to the experimental parameters listed in Table 1. The simulation data was uniformly and randomly generated within the range shown in the table. Perceptual task data for users was randomly generated. Data from normal users followed a Gaussian distribution near the true values, while data from colluding users uniformly deviated from the true values, forming a cluster of false data. To ensure the experiment's validity, a trust growth control coefficient was used. Historical contribution weighting coefficient The privacy utility adjustment parameter and privacy utility decay factor are both set to a neutral value of 0.5. Since this scheme focuses on validating cold start scenarios, the trust value weighting coefficient... , , Set to 0.4, 0.3, and 0.3 respectively. Number of sliding window checks. Set it to 10.

[0284] Table 1 Parameter Settings

[0285] parameter value describe n [0,50] <![CDATA[The number of times user u i submits reliable data]]> m [0,50] <![CDATA[User u i Total number of times the task is completed]]> <![CDATA[t i j ]]> [2,6] <![CDATA[User u i for task t j completion time]]> <![CDATA[ε i j ]]> [0,10] <![CDATA[User u i for task t j 's privacy budget]]>

[0286] The experiment was conducted in three parts. In Part 1, this invention demonstrated the impact of three methods on the accuracy of trust assessment and data verification under different conditions. First, with a fixed user base of 1000, the proportion of colluding users was kept constant, while the proportion of new users was gradually increased to 80%. Then, with the proportion of new users constant, the proportion of colluding users was gradually increased to 50%. The effectiveness of the TDA scheme was verified using accuracy, fake data filtering rate (CDR), and mean absolute error (MAE). In Part 2, with a fixed ratio of new users to colluding users, the runtime of the three methods was compared under different user scales. In the final part of the experiment, the impact of each component on TDA performance was compared.

[0287] 5.2 Experimental Results and Analysis

[0288] like Figure 3 As shown, at the initial stage, TDA's accuracy was slightly lower than DTI and PRS. As the proportion of new users gradually increased, TDA's curve became almost horizontal, while DTI and RPS showed a downward trend. Because new users lack historical records, TDA provides two trust exchange channels for them: users with low privacy sensitivity can increase their privacy budget in exchange for trust points, while users with high privacy sensitivity can exchange trust points through data similarity. Even as the proportion of new users gradually increases, TDA can still assess the reliability of new users through these two methods, avoiding a sharp increase in trust assessment errors due to the lack of historical records. RPS relies more on historical records, and therefore performs worse than DTI.

[0289] like Figure 4 , Figure 5As shown, DTI utilizes peer-to-peer trust recommendations to achieve a decentralized trust assessment mechanism. However, if users launch a conspiracy attack, DTI will be significantly affected. Therefore, as the proportion of conspiracy users gradually increases, DTI's CDR drops sharply. TDA uses consistency checks to filter out false data uploaded by malicious users in conspiracy, so its CDR is higher than DTI's. In addition, TDA does not rely on prior data, while RPS uses historical data as verification evidence to identify conspiracy attackers, so its CDR is slightly lower than TDA's. Therefore, DTI's MAE is higher than both TDA and RPS, while TDA is least affected by conspiracy attacks.

[0290] Figure 6 The graph shows the runtime of the three methods at different user scales. As can be seen from the graph, the runtime of all three methods is relatively short, and overall, all three methods are effective. TDA runs faster than DTI and RPS because DTI requires pairwise comparison of user trust relationships (trust cascading), while RPS requires iterative truth discovery across the entire dataset, making both have higher time complexity.

[0291] from Figure 7 It can be seen that the complete TDA maintains a high accuracy rate, but if privacy utility or data similarity is turned off, the overall performance of TDA trust assessment decreases. When users cannot exchange trust through privacy utility...

[0292] Both users with low and high privacy sensitivity can choose to exchange data similarity for trust value. However, when data similarity is turned off, users with high privacy sensitivity may not be willing to choose to expose their privacy in exchange for trust value. Therefore, TDA trust assessment is more affected when only data similarity is turned off.

[0293] from Figure 8 It can be seen that trust pruning and grouping, along with LSH, significantly improve the data verification efficiency of TDA. Whether using LSH directly to process all data or only performing trust pruning and grouping, the runtime is slower than the complete TDA scheme. This is because as the proportion of colluding users increases, without trust pruning, TDA needs to process more user data, while directly performing consistency checks would multiply the number of checks. Combining the two improves the computational efficiency of TDA.

[0294] This invention belongs to the field of data verification technology, and particularly relates to a trust-based data verification method and system for crowdsourced sensing. Against the backdrop of the widespread adoption of mobile smart devices and the rapid development of IoT technology, this technical solution addresses the core pain points in crowdsourced sensing scenarios, such as inconsistent data quality, difficulty in protecting user privacy, and frequent malicious collusion attacks, demonstrating broad commercial application potential. In the field of environmental monitoring, it can empower distributed environmental monitoring platforms to screen reliable user data through three-dimensional trust assessment, ensuring the authenticity of monitoring results such as air quality and water pollution. In the field of public health, it can support public health data collection systems, effectively identifying false epidemic data and providing accurate epidemic monitoring data for health departments. In the field of intelligent transportation, it can be applied to intelligent transportation data hubs to quickly screen reliable road condition information and improve traffic scheduling efficiency. Simultaneously, trust data records based on localized differential privacy can generate value-added data security services, including user credit assessment, privacy protection level certification, and as a qualification credential for participating in advanced data collaboration tasks. This is expected to drive the rapid development of multiple sub-markets such as environmental monitoring, public health, intelligent transportation, and data security services, forming new economic growth points and industrial ecosystems.

[0295] Evidence related to the technical effects obtained by the embodiments of the present invention.

[0296] Significantly improves the data quality and reliability of the crowd perception system (accuracy maintained above 0.85), and ensures fairness in participation for both new and existing users through a three-dimensional trust assessment mechanism. The trust assessment accuracy for both new and existing users is superior to traditional methods, as demonstrated by theoretical analysis and experimental data (e.g. Figure 3 , Figure 4 , Figure 5 This jointly validated the rationality of the trust assessment and the effectiveness of the data verification. Experiments showed that the system maintained a stable accuracy rate even in a cold start scenario where the proportion of new users was as high as 80%. Figure 3 The False Data Filtering Rate (CDR) significantly outperforms DTI and RPS methods; in attack scenarios where the proportion of colluding users reaches 50%, the CDR maintains its leading advantage. Figure 4 The mean absolute error (MAE) is least affected by attacks. Figure 5 This meets the security requirements of the crowd-sensing system. Theoretical analysis further shows that localized differential privacy technology ensures the privacy protection of user data (Theorem 1), social exchange theory compatibility guarantees the fairness of trust assessment (Theorem 2), and the anti-collusion attack mechanism can still effectively identify false data when the proportion of malicious users is ≤50% (Theorem 3). Regarding system efficiency, trust pruning and LSH acceleration technology (…) Figure 8 Significantly improves data verification efficiency, maintaining reasonable runtime even in large-scale scenarios with up to 1000 users. Figure 6 Component analysis experiment () Figure 7This verifies the importance of privacy utility and data similarity module. The complete TDA scheme performs optimally in all performance indicators, fully demonstrating the advanced nature and practicality of the technical solution of this invention.

[0297] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0298] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A trust-based data verification method in crowd perception, characterized in that, Includes the following steps: Step 1: Trust assessment includes four parts: historical contribution, privacy utility, data similarity and trust integration. These parts together constitute a multi-dimensional trust measurement system. The two parties involved in the trust assessment exchange are the platform and the user, which is used to solve the trust initialization problem for cold start users. Step 2: Use a localized differential privacy mechanism to perturb the user-uploaded data, quantify data similarity, and achieve a dynamic balance between privacy protection and trust assessment; Step 3: Use a user data consistency determination method based on Mann-Whitney U test, and combine it with locality-sensitive hashing algorithm to improve the test efficiency, thereby filtering out false or collusive data uploaded by malicious users; The calculation of the historical contribution includes: Quantitative results of users' long-term exchange behavior are obtained by recording users' historical tasks; The reliability of fulfilling data quality commitments is characterized by user integrity score and activity score. The integrity score is calculated based on the ratio of the number of times a user submits reliable data to the total number of tasks. The activity value is calculated by exponential decay based on the difference between the task completion time and the median task time, in order to measure the timeliness of the user. The integrity value and activity value are weighted and summed according to a preset weighting coefficient to obtain the user's historical contribution value, which is used to reflect the stability of the user's long-term behavior. The privacy utility is calculated using a localized differential privacy mechanism. Random noise following a Laplace distribution is introduced into the user's original data to meet the privacy budget constraint. The short-term privacy utility value is calculated based on the privacy budget of a single task, and the long-term privacy utility value is obtained by smoothing historical tasks using an exponential decay function. Finally, the short-term or long-term privacy utility value is selected as the privacy utility result of the current task based on whether the user has a historical record, in order to suppress the short-term speculative behavior of malicious users. The data similarity is calculated based on group consensus, which is generated from user data with historical contributions. Different weights are assigned to each user based on their historical trust value, with higher-trust users having higher weights. Calculate the Euclidean distance between each user's data and the group consensus data, and use the weighted group standard deviation as the normalization coefficient. Measure the consistency between user data and group data by the weighted distance ratio, thereby achieving the measurement of data authenticity. The trust fusion combines weighted fusion with nonlinear mapping. First, historical contribution, privacy utility, and data similarity are linearly weighted according to preset weights to obtain an initial trust value. Then, an S-shaped nonlinear mapping function is used to normalize the initial trust value to weaken the influence of extreme values ​​and capture the marginal change characteristics of trust. Finally, the user's trust value is output to guide subsequent task allocation and data filtering.

2. The system implemented by the method as described in claim 1, characterized in that, include: The trust exchange module is used to establish a multi-dimensional social exchange model between the platform and users and initialize trust parameters. The privacy and filtering module is used to perform localized differential privacy perturbations, calculate data similarity, and identify anomalous data by combining statistical testing methods. The analysis and verification module is used to calculate and verify the trust assessment results and detect malicious behavior.

3. The system as described in claim 2, characterized in that, The trust exchange module includes: The historical evaluation unit is used to calculate integrity and activity scores based on the user's historical task records. A privacy assessment unit is used to calculate short-term and long-term privacy utility based on privacy budget and noise parameters; The similarity calculation unit is used to calculate data similarity indicators based on the group consensus model.

4. The system as described in claim 2, characterized in that, The privacy and filtering module includes: Differential privacy processing unit is used to inject noise that follows a Laplace distribution into user data to achieve privacy protection; The consistency check unit is used to perform rank-based nonparametric statistical tests on user-uploaded data. The hash optimization unit is used to achieve fast matching and fake data filtering using the locality-sensitive hashing algorithm.

5. The system as described in claim 2, characterized in that, The analysis and verification module includes: The trust fusion unit is used to calculate the final trust value based on linear weighting and nonlinear mapping. The visualization evaluation unit is used to display the dynamic changes in trust among different users and to assist in system task scheduling; The malicious detection unit is used to identify and block malicious user groups who collude to upload false data.

Citation Information

Patent Citations

  • Privacy protection mobile crowd sensing method based on reinforcement learning

    CN113553614A

  • Federal learning-based identity authentication method and device, medium and program product

    CN120474679A