Bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing
By using a bilateral personalized information fusion mechanism, the problem of processing multi-source heterogeneous data in mobile crowd sensing was solved, enabling adaptive data quality assessment and privacy protection, and improving the system's applicability and user satisfaction.
Patent Information
- Application Number
- CN202310317192.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-03-28
AI Technical Summary
Existing mobile crowd sensing technologies fail to effectively consider the personalized needs of data collectors and users when processing multi-source heterogeneous data, resulting in insufficient system applicability and inability to be applied to complex sensing scenarios.
It adopts a bilateral personalized information fusion mechanism, including personalized privacy sampling, adaptive data quality assessment and personalized data quality sampling, combined with differential privacy technology, to adaptively process multi-source heterogeneous data and meet the personalized needs of data collectors and users.
It improves system scalability and user satisfaction, enhances the ability to process multi-source heterogeneous data, protects data privacy, and improves system efficiency.
Smart Images

Figure CN116527318B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile crowd sensing technology, specifically to a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing. Background Technology
[0002] In recent years, with the increasing demands of mobile devices on sensors, more and more types of sensors have been integrated into mobile intelligent systems, laying the foundation for the vigorous development of the Internet of Things (IoT) technology. Mobile crowd sensing collects sensing data through mobile intelligent devices carried by widely distributed individuals to complete complex sensing tasks, injecting new vitality into traditional IoT technology. Today, mobile crowd sensing is widely used in many areas such as environmental monitoring, intelligent transportation, and urban management, and has become an important component supporting the development of smart cities.
[0003] In Mobile Crowdsensing (MCS), roles within the system can be categorized into three types based on their responsibilities: data collectors, sensing platform servers, and data users. The sensing platform server assigns tasks to qualified data collectors based on their specific content. After completing their tasks, data collectors upload the collected data to the sensing platform server. Data users submit their usage requests to the sensing platform server, which then performs information fusion on the collected data based on these requests and returns the fusion results to the data users.
[0004] However, in mobile crowd sensing, the types of tasks are diverse and complex, and the data types that need to be processed in different scenarios are often heterogeneous, with varying data quality. Improving the information fusion mechanism of crowd sensing, enabling it to adaptively process isolated and heterogeneous sensing data, can greatly enhance the applicability of mobile crowd sensing systems and expand their application scenarios.
[0005] Existing technical solutions often fall into two categories: one focuses on the bilateral privacy and security of data collectors and data users, but neglects the personalized needs of data collectors and data users; the other only focuses on the personalized needs of data collectors, but neglects the personalized needs of data users. Both of these solutions are only for single data types and do not take into account that information aggregation in a crowd-sensing system may involve different types of sensing data and generate different types of aggregation results. Therefore, they are not applicable to sensing scenarios with multi-source heterogeneous sensing data. To address this, we propose a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd-sensing. Summary of the Invention
[0006] The purpose of this invention is to provide a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd intelligence sensing, thereby solving the problems in the background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing, including a system model of the bilateral personalized information fusion mechanism in mobile crowd sensing. The system model of the bilateral personalized information fusion mechanism in mobile crowd sensing includes a data collector, a crowd sensing platform server, and data users. The crowd sensing platform server includes the following steps:
[0008] S1: Personalized privacy sampling;
[0009] S2: Adaptive Data Quality Assessment;
[0010] S3: Personalized data quality sampling;
[0011] S4: Price Calculation.
[0012] Preferably, the crowd sensing platform server issues tasks to qualified data collectors, instructing them to collect specified information using mobile sensing devices.
[0013] Preferably, the data collector collects data and uploads it to the crowd-sensing platform server, and the data collector sets a privacy protection level for each piece of data.
[0014] Preferably, the crowdsourced sensing platform server will provide compensation to the data collector based on the quality of the data submitted. After receiving the data, the sensing platform server will adaptively select a data quality assessment mechanism to evaluate the data quality according to the data type and task requirements.
[0015] This invention provides a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing. This bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing has the following beneficial effects:
[0016] (1) The bilateral personalized information fusion mechanism for multi-source heterogeneous data in this crowd sensing can adaptively process multi-source heterogeneous data. It can adaptively select the data quality assessment mechanism according to the type of sensing data, such as numerical, text and image data, and adaptively select the differential privacy mechanism according to the type of query results. Therefore, this solution can be applied to different crowd sensing task environments, can cope with the increasingly complex application scenarios of crowd sensing systems, and enhance the scalability of crowd sensing systems.
[0017] (2) The bilateral personalized information fusion mechanism for multi-source heterogeneous data in the crowd sensing system takes into account the personalized needs of bilateral users in the crowd sensing system. On the basis of considering the personalized privacy needs of data collectors, it meets the personalized needs of data users for data quality, improves the satisfaction of bilateral users in the crowd sensing system, and enhances the scalability of the system.
[0018] (3) The bilateral personalized information fusion mechanism for multi-source heterogeneous data in this group intelligence perception uses differential privacy technology to protect privacy information. Compared with the scheme using traditional encryption technology, it requires less computation and has higher system efficiency. Attached Figure Description
[0019] Figure 1 This is a system model diagram of the present invention;
[0020] Figure 2 This is a system flowchart of the present invention. Detailed Implementation
[0021] like Figure 1-2 As shown, this invention provides a technical solution: a bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing, including a system model of the bilateral personalized information fusion mechanism in mobile crowd sensing. The system model includes data collectors, a crowd sensing platform server, and data users. The crowd sensing platform server issues tasks to qualified data collectors, instructing them to collect specified information using mobile sensing devices. Data collectors collect data and upload it to the crowd sensing platform server. Each data collector sets a privacy protection level for its data. The crowd sensing platform server pays data collectors based on the quality of the submitted data. After receiving the data, the sensing platform server adaptively selects a data quality assessment mechanism to evaluate the data quality based on the data type and task. The crowd sensing platform server includes the following steps:
[0022] S1: Personalized privacy sampling;
[0023] The goal of personalized privacy sampling is to provide each data collector's perceived data with privacy protection no less than their privacy needs. In practice, it's unrealistic to expect users to understand the detailed principles of privacy protection mechanisms and set a precise level of privacy protection for their perceived data accordingly. Differential privacy introduces the concept of a privacy budget to quantify the degree of privacy protection; a smaller privacy budget results in a higher level of privacy protection, but also introduces more noise during perturbations.
[0024] In our approach, we will use a privacy budget as a metric to measure the level of privacy protection. To improve system usability, we will provide data collectors with a user-friendly option so that users can clearly express their privacy needs: we allow perceptual users to use `r` to represent the privacy protection level of their perceptual data, where `r` ∈ {1, 2, 3, 4, 5, 6, 7, 8, 9}. We set nine levels to represent the degree of privacy protection, with higher numbers indicating greater privacy protection. For the perceptual task T, we assume its corresponding privacy budget range is [ε]. min ,ε max For privacy protection level r i Perceived data d i The privacy budget we allocated to it is:
[0025]
[0026] Here, `rand` is a random function for generating random numbers, and `rand(x,y)` outputs a random number between `x` and `y`. We believe that a fixed privacy budget can pose certain risks to user privacy protection. Attackers can use the privacy budget to calculate the parameters related to the noise added during perturbation, thereby obtaining the noise distribution and inferring the true value of the aggregation result. Therefore, we introduce randomness into the privacy budget process to better protect the privacy of perceived users. Furthermore, in this paper, we assume that different perceived tasks correspond to different privacy budget spaces for perceived data. Therefore, we do not define ε... min and ε max The value of is limited to a fixed value because we believe that the sensitivity of the corresponding perceptual data varies for different perceptual tasks, and the overall level of privacy protection required for the perceptual data set also varies. We can calculate a data set S = {ε1,...,ε} based on the privacy protection level chosen by the data collector and formula (1). M This dataset contains the privacy budget corresponding to each piece of sensing data, and there is a one-to-one mapping relationship between it and the sensing data in the sensing dataset D.
[0027] Based on the calculated overall privacy budget corresponding to all perceived data, we set an appropriate sampling threshold t. p We will perform personalized privacy sampling on the collections of sensory data uploaded by data collectors. We will independently sample data d with the following probabilities: i Perform sampling:
[0028]
[0029] Where ε i For sensing data d i The corresponding privacy budget;
[0030] After the personalized privacy sampling step, we will obtain a subset D1 = {d} of the entire perceptual data set D. a ,..,d b (a≥1, b≤M).
[0031] S2: Adaptive Data Quality Assessment;
[0032] In real-world scenarios of crowdsourced sensing, the quality of sensing data for a given sensing task is often inconsistent. In our solution, we will sample a subset of data D1 obtained through personalized privacy sampling, based on the user's needs and the data quality of each sensing data point, and use only a portion of this data for aggregation and computation.
[0033] Before sampling, we must first assess the data quality of each piece of perceived data. For tasks where data quality is difficult to assess due to the challenge of direct computation, user reputation or reliability can be used as a standard. However, this method typically relies on empirical task completion records, resulting in lower accuracy for users with limited task records and often lacking specificity; it is usually reserved as a last resort for data quality assessment. In our study, we adaptively employ corresponding data quality assessment mechanisms based on task type, using reputation or reliability mechanisms as the final means of data quality assessment.
[0034] S3: Personalized data quality sampling;
[0035] Similar to the humanized privacy protection levels we offer to data collectors, we provide data users with three data quality levels—{high, medium, low}—to choose from, allowing them to clearly indicate their data quality needs. Meanwhile, we assume that for data d... i ∈D, and its data quality is q i ∈Q. That is, for the data set D={d1,..,d...} M There exists a data quality set Q = {q1, ..., q}. M There is a one-to-one mapping between the data in set D and set Q. Let q be the maximum value in set Q. max The minimum value is denoted as q. min .
[0036] We will calculate the personalized data quality sampling threshold based on the data quality level ql selected by the data user using the following formula:
[0037]
[0038] Where Q1 is the first tertiary of set Q, and Q2 is the second tertiary of set Q. Then we will independently apply the following probabilities to the data d. j Perform sampling:
[0039]
[0040] After the personalized data quality sampling step, we will obtain the final perceptual data subset D2 = {d} used for aggregation calculation. c ,..,d d (c≥a,d≤b).
[0041] S4: Price Calculation;
[0042] In our model, the perception platform server will calculate the price payable by the data user based on the data quality of each piece of perception data in the actual used subset D2 of perception data. We use v u This represents the value of a unit of data quality. The larger the u value, the higher the value of the perceived data. For a piece of perceived data d... j Its value
[0043] For the perceptual data subset D2, its value is:
[0044]
[0045] After the perception platform server calculates the price to be paid, it sends the price value to the data user. After the data user completes the payment, the aggregated result is sent back to the data user.
Claims
1. A bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing, including a system model for a bilateral personalized information fusion mechanism in mobile crowd sensing, characterized by: The system model of the bilateral personalized information fusion mechanism in the mobile crowd sensing system includes a data collector, a crowd sensing platform server, and data users. The crowd sensing platform server issues tasks to qualified data collectors, instructing them to use mobile sensing devices to collect specified information; The data collector collects data and uploads it to the crowd intelligence sensing platform server, and the data collector sets a privacy protection level for each piece of data. The crowd intelligence sensing platform server includes the following steps: S1: Personalized privacy sampling; S2: Adaptive Data Quality Assessment; S3: Personalized data quality sampling; S4: Price Calculation; S1 includes: Using privacy budget as a metric to measure the degree of privacy protection, a perceptual user uses r to represent the privacy protection level of their perceptual data, where r∈{1,2,3,4,5,6,7,8,9}. For a perceptual task T, the corresponding privacy budget range is [ , For privacy protection level Perceived data Its allocated privacy budget is: Here, `rand` is a random number generator, and `rand(x,y)` outputs a random number between `x` and `y`. An appropriate sampling threshold is set based on the calculated overall privacy budget corresponding to all perceived data. Personalized privacy sampling will be performed on the collection of perceived data uploaded by data collectors, and the data will be independently sampled with the following probabilities. Perform sampling: , in For sensing data The corresponding privacy budget; S2 includes: The data quality of each piece of perceived data is evaluated, and the corresponding data quality evaluation mechanism is adaptively adopted according to the type of task. The reputation mechanism or reliability mechanism is used as the means of data quality evaluation. S3 includes: Provide data users with three data quality levels – {high, medium, low} – to choose from, assuming that for data… Its data quality is For the data set D={ ,.., }, there exists a data quality set Q={ ,.., There is a one-to-one mapping relationship between the data in set D and set Q. Let the maximum value in set Q be denoted as... The minimum value is denoted as ; Based on the data quality level selected by the data user The threshold for personalized data quality sampling is calculated using the following formula: , in It is the first tertile of set Q. It is the second tertile of set Q, which will independently affect the data with the following probability. Perform sampling: , The final subset of perceived data used for aggregation calculations is obtained; S4 includes: The perception platform server calculates the price payable by the data user based on the data quality of each piece of perception data in the actual perception data subset. After calculating the price, the server sends the price value to the data user and waits for the data user to complete the payment before sending the aggregated result back to the data user.
2. The bilateral personalized information fusion mechanism for multi-source heterogeneous data in crowd sensing according to claim 1, characterized in that: The collective intelligence sensing platform server will provide compensation to the data collector based on the quality of the data submitted by the data collector. After receiving the data, the sensing platform server will adaptively select a data quality assessment mechanism to assess the data quality according to the data type and task situation.
Citation Information
Patent Citations
Game method for position privacy protection and platform task allocation in mobile crowd sensing
CN111770454A
Differential privacy method for collecting real-time position information of user in mobile crowd sensing
CN113207120A