Multi-subject data sharing and incentive distribution method and system

By calculating data quality scores and extracting feature vectors, combining game theory models and behavior prediction technology, optimizing dynamic incentive factor weights, and generating tamper-proof benefit distribution records, the problem of balancing data privacy protection and model performance improvement in multi-subject data sharing is solved, achieving fair incentive distribution and full utilization of data value.

CN120612128APending Publication Date: 2025-09-09BEIJING JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510693898.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In multi-agent data sharing and collaboration, how to balance data privacy protection and model performance improvement, how to accurately evaluate the contributions of all parties and design dynamic incentive mechanisms to avoid short-sighted behavior, how to fairly link model contribution with profit distribution, and how to find a balance between privacy protection, model performance, incentive mechanisms and reputation assessment.

Method used

Data quality scores are calculated through statistical analysis and feature vector extraction technology, and short-term and long-term benefits are evaluated in combination with game theory models to generate dynamic incentive mechanisms. A preliminary allocation plan for incentive factors is generated using the subject behavior prediction model, and the weights of dynamic incentive factors are optimized through cost constraint analysis. Finally, the unalterable profit distribution records are recorded through smart contracts.

Benefits of technology

It effectively balances the interests of multiple subjects, increases the enthusiasm for data sharing, promotes the full exploration and utilization of data value, and provides new ideas for the design of incentive mechanisms in multi-subject collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612128A_ABST
    Figure CN120612128A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-subject data sharing and incentive distribution method and system, and the method comprises the steps: obtaining multi-subject shared data which comprises data quality, data volume and cooperation cost information; processing the shared data by adopting a data cleaning and statistical analysis technology to generate a quality score of each subject; according to the quality score and the cooperation cost information, calculating a cost distribution proportion, generating a boundary condition evaluation index, and determining a shared boundary condition of each subject; according to the initial parameter configuration and the historical cooperation data, a behavior prediction model is adopted to generate behavior tendency parameters, and an excitation factor distribution scheme is determined; according to the dynamic excitation factor distribution scheme and the quality score, a contribution evaluation technology is adopted to calculate the model contribution degree of each subject, and a profit distribution strategy is generated; and recording the income distribution strategy through the intelligent contract, and generating an income distribution record which cannot be tampered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to a multi-agent data sharing and incentive distribution method and system. Background Art

[0002] In multi-agent data sharing and collaboration, there is a technical contradiction in how to balance data privacy protection and model performance improvement. The use of a federated learning framework can protect the original data, but it is difficult to accurately evaluate the contributions of all parties. If a feature vector sharing mechanism is introduced, the model accuracy can be improved, but the risk of data leakage is increased. On the other hand, the design of a dynamic incentive mechanism also faces challenges. Short-term benefits are easy to quantify, but may lead to short-sighted behavior of participants. Considering long-term benefits helps maintain the stability of cooperation, but it is difficult to accurately predict. In addition, how to link model contribution with fair distribution of benefits is also a key issue. If the contribution evaluation is too frequent, it will increase the computing overhead, and if the interval is too long, the adjustment opportunity may be missed. On this basis, the introduction of a dynamic reputation mechanism can restrain malicious behavior, but overly strict punishment measures may undermine participation enthusiasm. How to find a balance between privacy protection, model performance, incentive mechanism and reputation evaluation is a technical problem that needs to be solved urgently. Summary of the Invention

[0003] In order to solve the problems existing in the above-mentioned prior art, the purpose of this application is to provide a multi-agent data sharing and incentive distribution method and system.

[0004] The multi-agent data sharing and incentive allocation method described in this application includes:

[0005] S101. Obtain data quality, data volume, and collaboration cost allocation information from multi-agent shared data. Calculate the data quality score for each agent using statistical analysis and feature vector extraction techniques. Combined with the collaboration cost allocation, generate boundary condition evaluation indicators to determine the shared boundary conditions for each agent.

[0006] S102. Based on the boundary condition evaluation indicators, a game theory model is used to construct a payoff matrix to calculate the short-term payoff calculation results and long-term payoff evaluation values ​​of each entity in cooperative and non-cooperative scenarios. If the data quality score is higher than the preset threshold, the payoff matrix weight is adjusted to generate the initial parameter configuration of the dynamic incentive mechanism;

[0007] S103. Obtain short-term benefit calculation results and long-term benefit evaluation values ​​from the initial parameter configuration of the dynamic incentive mechanism, use the subject behavior prediction model to analyze the behavioral tendencies of each subject through historical collaboration data, generate behavior prediction parameters, and determine a preliminary allocation plan for incentive factors;

[0008] S104. Based on the preliminary incentive factor allocation plan and the collaboration cost allocation ratio, cost constraint analysis technology is used to optimize the dynamic incentive factor weights by weighing the data quality score and cost constraints, and generate a dynamic incentive factor allocation plan;

[0009] S105. Obtain incentive factor weights from the dynamic incentive factor allocation plan and data quality scores, use contribution assessment technology, and calculate the model contribution record value of each entity through data substitutability and feature vector extraction analysis to generate the initial weights of the income distribution strategy;

[0010] S106. Obtain the model contribution record value from the initial weight of the profit distribution strategy, and record the profit constraints and distribution ratio of each entity through the smart contract. If the profit constraints meet the preset rules, an unalterable profit distribution record is generated.

[0011] Preferably, in step S101, the step of processing the shared data using data cleaning and statistical analysis techniques to generate a quality score for each subject includes:

[0012] Processing the shared data using preset data cleaning rules to obtain a standardized data set;

[0013] Determine whether the data quality of the standardized data set is lower than a preset threshold; if so, calculate the distribution characteristics of the data quality using a statistical analysis method to obtain a data quality vector;

[0014] Extract key features from the data quality vector using feature vector extraction technology, and calculate the quality score of each subject in combination with the data volume information;

[0015] If the quality score does not meet the preset conditions, the weight of the feature vector is adjusted, and the quality score is recalculated to obtain an optimized quality score.

[0016] Preferably, in step S101, calculating the cost allocation ratio according to the quality score and the collaboration cost information includes:

[0017] Calculate the cost allocation ratio of each subject using a linear weighted method based on the quality score and the collaboration cost information to obtain a cost allocation result;

[0018] Determine whether the variance of the cost allocation ratio is greater than a preset threshold; if so, adjust the weight of the feature vector, recalculate the cost allocation ratio, and obtain an optimized cost allocation result;

[0019] Based on the optimized cost allocation result and the quality score, a cluster analysis method is used to generate boundary condition evaluation indicators to determine shared boundary conditions.

[0020] Preferably, in step S102, the use of a game theory model to construct a benefit matrix and calculate the benefit evaluation value of each entity includes:

[0021] Obtaining an initial payoff matrix of the game theory model, wherein the initial payoff matrix includes short-term payoffs and long-term payoffs under cooperative and non-cooperative scenarios;

[0022] Determining whether the quality score is higher than a preset threshold, and if so, adjusting the weights of the initial profit matrix according to the quality score to obtain an updated profit matrix;

[0023] Recalculating the short-term and long-term benefits of each entity using the updated benefit matrix to obtain an adjusted benefit assessment value;

[0024] According to the adjusted benefit evaluation value, a genetic algorithm is used to optimize parameter configuration to obtain an optimized incentive parameter set.

[0025] Preferably, in step S103, generating behavior tendency parameters using a behavior prediction model based on the initial parameter configuration and historical collaboration data includes:

[0026] Obtain historical collaboration data from the storage system and use data cleaning technology to obtain a standardized collaboration data set;

[0027] extracting subject behavior features from the standardized collaborative data set using a feature extraction algorithm to obtain a subject behavior feature vector;

[0028] Processing the subject behavior feature vector through a pre-trained behavior prediction model to obtain behavior tendency parameters;

[0029] According to the behavioral tendency parameters and the initial parameter configuration, the short-term benefit value and the long-term benefit evaluation value are calculated using the benefit calculation formula. The benefit calculation formula is S=w1P+w2Q, where S represents the benefit value, P represents the behavioral tendency parameter, Q represents the initial parameter configuration, and w1 and w2 represent weight coefficients.

[0030] Preferably, in step S104, the step of optimizing the incentive factor weights using cost constraint analysis technology to generate a dynamic incentive factor allocation plan includes:

[0031] Obtaining the incentive factor weights and the cost allocation ratio, and determining an initial allocation plan;

[0032] Determining whether the data quality score of the initial allocation plan is lower than a preset threshold; if so, adjusting the incentive factor weights using a trade-off analysis technique to obtain an optimized weight set;

[0033] According to the optimized weight set, the cost allocation ratio is updated using a dynamic weight adjustment technology to generate a dynamic incentive factor allocation plan;

[0034] It is determined whether the cost constraint condition of the dynamic incentive factor allocation scheme meets the preset cost allocation ratio. If not, the incentive factor weight is re-optimized to obtain an adjusted allocation scheme.

[0035] Preferably, in step S105, the step of calculating the model contribution of each subject using contribution evaluation technology includes:

[0036] Obtaining a data quality score matrix, wherein the data quality score matrix includes a data quality score vector for each subject;

[0037] Using a linear regression algorithm to process the data quality score matrix to obtain an incentive factor weight vector;

[0038] Calculating the subject participation vector using a weighted average method based on the incentive factor weight vector and the dynamic incentive factor matrix;

[0039] If there is a subject whose participation degree is lower than a preset threshold in the subject participation vector, a principal component analysis algorithm is used to process the subject participation vector and the data feature matrix to obtain a feature vector;

[0040] According to the characteristic vector and the data quality score matrix, data substitutability analysis is used to calculate the data substitutability score and determine a list of highly substitutable subjects.

[0041] Preferably, in step S106, recording the profit distribution strategy through a smart contract includes:

[0042] Extract the model contribution value from the initial weight of the income distribution strategy and calculate the contribution value set using a weighted algorithm;

[0043] Deploy revenue constraints through smart contracts, obtain the identities of each entity and its distribution ratio from the contribution value set, and generate a constraint set;

[0044] Determine whether the allocation ratio in the constraint condition set meets the preset rules. If so, generate a compliant constraint condition set through the verification mechanism of the smart contract;

[0045] Use hash algorithm to generate tamper-proof profit distribution records;

[0046] The storage address of the allocation record is obtained through on-chain storage technology to generate a permanently stored allocation log.

[0047] The multi-agent data sharing and incentive distribution system described in this application includes:

[0048] The data sharing module is used to input shared data from multiple entities, including information on each entity's data quality, data volume, and collaboration costs. It uses data cleaning rules to standardize the original shared data and generate a standardized data set. It uses statistical analysis techniques to evaluate the data quality distribution characteristics and generate a data quality vector. It combines feature vector extraction technology with data volume information to calculate the quality score of each entity. Based on the quality score and collaboration cost, it uses a linear weighted method to generate a cost allocation ratio, and uses cluster analysis to determine the sharing boundary conditions of each entity.

[0049] The utility calculation module is used to construct the payoff matrix of the game theory model based on shared boundary conditions, and calculate the short-term and long-term payoff assessment values ​​of each entity in cooperative and non-cooperative scenarios. The payoff matrix weights are dynamically adjusted according to the data quality score to generate the initial parameter configuration. The parameter configuration is optimized using a genetic algorithm to generate an optimized incentive parameter set to ensure the stability and long-term nature of the payoff distribution.

[0050] The incentive mechanism module is used to extract subject behavior characteristics from historical collaboration data and generate behavior feature vectors. It uses a predictive model to analyze behavioral tendency parameters and calculate short-term and long-term benefits based on initial parameter configurations. Based on cost constraint analysis technology, it dynamically adjusts the weights of incentive factors and generates an optimized incentive factor allocation plan.

[0051] The profit distribution module is used to calculate the model contribution record value of each entity through data substitution analysis and feature vector extraction technology; combine the incentive factor weights and data quality scores to generate the initial weights of the profit distribution strategy; and use normalization and weighted summation methods to determine the final profit distribution ratio;

[0052] The storage module is used to deploy revenue constraints through smart contracts and record the distribution ratio and contribution value of each entity; a hash algorithm is used to generate tamper-proof revenue distribution records; and on-chain storage technology is used to permanently store the distribution log on the blockchain node to ensure data transparency and traceability.

[0053] The multi-subject data sharing and incentive distribution method and system described in the present application has the advantages of calculating the data quality score and sharing boundary conditions of each subject by analyzing data quality, data volume and collaboration costs, using game theory models to evaluate short-term and long-term benefits, and generating a preliminary allocation plan for incentive factors in combination with the subject behavior prediction model. The dynamic incentive factor weights are further optimized through cost constraint analysis, and the model contribution is calculated using contribution evaluation technology to form a profit distribution strategy. Finally, the present invention records the profit constraints and distribution ratios through smart contracts to generate an unalterable profit distribution record. This method effectively balances the interests of multiple subjects, improves the enthusiasm for data sharing, promotes the full exploration and utilization of data value, and provides new ideas for the design of incentive mechanisms in multi-subject collaboration. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is the process of a multi-agent data sharing and incentive distribution method described in this application Figure 1 ;

[0055] Figure 2 This is the process of a multi-agent data sharing and incentive distribution method described in this application Figure 2 . DETAILED DESCRIPTION

[0056] like Figure 1-Figure 2 As shown, the multi-agent data sharing and incentive allocation method described in this application includes the following steps:

[0057] S101. Obtain data quality, data volume, and collaboration cost allocation information from multi-agent shared data. Calculate the data quality score for each agent using statistical analysis and feature vector extraction techniques. Combined with the collaboration cost allocation, generate boundary condition evaluation indicators to determine the shared boundary conditions for each agent.

[0058] S102. Based on the boundary condition evaluation indicators, a game theory model is used to construct a payoff matrix to calculate the short-term payoff calculation results and long-term payoff evaluation values ​​of each entity in cooperative and non-cooperative scenarios. If the data quality score is higher than the preset threshold, the payoff matrix weight is adjusted to generate the initial parameter configuration of the dynamic incentive mechanism;

[0059] S103. Obtain short-term benefit calculation results and long-term benefit evaluation values ​​from the initial parameter configuration of the dynamic incentive mechanism, use the subject behavior prediction model to analyze the behavioral tendencies of each subject through historical collaboration data, generate behavior prediction parameters, and determine a preliminary allocation plan for incentive factors;

[0060] S104. Based on the preliminary incentive factor allocation plan and the collaboration cost allocation ratio, cost constraint analysis technology is used to optimize the dynamic incentive factor weights by weighing the data quality score and cost constraints, and generate a dynamic incentive factor allocation plan;

[0061] S105. Obtain incentive factor weights from the dynamic incentive factor allocation plan and data quality scores, use contribution assessment technology, and calculate the model contribution record value of each entity through data substitutability and feature vector extraction analysis to generate the initial weights of the income distribution strategy;

[0062] S106. Obtain the model contribution record value from the initial weight of the profit distribution strategy, and record the profit constraints and distribution ratio of each entity through the smart contract. If the profit constraints meet the preset rules, an unalterable profit distribution record is generated.

[0063] like Figure 1-Figure 2 As shown, in step S101, data quality, data volume, and collaboration cost allocation information are obtained from the shared data of multiple subjects. Statistical analysis and feature vector extraction technology are used to calculate the data quality score of each subject. Combined with the collaboration cost allocation, the boundary condition evaluation index is generated to determine the shared boundary conditions of each subject.

[0064] Furthermore, in step S101 , multi-agent shared data is obtained, including the data source, data quality, data volume, and collaboration cost information of each agent, to obtain an original shared data set;

[0065] The original shared dataset is processed using the preset data cleaning rules to obtain a standardized dataset. If the data quality of the standardized dataset is lower than the preset threshold, the distribution characteristics of the data quality are calculated using statistical analysis methods to obtain a data quality vector.

[0066] Key features are extracted from the data quality vector using feature vector extraction technology. Combined with data volume information, the quality score of each subject is calculated. Based on the quality score and collaboration cost information, the cost allocation ratio of each subject is calculated using a linear weighted method to obtain the cost allocation result.

[0067] If the variance of the cost allocation ratio is greater than the preset threshold, the weight of the feature vector is adjusted, and the quality score and cost allocation ratio are recalculated to obtain the optimized cost allocation result;

[0068] Based on the optimized cost allocation results and quality scores, cluster analysis method is used to generate boundary condition evaluation indicators for each entity and determine the shared boundary conditions;

[0069] Obtain the dynamic incentive factor allocation plan and data quality score, use contribution assessment technology, through data substitutability and feature vector extraction analysis, calculate the model contribution record value of each subject, and obtain the initial weight of the income distribution strategy;

[0070] According to the initial weight of the profit distribution strategy and combined with the shared boundary conditions, the weighted average method is used to adjust the final profit distribution ratio of each entity to determine the profit distribution result.

[0071] Specifically, in step S101, when acquiring multi-subject shared data, the data source records, data quality scores (completeness 95%, accuracy 92%), data volume (subject A provides 10TB, subject B provides 8TB) and collaboration costs (subject A costs 5,000 yuan, subject B costs 4,500 yuan) of each subject are extracted from the database to form the original shared data set; the data tuples in the original data acquisition and cleaning are defined as Mi = (Si, Qi, Vi, Ci), Si: data source code (example: subject A = IoT_01), Qi = [q1, q2]: quality score vector (q1 = 95%, q2 = 92%), Vi: data volume (unit: TB, VA = 10 13 bytes), Ci: collaboration cost (unit: yuan, CA = 5×10 3 ), Qi: data quality score vector, Si: data source code.

[0072] The original data set was processed using data cleaning rules (missing values ​​were filled with the mean and outliers were removed with three times the standard deviation) to generate a standardized data set;

[0073] If the mean quality score of the standardized data set is lower than the threshold (85%), the distribution characteristics of the data quality, such as skewness (-0.3) and kurtosis (2.1), are calculated to construct a data quality vector. In the quality vector construction of the quality assessment system: m 3 =E[(q-μ) 3 ]=-0.3(skewness),m 4 =E[(q-μ) 4 ] = 2.1 (kurtosis);

[0074] The first three principal components (90% variance contribution) are extracted from the data quality vector through principal component analysis (PCA). Combined with the data volume weight (subject A accounts for 55% of the data volume), the quality score is calculated (subject A scores 88 points and subject B scores 76 points). The quality score calculation for cost allocation is: ω=[0.6,0.3,0.1](principal component weight), substitute the data:

[0075] Based on the quality score and collaboration cost, linear weighting (quality weight 0.6, cost weight 0.4) is used to calculate the cost allocation ratio (52% allocated to entity A and 48% allocated to entity B). In the cost allocation optimization: Parameter constraint: α+β=1(α=0.6,β=0.4), calculation example:

[0076] If the variance of the distribution ratio exceeds the threshold (0.05), the PCA weight is adjusted (the weight of principal component 1 increases from 0.5 to 0.6), and the score and distribution ratio are recalculated. The weight iterative optimization in the dynamic adjustment mechanism is: minw||Var(ρ(w))-0.05||2, gradient descent: (For example, when the weight of principal component 1 changes from 0.5 to 0.6, the variance drops to 0.048).

[0077] K-means clustering (k=3) is used to classify the optimized allocation results and quality scores to generate boundary condition evaluation indicators (the boundary condition of subject A is "data quality ≥ 85% and cost ratio ≤ 55%"); the boundary conditions of the dynamic adjustment mechanism are generated as follows: Eigenvector xi = (ρi, QSi / 100), example: the cluster center of subject A class μ1 = (0.55, 0.85).

[0078] Combining the dynamic incentive factor (timeliness factor weight 0.3) and contribution assessment technology (subject A’s contribution in alternative analysis is 70%), calculate the initial profit distribution weight (subject A’s weight 60%).

[0079] Finally, the profit distribution result (58% of the profit of subject A) is determined by weighted average (boundary condition weight 0.4, initial weight 0.6). In the final distribution calculation,

[0080] γ=0.6 (dynamic excitation factor),

[0081] Formula specification description: Data flow: Mi→Qvec→PCk→QSi→ρi→Wi.

[0082] In one embodiment, the principal component analysis (PCA) extracts the first three principal components from the data quality vector by principal component analysis (PCA), and the specific steps include:

[0083] Calculate the covariance matrix of the data quality vector;

[0084] Perform eigenvalue decomposition on the covariance matrix and sort it in descending order of eigenvalues;

[0085] The first three principal components (cumulative variance contribution ≥ 90%) were selected with weight distribution of [0.6, 0.3, 0.1].

[0086] In one embodiment, the clustering method is specifically implemented using the K-means clustering algorithm (number of clusters k=3, Euclidean distance metric, and a maximum number of iterations of 100). The optimized cost allocation results and quality scores are input as feature vectors to generate three types of boundary conditions:

[0087] Category 1: Data quality ≥ 85% and cost ratio ≤ 55%;

[0088] Category 2: Data quality ≥ 75% and cost ratio ≤ 65%;

[0089] Category 3: Other situations.

[0090] like Figure 1-Figure 2 As shown, in step S102, based on the boundary condition evaluation indicators, a game theory model is adopted to construct a profit matrix to calculate the short-term profit calculation results and long-term profit evaluation values ​​of each entity in cooperation and non-cooperation scenarios. If the data quality score is higher than the preset threshold, the profit matrix weight is adjusted to generate the initial parameter configuration of the dynamic incentive mechanism.

[0091] Furthermore, in step S102, based on the boundary condition evaluation indicators, a game theory model is used to construct a payoff matrix to calculate the short-term and long-term benefits of each entity in the cooperation and non-cooperation scenarios, and obtain an initial benefit evaluation value;

[0092] Obtain the short-term and long-term profit data of each entity in the initial profit matrix, combine it with historical collaboration data, use data quality assessment algorithms, calculate the data quality score, and determine the data quality score result;

[0093] Determine whether the data quality score is higher than the preset threshold. If it is higher than the preset threshold, adjust the weight of the initial profit matrix according to the data quality score to obtain an updated profit matrix;

[0094] Using the updated benefit matrix, recalculate the short-term and long-term benefits of each entity under cooperative and non-cooperative scenarios to obtain the adjusted benefit assessment value;

[0095] Based on the adjusted benefit evaluation value, generate the initial parameter configuration of the dynamic incentive mechanism and determine the incentive parameter set;

[0096] For the incentive parameter set, a genetic algorithm is used to optimize the parameter configuration, and the optimized incentive parameter set is obtained through iterative calculation and fitness evaluation;

[0097] Calculate the revenue growth rate of each entity in the cooperation scenario under the optimized incentive parameter set, determine whether the revenue growth rate is higher than the preset stability threshold, and determine the revenue stability result;

[0098] If the revenue growth rate exceeds the preset stability threshold, the final dynamic incentive mechanism configuration is generated through the optimized incentive parameter set to determine the long-term incentive plan;

[0099] The short-term benefit calculation results and long-term benefit evaluation values ​​are extracted from the final dynamic incentive mechanism configuration. The subject behavior prediction model is used to analyze the behavioral tendencies of each subject through historical collaboration data, generate behavior prediction parameters, and determine the preliminary allocation plan of incentive factors.

[0100] Specifically, in step S102, based on the boundary condition evaluation index, the Nash equilibrium theory in the game theory model is used to construct a 2×2 payoff matrix including agent A and agent B. The short-term payoff in the cooperative scenario is set to [5, 5], and in the non-cooperative scenario to [2, 3]. The initial payoff evaluation value is calculated to be [3.5, 4]. A 2×2 symmetric game matrix is ​​constructed in the Nash equilibrium game matrix constructed from the initial payoff matrix:

[0101] a11=5, [Short-term benefits of cooperation between subject A and subject B]

[0102] a12=2, [Short-term benefits of cooperation by subject A / non-cooperation by subject B]

[0103] a21=3, [Short-term benefits of non-cooperation of subject A / cooperation of subject B]

[0104] a22=0, [Short-term benefits of non-cooperation of subject A / non-cooperation of subject B]

[0105] After obtaining the initial profit matrix data, combined with the historical collaboration data set D = {d1, d2, …, dn}, the data quality assessment algorithm based on information entropy is used to calculate the data quality score Q = 0.85, with a preset threshold of 0.8;

[0106] If Q>0.8, update the benefit matrix according to the weight adjustment formula W_new=W_old×(1+Q) to obtain the adjusted cooperation scenario benefit [5.5,5.5];

[0107] Recalculate the adjusted benefit estimate using the updated matrix to [4.2, 4.7]. Generate the initial parameter configuration based on this value, and set the incentive parameter set P = {p1 = 0.6, p2 = 0.4};

[0108] Using genetic algorithm, we set the population size to 50, the crossover probability to 0.8, and the mutation probability to 0.1. After 100 iterations, we obtained the optimized set P_opt = {p1 = 0.72, p2 = 0.28}.

[0109] Calculate the revenue growth rate R = 12% in the cooperation scenario, and set the preset stability threshold to 10%. If R > 10%, output the long-term incentive plan L = {P_opt, period T = 30 days};

[0110] Extract short-term benefits [6.1, 6.1] and long-term benefits [7.3, 7.3] from plan L, use the LSTM behavior prediction model, input historical collaboration data D, output behavior tendency parameters B = {b1 = 0.8, b2 = 0.6}, and generate the incentive factor allocation plan F = {f1 = 0.7, f2 = 0.3}.

[0111] In one embodiment, a specific example is used below to illustrate, but the implementation of the present invention is not limited thereto: entity A provides 10 TB of data with a quality score of 95%, and entity B provides 8 TB of data with a quality score of 92%.

[0112] In one embodiment, the genetic algorithm is used to optimize the excitation parameters, and the parameters are set as follows:

[0113] Population size: 50, crossover probability: 0.8, mutation probability: 0.1, maximum number of iterations: 100, fitness function is the income growth rate, and the iteration is terminated when the growth rate exceeds a preset threshold (such as 10%).

[0114] like Figure 1-Figure 2 As shown, in step S103, the short-term benefit calculation results and long-term benefit evaluation values ​​are obtained from the initial parameter configuration of the dynamic incentive mechanism, and the subject behavior prediction model is adopted to analyze the behavioral tendencies of each subject through historical collaboration data, generate behavior prediction parameters, and determine the preliminary allocation plan of the incentive factors.

[0115] Furthermore, in step S103, historical collaboration data is obtained from the storage system, and data cleaning technology is used to remove noise data to obtain a standardized collaboration data set;

[0116] A feature extraction algorithm is used to extract subject behavior features from the standardized collaborative dataset to obtain the subject behavior feature vector;

[0117] The behavioral feature vectors of the subjects are processed by the pre-trained behavior prediction model to obtain the behavioral tendency parameters of each subject;

[0118] According to the initial configuration parameters and behavioral tendency parameters of the dynamic incentive mechanism, the profit calculation formula S = w1P + w2Q is used to calculate the short-term profit value and long-term profit evaluation value;

[0119] If the short-term profit value is lower than the preset threshold, the initial configuration parameters are adjusted and the profit calculation formula is re-used to obtain the optimized profit data;

[0120] Based on the optimized income data and behavioral tendency parameters, an allocation algorithm is used to generate a preliminary allocation plan for incentive factors;

[0121] Use simulation verification technology to test the preliminary allocation plan and obtain the final incentive factor allocation plan;

[0122] Obtain the initial payoff matrix of the game theory model, use the initial payoff matrix to calculate the payoffs of each subject in cooperative and non-cooperative scenarios, and obtain the initial payoff evaluation value;

[0123] If the data quality score is higher than the preset threshold, the weight of the initial benefit matrix is ​​adjusted according to the data quality score, and the benefits of each entity are recalculated to obtain the adjusted benefit evaluation value.

[0124] Specifically, in step S103, historical collaboration data for the past three months is obtained from the MySQL database in the storage system. The data is cleaned using the Python-based Pandas library, records with more than 30% missing values ​​are removed, and numeric fields are normalized using the Z-score to obtain a standard collaboration dataset containing 10,000 records.

[0125] The Random Forest feature importance algorithm was used to screen the top 10 behavioral features from the standardized dataset, including indicators such as task response time, collaboration frequency, and contribution volatility, forming a feature vector with a dimension of 10.

[0126] The feature vector is input into the pre-trained XGBoost behavior prediction model. The model outputs the cooperation tendency parameter of each subject in the interval [0,1]. The subject with a parameter exceeding 0.7 is identified as a subject with high cooperation tendency.

[0127] Based on the initial configuration parameters α = 0.5 and β = 0.3, combined with the behavioral tendency parameter P, the linear weighted formula S = 0.6*P + 0.4*Q is used to calculate the return value. When the short-term return value falls below the threshold of 0.65, the β parameter is increased by 0.1 and recalculated;

[0128] Based on the optimized revenue data, the Shapley value allocation algorithm is used to calculate the contribution weight of each entity and generate a preliminary plan with five incentive levels;

[0129] The solution was tested 1,000 times through Monte Carlo simulation, and the allocation items with incentive factor deviation exceeding 15% were adjusted to output the final solution.

[0130] The 2×2 game payoff matrix is ​​read from the NoSQL database. The initial cooperative payoff is [5,5], and the non-cooperative payoff is [3,1]. The average payoff evaluation value in the cooperative scenario is calculated to be 4.2.

[0131] When the data quality score reaches 0.85, the entropy weight method is used to increase the cooperation benefit weight from 0.5 to 0.7. After recalculation, the cooperation benefit evaluation value is increased to 4.6. The dynamic benefit formula in the benefit calculation model is: S i =α·P i +β·Q i , Si is the return value, α is the behavioral tendency weight, β is the quality factor weight, α=0.6: behavioral tendency weight, β=0.4: quality factor weight, when Si<0.65, the parameter adjustment is triggered: β′=β+Δβ.

[0132] like Figure 1-Figure 2 As shown, in step S104, according to the preliminary allocation plan of incentive factors and the allocation ratio of collaboration costs, cost constraint analysis technology is adopted to optimize the weights of dynamic incentive factors by weighing the data quality score and cost constraints, and generate a dynamic incentive factor allocation plan.

[0133] Furthermore, in step S104, a preliminary allocation plan of incentive factors and a collaborative cost allocation ratio are obtained, and the incentive factor weights and cost allocation ratios are extracted using cost constraint analysis technology to determine the initial allocation plan;

[0134] According to the initial allocation plan, the data quality score is calculated, and whether the data quality score is lower than the preset threshold is determined to determine whether the quality requirements are met;

[0135] If the data quality score is lower than the preset threshold, the weight of the incentive factor is adjusted through the trade-off analysis technology to obtain the optimized weight set;

[0136] Based on the optimized weight set, dynamic weight adjustment technology is used to update the collaboration cost allocation ratio and generate a dynamic incentive factor allocation plan;

[0137] Extract cost constraints from the dynamic incentive factor allocation plan, determine whether the cost constraints meet the preset cost allocation ratio, and determine whether the allocation plan meets the constraints;

[0138] If the cost constraint condition does not meet the preset cost allocation ratio, the dynamic incentive factor weights are re-optimized through the constraint condition processing technology to obtain an adjusted allocation plan;

[0139] Generate the final dynamic incentive factor allocation plan based on the adjusted allocation plan and determine the output result of the allocation plan;

[0140] Obtain the incentive factor weights from the final dynamic incentive factor allocation scheme and data quality scores. Use contribution evaluation techniques to calculate the model contribution degrees of each entity through data substitutability analysis, and obtain the initial contribution degree record values;

[0141] According to the initial contribution degree record values, optimize the model contribution degrees of each entity through eigenvector extraction analysis to generate the initial weights of the revenue distribution strategy.

[0142] Specifically, in step S104, obtain the preliminary incentive factor allocation scheme and the collaborative cost allocation ratio. Use the linear programming model in cost constraint analysis techniques, set the objective function to minimize the total collaborative cost, and the constraint conditions include the incentive factor weight range (such as 0.1 ≤ w_i ≤ 0.5), and solve to obtain the initial allocation scheme;

[0143] According to the initial allocation scheme, use a data quality evaluation algorithm (such as a scoring model based on information entropy) to calculate the data quality score S_q (such as S_q = 0.85), preset the threshold T_q = 0.8, and judge whether S_q < T_q holds. If it holds, adjust the incentive factor weights through the gradient descent method in the trade-off analysis technique, and iterate and optimize with a step size of Δw_i = 0.01 to obtain the optimized weight set W' = {w_1 = 0.25, w_2 = 0.35};

[0144] Based on W', use the proportional allocation algorithm in the dynamic weight adjustment technique to update the collaborative cost allocation ratio according to w_i' / (∑w_i') to generate a dynamic incentive factor allocation scheme;

[0145] Extract the cost constraint condition C_max = 1000 from the scheme, and judge whether the actual cost C = 950 satisfies C ≤ C_max. If not, re-optimize the weights through the Lagrange multiplier method in the constraint condition processing technique to obtain the adjusted allocation scheme;

[0146] According to the adjustment result, generate the final allocation scheme F = {w_1 = 0.23, w_2 = 0.32, C = 980};

[0147] Extract the weights from F and the data quality score S_q, and use the Shapley value algorithm in the contribution evaluation technique to calculate the contribution degrees V_i of each entity (such as V_1 = 0.4, V_2 = 0.6);

[0148] Based on V_i, optimize the contribution degrees through eigenvector extraction analysis (PCA dimensionality reduction) to generate the initial weights R = {r_1 = 0.38, r_2 = 0.62} of the revenue distribution strategy.

[0149] Such as Figure 1-Figure 2As shown, in step S105, the incentive factor weight is obtained from the dynamic incentive factor allocation plan and the data quality score, and the contribution evaluation technology is used to calculate the model contribution record value of each subject through data substitutability and feature vector extraction analysis to generate the initial weight of the profit distribution strategy.

[0150] Furthermore, in step S105, the initial incentive factor weight is obtained from the dynamic incentive factor allocation plan, and the comprehensive weight is calculated using the weighted average method in combination with the data quality score to obtain the incentive factor weight value of each subject;

[0151] According to the weight value of the incentive factor, the contribution evaluation technology is used to calculate the substitutability score of each subject's data to the model through the data substitutability analysis method to obtain the data substitutability record value;

[0152] By using feature vector extraction technology, key features are extracted from the data alternative record values. Combined with the data quality score, the model contribution of each subject is calculated using the linear regression method to obtain the model contribution record value;

[0153] Based on the recorded values ​​of model contribution, the contribution is standardized using the normalization method. Combined with the weight values ​​of the incentive factors, the initial weights of the income distribution strategy are generated using the weighted summation method to obtain the initial weight data set.

[0154] Acquire multi-subject shared data, process the shared data using preset data cleaning rules, combine data source and data volume information, generate standardized data sets, and obtain cleaned shared data records;

[0155] If the data quality of the standardized data set is lower than the preset threshold, the statistical analysis method is used to calculate the distribution characteristics of the data quality, and combined with the data volume information, a data quality vector is generated to obtain a data quality feature record;

[0156] By using feature vector extraction technology, key features are extracted from data quality feature records. Combined with collaboration cost information, a linear weighted method is used to calculate the quality score of each subject to obtain a quality score dataset.

[0157] Based on the quality score data set, the cluster analysis method is used to group the quality scores of each subject. Combined with the initial weight data set, the boundary condition evaluation index of each subject is generated to obtain the shared boundary condition record;

[0158] Through simulation verification technology, the shared boundary condition records and initial weight data sets are tested, and the weight parameters are adjusted using an iterative optimization algorithm to generate the final profit distribution strategy and obtain the optimized distribution plan.

[0159] Specifically, in step S105, the initial incentive factor weights (weight coefficients w1 = 0.6, w2 = 0.4) are extracted from the dynamic incentive factor allocation scheme. Combined with the data quality score (subject A scores 85 points), the weighted average method (formula: comprehensive weight = w1 × incentive factor weight + w2 × data quality score) is used to calculate the comprehensive weight to obtain the incentive factor weight value of each subject (e.g., the comprehensive weight of subject A is 0.72);

[0160] Based on the weight value of the incentive factor, the entropy weight method in the contribution evaluation technology is used to calculate the data substitutability score (the data substitutability score of subject A is 0.8). The key features (the principal component with a variance contribution rate greater than 90%) are extracted by combining the feature vector extraction technology (PCA principal component analysis). The model contribution is calculated by the linear regression method (least squares fitting) (for example, the contribution of subject A is 0.75). Specifically, the model contribution is calculated by the linear regression model (least squares method):

[0161] Contribution = 0.7 × quality score + 0.3 × alternative score;

[0162] Normalize the model contribution records (Min-Max normalization to the [0, 1] interval), combine them with the incentive factor weight (the weight of entity A is 0.72), and use weighted summation (formula: initial weight = normalized contribution × 0.7 + incentive factor weight × 0.3) to generate the initial weight of the profit distribution strategy (the initial weight of entity A is 0.74);

[0163] Acquire shared data from multiple subjects (the data volume of subject B is 1000), process the data using data cleaning rules (eliminating fields with more than 20% missing values), and generate a standardized data set (Z-score standardization) based on the data source (the data source of subject B is a public database);

[0164] If the data quality is lower than the threshold (F1-score < 0.7), the distribution characteristics of the data quality (skewness, kurtosis) are calculated to generate a data quality vector ([0.8, 0.2, 0.5]);

[0165] Key features (discriminant weight vector [0.6, 0.3, 0.1]) were extracted through feature extraction (LDA linear discriminant analysis). Combined with the collaboration cost (the cost of subject B was 500 yuan), a linear weighting method (formula: quality score = feature vector · weight vector + cost coefficient × 0.2) was used to calculate the quality score (subject B's score was 0.65).

[0166] Based on the quality score dataset (subject A = 0.8, subject B = 0.65), K-means clustering (k = 2) was used to group the subjects and generate boundary condition evaluation indicators based on the initial weights (the boundary condition of subject A is [0.7, 0.9]);

[0167] The boundary condition records were tested through simulation verification (Monte Carlo simulation 1000 times), and the weight parameters were adjusted using the gradient descent method (learning rate η = 0.01) to generate the final distribution plan (the income weight of subject A is 0.78.

[0168] like Figure 1-Figure 2 As shown, in step S106, the model contribution record value is obtained from the initial weight of the profit distribution strategy, and the profit constraints and distribution ratios of each entity are recorded through the smart contract. If the profit constraints meet the preset rules, an unalterable profit distribution record is generated.

[0169] Furthermore, in step S106, the contribution record value of each model is extracted from the initial weight of the profit distribution strategy, and the contribution value of each model is calculated using a preset weighting algorithm to obtain a contribution value set;

[0170] Deploy revenue constraints through smart contracts, obtain the identifiers of each entity and their corresponding distribution ratios from the contribution value set, and generate a constraint set;

[0171] If the allocation ratio in the constraint set meets the preset rules, the smart contract verification mechanism is used to determine whether the constraints are compliant and obtain a compliant constraint set;

[0172] Based on the set of compliant constraints, a hash algorithm is used to generate an unalterable profit distribution record to obtain a profit distribution record set;

[0173] Extract the profit distribution data of each entity from the profit distribution record set, and determine the profit share set of each entity through the distribution mechanism of the smart contract;

[0174] For the revenue share set, we use on-chain storage technology to obtain the storage address of the distribution record and generate a permanently stored distribution log;

[0175] Obtain access rights from the allocation log through the query interface of the blockchain network to determine the verifiable income distribution results of each entity;

[0176] The incentive factor weights are obtained from the dynamic incentive factor allocation plan and the data quality score. The contribution evaluation technology is used to calculate the model contribution record value of each subject through data substitution and feature vector extraction analysis to obtain the updated contribution record value set.

[0177] According to the updated contribution record value set, the initial weight of the profit distribution strategy is adjusted, and the adjusted profit constraints and distribution ratio are recorded through the smart contract to generate a new profit distribution record set.

[0178] Specifically, in step S106, the contribution record value of each model is extracted from the initial weight of the profit distribution strategy. For example, the initial weight of model A is 0.3, model B is 0.5, and model C is 0.2. The contribution value of each model is calculated using a weighted algorithm, and the contribution value set obtained is {model A: 0.3, model B: 0.5, model C: 0.2};

[0179] Deploy revenue constraints through smart contracts, obtain the identifiers of each subject and their corresponding distribution ratios from the contribution value set. For example, if the distribution ratio of subject X is 0.4 and that of subject Y is 0.6, generate the constraint set {subject X: 0.4, subject Y: 0.6};

[0180] If the allocation ratio in the constraint set meets the preset rules, for example, the allocation ratio of subject X does not exceed 0.5, the smart contract verification mechanism is used to determine whether the constraint is compliant, and the compliant constraint set {subject X: 0.4, subject Y: 0.6} is obtained;

[0181] Based on the set of compliant constraints, the SHA-256 hash algorithm is used to generate an unalterable profit distribution record, resulting in a profit distribution record set {record 1: hash value 1, record 2: hash value 2};

[0182] Extract the profit distribution data of each subject from the profit distribution record set. For example, if the profit of subject X is 100 and the profit of subject Y is 150, the profit share set of each subject is determined through the distribution mechanism of the smart contract {subject X: 100, subject Y: 150};

[0183] For the revenue share set, use on-chain storage technology to obtain the storage address of the allocation record, such as address 0x1234, and generate a permanently stored allocation log {log 1: address 0x1234};

[0184] Through the query interface of the blockchain network, access rights are obtained from the allocation log to determine the verifiable income distribution results of each entity {Entity X: 100, Entity Y: 150};

[0185] The incentive factor weight is obtained from the dynamic incentive factor allocation scheme and the data quality score. For example, if the data quality score is 0.8, the incentive factor weight is 0.2. The contribution evaluation technology is used to calculate the model contribution record value of each subject through data substitutability and feature vector extraction analysis, and the updated contribution record value set {Model A: 0.35, Model B: 0.45, Model C: 0.20} is obtained;

[0186] According to the updated contribution record value set, the initial weight of the profit distribution strategy is adjusted. For example, the weight of model A is adjusted to 0.35. The adjusted profit constraints and distribution ratio are recorded through the smart contract to generate a new profit distribution record set {Record 3: Hash value 3, Record 4: Hash value 4}.

[0187] In one embodiment, the profit distribution strategy is recorded through a smart contract, and the specific implementation includes:

[0188] Contract deployment: Based on the Ethereum platform, use the Solidity language to write contract logic and define distribution rules (contribution ≥ 50% can only obtain benefits);

[0189] Hash generation: Use the SHA-256 algorithm to perform hash operations on the allocation record to generate a unique identifier;

[0190] On-chain storage: Store the hash value and distribution details on IPFS, and record the storage address on the Ethereum blockchain.

[0191] The multi-agent data sharing and incentive distribution system described in this application includes:

[0192] The data sharing module is used to input shared data from multiple entities, including information on each entity's data quality, data volume, and collaboration costs. It uses data cleaning rules to standardize the original shared data and generate a standardized data set. It uses statistical analysis techniques to evaluate the data quality distribution characteristics and generate a data quality vector. It combines feature vector extraction technology with data volume information to calculate the quality score of each entity. Based on the quality score and collaboration cost, it uses a linear weighted method to generate a cost allocation ratio, and uses cluster analysis to determine the sharing boundary conditions of each entity.

[0193] The utility calculation module is used to share boundary conditions, construct the payoff matrix of the game theory model, and calculate the short-term and long-term payoff assessment values ​​of each entity in cooperative and non-cooperative scenarios. The payoff matrix weights are dynamically adjusted based on the data quality score to generate the initial parameter configuration. The parameter configuration is optimized using a genetic algorithm to generate an optimized incentive parameter set to ensure the stability and long-term nature of the payoff distribution.

[0194] The incentive mechanism module is used to extract subject behavior characteristics from historical collaboration data and generate behavior feature vectors. It uses a predictive model to analyze behavioral tendency parameters and calculate short-term and long-term benefits based on initial parameter configurations. Based on cost constraint analysis technology, it dynamically adjusts the weights of incentive factors and generates an optimized incentive factor allocation plan.

[0195] The profit distribution module is used to calculate the model contribution record value of each entity through data substitution analysis and feature vector extraction technology; combine the incentive factor weights and data quality scores to generate the initial weights of the profit distribution strategy; and use normalization and weighted summation methods to determine the final profit distribution ratio;

[0196] The storage module is used to deploy revenue constraints through smart contracts and record the distribution ratio and contribution value of each entity; a hash algorithm is used to generate tamper-proof revenue distribution records; and on-chain storage technology is used to permanently store the distribution log on the blockchain node to ensure data transparency and traceability.

[0197] Those skilled in the art can make various other corresponding changes and deformations based on the technical solutions and concepts described above, and all of these changes and deformations should fall within the scope of protection of the claims of this application.

Claims

1. A multi-agent data sharing and incentive distribution method, characterized in that: include: Obtain the quality, scale, and collaboration cost parameters of multi-agent shared data, standardize it through data cleaning rules, construct a data quality vector containing skewness and kurtosis characteristics, use principal component analysis to extract feature dimensions, and calculate weighted quality scores; A linear weighted model was constructed based on quality scores and collaboration cost parameters. Sharing boundary conditions were determined through cluster analysis. A game theory payoff matrix was established and a genetic algorithm was used to optimize incentive parameters. After the game theory payoff matrix outputs the initial parameter configuration, the behavior prediction model is trained based on historical collaboration data. The weights of incentive factors are dynamically adjusted based on the quality parameters and prediction results. The linear programming model is used to generate an allocation plan that meets the constraints. After allocating the plan, the Shapley value algorithm is used to calculate the subject contribution, and the alternative evaluation model is established in combination with the principal component eigenvector to generate normalized contribution records; Profit constraints are deployed through smart contracts, and hash algorithms are used to generate tamper-proof distribution records.

2. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The data cleaning rules include: standardizing the original shared data set by using missing value filling mean and outlier three times standard deviation elimination technology.

3. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The specific steps of the principal component analysis (PCA) include: calculating the covariance matrix of the data quality vector, performing eigenvalue decomposition, selecting the first three principal components with cumulative variance contribution rates ≥ 90%, and assigning weights of 0.6, 0.3, and 0.

1.

4. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The cluster analysis adopts K-means algorithm, with the number of clusters k=3, Euclidean distance as the metric, and a maximum number of iterations of 100 to generate three types of boundary conditions.

5. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The parameters of the genetic algorithm are set as follows: population size 50, crossover probability 0.8, mutation probability 0.1, maximum number of iterations 100, fitness function is the income growth rate, and the iteration is terminated when the growth rate exceeds 10%.

6. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The behavior prediction model is an XGBoost model, which inputs the top 10 behavior feature vectors extracted from historical collaboration data and outputs a cooperation tendency parameter in the range of 0 and 1. If the parameter exceeds 0.7, it is determined to be a subject with high cooperation tendency.

7. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The smart contract is based on the Ethereum platform and is written in Solidity. The distribution rule is that only those with a contribution of ≥50% can obtain benefits, and the hash value is generated by the SHA-256 algorithm and stored in IPFS.

8. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The Shapley value algorithm is used to calculate the subject contribution, and the principal component eigenvector is combined to establish an alternative evaluation model, and the contribution is normalized to the interval of 0, 1 through Min-Max standardization.

9. A multi-agent data sharing and incentive distribution method according to claim 1, characterized in that: The dynamic incentive factor weight is adjusted using a gradient descent method, with a learning rate of η=0.01, an objective function of minimizing the total collaboration cost, and a constraint condition of an incentive factor weight range of 0.1≤w_i≤0.

5.

10. A multi-agent data sharing and incentive distribution system, characterized in that: include: The data sharing module is used for multi-agent data sharing. After standardization and quality assessment, cost allocation is performed and cluster analysis is performed to determine the boundaries. The utility calculation module is used to construct a game theory payoff matrix based on shared boundaries, dynamically adjust weights, and optimize using genetic algorithms to ensure stable payoff distribution. The incentive mechanism module is used to extract behavioral characteristics from historical data, analyze trends using predictive models, and dynamically adjust the weight of incentive factors under cost constraints; The income distribution module is used for alternative analysis and feature extraction to calculate contribution, combine incentive factors and quality scores to generate initial weights, and normalize the weighted sum to determine the distribution ratio; The storage module is used for smart contract constraints on revenue distribution and on-chain storage.

Citation Information

Cited By

  • Digital archive sharing cooperation method and system based on intelligent contract

    CN121327864A