Intelligent weighted consensus method and system for multi-agent decision making

CN122718136APending Publication Date: 2026-09-08钰兔科技集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610858223.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

当智能体针对不同子目标形成多个优势子群时,单一加权平均会湮灭子群内的局部共识,导致最终决策丧失对关键少数意见的代表性,降低系统在复杂任务中的鲁棒性和适应能力

Benefits of technology

[0052]本发明实施例基于智能加权共识方法,各智能体的长期可靠度与决策互补性被充分融合为动态信任权重,有效抑制不可靠或恶意智能体对共识结果的负面影响,显著提升决策准确性与鲁棒性。历史贡献记录与偏差修正记录计算可信指数,结合协作关系图谱捕获互补性与冲突性,使信任分配兼顾个体历史表现与团队协作价值,在分布式决策环境中实现去中心化的信任评估与权重动态调整,避免单一信任来源导致的偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122718136A_ABST
    Figure CN122718136A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of multi-agent decision-making, and particularly relates to an intelligent weighted consensus method and system for multi-agent decision-making, which collects each agent decision-making scheme and historical contribution deviation record to calculate a trust index, extracts a decision-making target to analyze complementarity and conflict, constructs a cooperation relationship graph, generates a cooperation contribution factor and a trust index nonlinear fusion to obtain a dynamic trust weight, identifies a high-density decision-making area through kernel density estimation to form a decision-making peak, and generates a peak value representing a decision-making by weighting with the dynamic trust weight in each peak, constructs a fusion optimization model with global consistency and decision-making peak representation retention maximization as double optimization objectives to search for a Pareto optimal solution space to generate a final consensus decision, and reversely updates the trust index according to the execution effect. The present application improves the accuracy and robustness of multi-agent decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent decision-making technology, and in particular to an intelligent weighted consensus method and system for multi-agent decision-making. Background Technology

[0002] Multi-agent decision-making systems are widely used in distributed artificial intelligence, the Internet of Things, and cooperative control. One of their core tasks is to coordinate local decision proposals submitted by multiple agents to form a global consensus. Existing conventional approaches typically employ voting-based consensus mechanisms or weighted averaging methods, where weight allocation often relies on static reputation scores or simple historical success rate statistics of agents. For example, fusion strategies based on Byzantine fault tolerance algorithms or norm averaging achieve decision integration by setting fixed weight thresholds or by linearly weighting historical behaviors through a centralized server.

[0003] These conventional approaches suffer from two main drawbacks. First, weight update mechanisms mostly employ lagging periodic adjustments, lacking sensitivity to real-time behavioral biases and collaborative dynamics of agents. This allows malicious or faulty agents to accumulate false reputations through short-term performance, thus exerting a sustained negative influence in subsequent consensus processes. Second, the consensus process typically treats all decision options as independent, identically distributed samples, directly generating the final decision using a globally weighted average. This ignores the complementarity and conflict between different decision objectives and fails to identify the potential multimodal distribution characteristics in the decision space. When agents form multiple dominant subgroups for different sub-objectives, a single weighted average can annihilate local consensus within these subgroups, causing the final decision to lose representativeness of the key minority opinions and reducing the system's robustness and adaptability in complex tasks. Summary of the Invention

[0004] This invention provides an intelligent weighted consensus method and system for multi-agent decision-making, which can solve the problems in the prior art.

[0005] A first aspect of the present invention provides an intelligent weighted consensus method for multi-agent decision-making, comprising:

[0006] Collect decision proposals submitted by each agent and simultaneously obtain historical contribution records and deviation correction records of each agent;

[0007] Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a trust index; the decision target vector of the decision scheme is extracted, the decision complementarity and conflict between agents are analyzed and a cooperation relationship graph is constructed, the node centrality and edge weight distribution are calculated to generate a cooperation contribution factor, and the dynamic trust weight is obtained by nonlinear fusion with the trust index.

[0008] Based on the decision target vector of the decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weight to generate peak representative decisions.

[0009] A fusion optimization model is constructed based on the decisions of each peak representative. The dual optimization objectives are to maximize global consistency and maximize the retention of the representativeness of the decision peak. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of each peak representative decision to generate the final consensus decision.

[0010] The final consensus decision is sent to the execution layer and the execution effect is monitored. The credibility index of each agent is updated in reverse based on the execution effect.

[0011] Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a reliability index, including:

[0012] Extract the decision adoption frequency and decision execution success rate of each agent in each historical decision cycle from the historical contribution records, multiply them to obtain the contribution degree and construct a contribution degree sequence, apply a time decay function to the contribution degree sequence for weighting, and calculate the weighted historical contribution degree of each agent;

[0013] Extract the decision deviation magnitude and deviation correction response speed of each agent from the deviation correction record, construct the deviation behavior feature vector, and calculate the deviation penalty coefficient of each agent by mapping the deviation behavior feature vector to the normalization space.

[0014] The weighted historical contribution value and the deviation penalty coefficient are combined nonlinearly to generate a preliminary reliability score for each agent. The preliminary reliability score is then globally normalized and a stability constraint transformation is applied to obtain the credibility index for each agent.

[0015] Extract the decision objective vector of the decision scheme, analyze the complementarity and conflict of decisions among the agents and construct a cooperative relationship graph, calculate node centrality and edge weight distribution to generate a cooperative contribution factor, and perform nonlinear fusion with the credibility index to obtain a dynamic trust weight, including:

[0016] The decision schemes of each agent are mapped in the feature space to extract the decision target vector and the constraint boundary set. The proportion of the intersection between the projected correlation of the decision target vectors and the constraint boundary set between the agent pairs is calculated. When the projected correlation is positive and the intersection proportion is lower than the complementarity judgment threshold, it is marked as a complementary relationship. When the projected correlation is negative or the intersection proportion is higher than the conflict judgment threshold, it is marked as a conflict relationship.

[0017] Using each agent as a node and complementary and conflicting relationships as connecting edges, a cooperative relationship graph is constructed. For complementary relationships, the positive edge weight is calculated using the cooperative gain degree of the decision objective vector, and for conflicting relationships, the negative edge weight is calculated using the overlapping conflict intensity of the constraint boundary, forming a signed edge weight distribution.

[0018] Based on the edge weight distribution, perform multi-hop neighborhood aggregation operation, accumulate the positive edge weights and negative edge weights received by each node, and calculate the node centrality by combining the node in-degree ratio.

[0019] The collaboration contribution factor is generated by weighted fusion of the variance features of node centrality and edge weight distribution; a nonlinear fusion mapping function is constructed, which takes the collaboration contribution factor and the trust index as two-dimensional inputs and achieves nonlinear coupling through hyperbolic tangent transformation to generate dynamic trust weights.

[0020] Based on the edge weight distribution, a multi-hop neighborhood aggregation operation is performed, accumulating the positive and negative edge weights received by each node, and calculating the node centrality by combining the node's in-degree ratio, including:

[0021] Starting from each node, a breadth-first traversal is performed along the directed edges in the collaboration graph to record the shortest hop count to reach neighboring nodes. Based on the shortest hop count, the neighboring nodes are divided into different levels of neighborhood rings to construct a hierarchical neighborhood structure. A weight decay factor based on topological distance is set for each level of neighborhood ring.

[0022] For each node, traverse the incoming edges it receives, identify the positive and negative edge weights carried by the incoming edges, obtain the corresponding weight decay factor according to the neighborhood ring level to which the starting point of the incoming edge belongs, perform weighted calculations to obtain weighted positive and weighted negative values, accumulate all weighted positive values ​​to obtain the positive aggregation amount, and accumulate all weighted negative values ​​to obtain the negative aggregation amount.

[0023] The ratio of the number of incoming edges to the number of outgoing edges of each node is calculated to obtain the node's in-degree ratio. The net influence is obtained by subtracting the negative aggregation from the positive aggregation. The net influence is then weighted and merged with the node's in-degree ratio to generate node centrality.

[0024] Based on the decision objective vector of the aforementioned decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weights to generate peak-representative decisions, including:

[0025] The decision target vectors of each agent are mapped to a unified decision feature space. The feature distance between the decision target vectors is calculated and a distance matrix is ​​constructed. The kernel function bandwidth parameter is determined based on the distance matrix. Local kernel functions are constructed with each decision target vector as the core and superimposed to generate a global kernel density estimation field.

[0026] The global kernel density estimation field is spatially gridded and sampled. Density estimates are calculated at grid nodes to form a density scalar field. Morphological watershed segmentation is performed on the density scalar field to identify the boundaries of the density gradient convergence basins. The maximum density point of each convergence basin is extracted as the peak of the decision peak. The decision peak region is extended along the peak of the decision peak to the boundary of the convergence basin.

[0027] The dynamic trust weights of agents within each decision peak region are extracted. Weighted centroid calculation is performed on the decision target vector within the decision peak region to obtain the centroid decision vector. The spatial variance and distribution skewness of the decision target vector within the decision peak region are calculated to construct a regional stability correction factor. The centroid decision vector is then corrected to obtain the peak representative decision of each decision peak region.

[0028] A fusion optimization model is constructed based on the representative decisions of each peak, with the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of the representative decisions of each peak, generating the final consensus decision, including:

[0029] Each peak represents a decision and is constructed into a peak decision matrix. Singular value decomposition is then performed, and the left singular vectors corresponding to the principal singular values ​​are spanned into a globally consistent constraint subspace.

[0030] Extract the cumulative dynamic trust weight of the agent within the decision peak region corresponding to each peak and the peak height of the kernel density. Multiply the cumulative dynamic trust weight by the peak height of the kernel density and apply a logarithmic transformation to obtain the peak influence measure.

[0031] The fusion coefficients of each peak representing the decision are initialized with equal weights. The first objective function is constructed as the projection modulus of the fused decision in the global consistency constraint subspace. The second objective function is constructed as the inner product of the fusion coefficients and the peak influence measure. The constraint condition is set that the sum of the fusion coefficients is a unit value and each component is non-negative, thus obtaining the bi-objective optimization problem.

[0032] A multi-objective evolutionary algorithm is used to solve the bi-objective optimization problem. The population is screened by fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, thus obtaining the Pareto optimal solution set.

[0033] In the Pareto optimal solution set, the solution with the largest sum of the normalized values ​​of the two objective functions is selected, and the corresponding fusion coefficients are extracted. A convex combination operation is then performed on each peak to represent the decision to generate the final consensus decision.

[0034] A multi-objective evolutionary algorithm is employed to solve the bi-objective optimization problem. The population is screened using fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, yielding a Pareto optimal solution set, including:

[0035] An initial population is generated, each individual is encoded as a fusion coefficient vector, the first objective function value and the second objective function value of each individual are calculated, a domination counter is established to record the number of times each individual is dominated, individuals with a counter of zero are added to the first non-dominated level, the domination list of individuals in the first non-dominated level is traversed and the counter of dominated individuals is decremented by one, and the process is repeated to form subsequent non-dominated levels.

[0036] The sum of the absolute values ​​of the differences between the individual and its neighbor in the two objective function values ​​within each non-dominated hierarchy is calculated as the crowding distance, and an infinite crowding distance is assigned to the individuals at the end of the hierarchy.

[0037] The crowding degree comparison selection operator is used to select parent individuals, and a uniform crossover operation is performed on the parent individuals to generate intermediate individuals and an adaptive mutation operation is performed. The mutated individuals are then processed into units to generate the offspring population.

[0038] The parent and offspring populations are merged and non-dominated sorting and crowding distance calculation are re-executed. Individuals at each non-dominated level are selected sequentially according to the elite strategy until the preset population size is reached. The last non-dominated level is truncated according to the crowding distance to form a new generation population.

[0039] Calculate the maximum movement distance of an individual in the first non-dominated level in the target space relative to the previous generation. When the maximum movement distance is lower than a preset threshold for multiple consecutive generations, convergence is determined, and all individuals in the first non-dominated level are output as the Pareto optimal solution set.

[0040] A second aspect of this invention provides an intelligent weighted consensus system for multi-agent decision-making, comprising:

[0041] The data acquisition unit is used to collect the decision-making schemes submitted by each agent and simultaneously obtain the historical contribution records and deviation correction records of each agent.

[0042] The credibility index calculation unit is used to calculate the long-term reliability of each agent based on the historical contribution record and the deviation correction record to obtain the credibility index.

[0043] The trust weight calculation unit is used to extract the decision target vector of the decision scheme, analyze the decision complementarity and conflict between the agents and construct a cooperative relationship graph, calculate the node centrality and edge weight distribution to generate a cooperative contribution factor, and perform nonlinear fusion with the trust index to obtain a dynamic trust weight.

[0044] The peak decision unit is used to identify high-density decision regions and form multiple decision peaks based on the decision target vector of the decision scheme through kernel density estimation, and to perform decision weighting in each decision peak region using the dynamic trust weight to generate peak representative decisions.

[0045] The consensus optimization unit is used to build a fusion optimization model based on the representative decisions of each peak. With the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks, it searches the Pareto optimal solution space by adaptively adjusting the fusion coefficients of the representative decisions of each peak to generate the final consensus decision.

[0046] The execution feedback unit is used to send the final consensus decision to the execution layer and monitor the execution effect, and update the credibility index of each agent in reverse according to the execution effect.

[0047] A third aspect of the present invention provides an electronic device, comprising:

[0048] processor;

[0049] Memory used to store processor-executable instructions;

[0050] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0051] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0052] This invention, based on an intelligent weighted consensus method, fully integrates the long-term reliability and decision complementarity of each agent into dynamic trust weights. This effectively suppresses the negative impact of unreliable or malicious agents on the consensus results, significantly improving decision accuracy and robustness. Historical contribution records and deviation correction records are used to calculate a trust index. Combined with a collaborative relationship graph, complementarity and conflict are captured, ensuring that trust allocation considers both individual historical performance and team collaboration value. This achieves decentralized trust assessment and dynamic weight adjustment in a distributed decision-making environment, avoiding bias caused by a single source of trust.

[0053] By using kernel density estimation to identify high-density regions of decision schemes, multiple decision peaks are formed, avoiding the oversimplification of multimodal decision distribution by traditional weighted averaging. Within each decision peak, dynamic trust weights are used to generate peak representative decisions, preserving the representativeness of the viewpoints of different agent groups while suppressing marginal viewpoints, ensuring the rationality of each potential consensus direction.

[0054] A fusion optimization model is established with the dual optimization objectives of maximizing global consistency and maximizing the representativeness retention of the decision peak. The final consensus decision is generated by adaptively searching the Pareto optimal solution space. The fusion coefficients are automatically adjusted without the need for pre-setting fusion rules, achieving a synergistic trade-off between collective wisdom and individual strengths. The consensus result achieves a balance between global and local factors in different task scenarios, avoiding excessive local optima or global compromise. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the intelligent weighted consensus method for multi-agent decision-making according to an embodiment of the present invention.

[0056] Figure 2 This is a flowchart illustrating the method for calculating the dynamic trust weights of each agent in an embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0059] Figure 1 This is a flowchart illustrating an intelligent weighted consensus method for multi-agent decision-making according to an embodiment of the present invention. The present invention provides an intelligent weighted consensus method for multi-agent decision-making, comprising:

[0060] Collect decision proposals submitted by each agent and simultaneously obtain historical contribution records and deviation correction records of each agent;

[0061] Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a trust index; the decision target vector of the decision scheme is extracted, the decision complementarity and conflict between agents are analyzed and a cooperation relationship graph is constructed, the node centrality and edge weight distribution are calculated to generate a cooperation contribution factor, and the dynamic trust weight is obtained by nonlinear fusion with the trust index.

[0062] Based on the decision target vector of the decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weight to generate peak representative decisions.

[0063] A fusion optimization model is constructed based on the decisions of each peak representative. The dual optimization objectives are to maximize global consistency and maximize the retention of the representativeness of the decision peak. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of each peak representative decision to generate the final consensus decision.

[0064] The final consensus decision is sent to the execution layer and the execution effect is monitored. The credibility index of each agent is updated in reverse based on the execution effect.

[0065] Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a reliability index, including:

[0066] Extract the decision adoption frequency and decision execution success rate of each agent in each historical decision cycle from the historical contribution records, multiply them to obtain the contribution degree and construct a contribution degree sequence, apply a time decay function to the contribution degree sequence for weighting, and calculate the weighted historical contribution degree of each agent;

[0067] Extract the decision deviation magnitude and deviation correction response speed of each agent from the deviation correction record, construct the deviation behavior feature vector, and calculate the deviation penalty coefficient of each agent by mapping the deviation behavior feature vector to the normalization space.

[0068] The weighted historical contribution value and the deviation penalty coefficient are combined nonlinearly to generate a preliminary reliability score for each agent. The preliminary reliability score is then globally normalized and a stability constraint transformation is applied to obtain the credibility index for each agent.

[0069] Extract the frequency of decision adoption and the success rate of decision execution for each agent within each historical decision-making cycle from historical contribution records. Multiply these two values ​​to obtain the agent's contribution within the corresponding cycle. Let the th... The agent in the th... The frequency of decision adoption within a historical decision-making cycle is The success rate of decision execution is The contribution during this period is The contributions of each historical period are arranged in chronological order to form a contribution sequence. ,in This represents the total number of historical decision-making cycles. The frequency of decision adoption reflects the proportion of solutions proposed by the agent that are accepted by the overall consensus process, while the success rate of decision execution measures the degree to which the adopted solutions achieve the expected goals during actual execution. The product of the two reflects the agent's substantial contribution to historical decision-making.

[0070] Applying a time decay function to the contribution sequence for weighting aims to give greater influence to recent decisions on reliability assessment, while gradually reducing the weight of older historical records over time. This allows the reliability index to dynamically reflect the agent's current state rather than relying solely on distant history. The time decay function uses an exponential decay form, where the latest period number corresponding to the current moment is denoted as . , No. The time interval between each cycle and the current period is The attenuation coefficient is ( ), then the first The time decay weight of the period is The weighted historical contribution of each agent is calculated as follows: This involves weighting the contribution sequence according to time decay weights, so that agents with good recent performance receive higher weighted historical contributions. Decay coefficient. The value of can be adjusted according to the rate of change of the decision-making environment in the specific application scenario. When the environment changes rapidly, a smaller value is taken to increase the proportion of recent weights, while a larger value can be taken when the environment is relatively stable to make full use of long-term historical information.

[0071] Extract the decision deviation magnitude and deviation correction response speed of each agent from the deviation correction records to construct a deviation behavior feature vector. The decision deviation magnitude measures the distance between the agent's submitted decision and the final consensus decision, reflecting the degree of deviation between the agent's decision tendency and the group consensus. The deviation correction response speed measures the timeliness of the agent's adjustment of its decision after receiving feedback signals; a faster response speed indicates a stronger self-correction ability. Let the... The agent in the th... The magnitude of decision deviation within a historical period is The deviation correction response speed is Then the feature vector of the deviation behavior is It covers the deviation magnitude and response speed information for each historical period.

[0072] The deviation behavior feature vector is mapped to a normalized space to eliminate the interference of different dimensions and numerical ranges on subsequent calculations. For the deviation amplitude component, min-max normalization is used to map it to... For the interval, the larger the deviation amplitude, the higher the normalized value. For the deviation correction response speed component, minimum-maximum normalization is also performed, but the faster the response speed, the higher the normalized value. After normalization, the average deviation amplitude across all historical periods is considered. With average response speed Calculate the deviation penalty coefficient The calculation logic for the deviation penalty coefficient is as follows: the larger the deviation, the heavier the penalty on reliability; while a faster response speed can partially offset the negative impact of the deviation. Specifically, ,in The penalty weight for the deviation magnitude, Both are positive hyperparameters used as compensation weights for response speed, and typically satisfy the following conditions: The primary focus is on ensuring the effectiveness of punishment for deviant behavior. After calculation, [the following is done / then...] Perform truncation to ensure its range is within Within the specified range, avoid negative or zero values ​​that could cause subsequent calculations to fail.

[0073] Weighted historical contribution With deviation penalty coefficient Perform nonlinear combination operations to generate the first... Preliminary reliability score of the agent The nonlinear combination adopts an exponential fusion form. ,in and These are fusion indices representing contribution and bias penalty coefficients, respectively, both positive parameters. Using a power-law approach instead of simple linear weighting allows for higher scores when both dimensions are high, while negative performance in either dimension strongly suppresses the overall score, thus more rigorously selecting agents that perform well in both contribution and bias behavior. Fusion Index and The relative size of the two dimensions determines their relative importance in the initial reliability score, and the emphasis on contribution and stability can be adjusted according to the application scenario.

[0074] The initial reliability scores of each agent are globally normalized, mapping all agent scores to a unified numerical range to eliminate the incomparability problem caused by differences in historical record length or data size among different agents. Global normalization employs min-max normalization on the set of initial reliability scores for all agents to obtain normalized scores. Based on normalization, a further stability constraint transformation is applied to prevent drastic fluctuations in the credibility index between adjacent evaluation periods, thereby ensuring the continuity and stability of the consensus decision-making process. The stability constraint transformation introduces a smoothing factor. ( The normalized score of the current period is weighted and fused with the credibility index of the previous period: Let the first period be... The credibility index of each agent in the previous evaluation period was: Then the credibility index for the current period is updated to Smoothing factor The larger the value, the more direct the impact of the current rating on the credibility index; The smaller the value, the stronger the inertia of the historical reliability index, and the smoother the changes. For agents participating in decision-making for the first time, its... Initialize the trust index to the mean of the normalized scores of all agents to avoid excessively high or low initial trust indices due to a lack of historical records, and ensure that newly added agents can participate in subsequent dynamic trust weight calculations from a reasonable starting point.

[0075] Figure 2This is a flowchart illustrating the method for calculating dynamic trust weights for each agent according to an embodiment of the present invention. The method involves extracting the decision objective vector of the decision scheme, analyzing the complementarity and conflict of decisions among the agents and constructing a cooperative relationship graph, calculating node centrality and edge weight distribution to generate a cooperative contribution factor, and nonlinearly fusing this factor with the trust index to obtain the dynamic trust weights. The process includes:

[0076] The decision schemes of each agent are mapped into the feature space to extract the decision objective vector and the set of constraint boundaries.

[0077] Calculate the proportion of the intersection between the projected correlation of the decision target vectors between agent pairs and the constraint boundary set. When the projected correlation is positive and the intersection proportion is lower than the complementarity determination threshold, it is marked as a complementary relationship. When the projected correlation is negative or the intersection proportion is higher than the conflict determination threshold, it is marked as a conflict relationship.

[0078] Using each agent as a node and complementary and conflicting relationships as connecting edges, a cooperative relationship graph is constructed. For complementary relationships, the positive edge weight is calculated using the cooperative gain degree of the decision objective vector, and for conflicting relationships, the negative edge weight is calculated using the overlapping conflict intensity of the constraint boundary, forming a signed edge weight distribution.

[0079] Based on the edge weight distribution, perform multi-hop neighborhood aggregation operation, accumulate the positive edge weights and negative edge weights received by each node, and calculate the node centrality by combining the node in-degree ratio.

[0080] The collaboration contribution factor is generated by weighted fusion of the variance features of node centrality and edge weight distribution; a nonlinear fusion mapping function is constructed, which takes the collaboration contribution factor and the trust index as two-dimensional inputs and achieves nonlinear coupling through hyperbolic tangent transformation to generate dynamic trust weights.

[0081] When performing feature space mapping on the decision schemes submitted by each agent, each decision scheme is projected from the original description space to a unified high-dimensional feature space, from which a decision objective vector and a set of constraint boundaries are extracted. The decision objective vector describes the direction and magnitude of the target that the agent expects to achieve in the current decision task, while the set of constraint boundaries describes the range of resource, temporal, and logical constraints imposed on the agent when executing the decision. Feature space mapping employs a combination of linear projection and nonlinear embedding to ensure that different types of decision schemes are comparable within the same space.

[0082] For any two intelligent agents and Calculate the projection correlation between their decision target vectors. Specifically, it calculates the cosine similarity between two target vectors in the feature space; it also calculates the proportion of the intersection of their constraint boundary sets. That is, the ratio of the metric of the intersection of two constraint boundary sets to the metric of their union. When and (in When the complementarity judgment threshold is reached, the agent is judged. and The existence of a complementary relationship means that the two have the same decision-making goals but little overlap in their constraint spaces, possessing the potential for synergistic complementarity; when or (in When the threshold for conflict determination is reached, a conflict is determined between the two, meaning that their decision-making objectives are contradictory or their constraint boundaries highly overlap, leading to resource competition. Complementarity determination threshold. Conflict determination threshold It can be pre-calibrated according to the task type and the size of the agent, or it can be dynamically adjusted based on historical distribution during operation.

[0083] After the agents have completed the relationship labeling, a collaborative relationship graph is constructed with each agent as a node. For node pairs marked as complementary, calculate the cooperative gain degree of their decision objective vectors. Specifically, it is the projection increment of the weighted superposition of the two target vectors onto the direction of the objective function, reflecting the additional benefit that can be obtained when the two make joint decisions compared to individual decisions; Assign the weight of the positive edge to the corresponding connection edge in the graph. For node pairs marked as conflicting, calculate the overlap conflict strength of their constraint boundaries. Specifically, it is the constraint density-weighted integral within the intersection region of the constraint boundaries of the two constraints, reflecting the degree of decision interference caused by constraint overlap; with The negative edge weights are assigned to the graph as negative edge weights, thus forming a signed edge weight distribution. The positive and negative edge weights together constitute the signed adjacency structure of the cooperative relationship graph, which can simultaneously characterize the cooperative potential and competitive risks among agents.

[0084] Based on the aforementioned signed edge weight distribution, a multi-hop neighborhood aggregation operation is performed on each node in the graph. In the first hop aggregation, the node... Sum the weights of the forward edges passed from all its direct neighboring nodes. The sum of the absolute values ​​of the negative edge weights In the second and higher hop aggregation, the contributions of indirect neighboring nodes are further incorporated, and an aggregation weight that decays with the number of hops is adopted to avoid excessive interference from distant nodes in the centrality calculation. This is combined with the in-degree ratio of the nodes. (i.e., node) The ratio of out-degree to in-degree reflects the agent's tendency to actively output in the cooperative network. This multi-hop aggregation result is then compared with... Weighted fusion yields node centrality. The calculation method is as follows ,in It is an adjustment index for the in-degree ratio, used to control the intensity of the influence of node initiative on centrality. A positive value indicates that the agent plays a dominant role in the cooperative network, while a negative value indicates that it causes more conflicts in the network or is in a passive receiving state.

[0085] Node centrality Variance characteristics of edge weight distribution (i.e., node) The variances of all associated edge weights (reflecting the stability and consistency of the agent's collaborative relationships) are weighted and fused to generate a collaborative contribution factor. The calculation method is as follows ,in This is the variance adjustment coefficient. When... When the value is large, it indicates that the cooperative relationships between the agent and its different neighbors vary significantly, and its cooperative contribution factor is somewhat suppressed, reflecting a preference for cooperative stability; when... When the value is relatively small, it indicates that the cooperative relationship of the agent is relatively balanced and the cooperative contribution factor is fully preserved.

[0086] Construct a nonlinear fusion mapping function to determine the collaborative contribution factor. With credibility index Assuming a two-dimensional input, nonlinear coupling is achieved through hyperbolic tangent transform to generate dynamic trust weights. Specifically, and After being linearly scaled, the inputs are given a hyperbolic tangent function, and the fusion method is as follows: ,in , These are the linear fusion coefficients of the collaboration contribution factor and the credibility index, respectively. The cross-coupling coefficients are used to capture the interaction enhancement or inhibition effects between collaborative contributions and credibility. The hyperbolic tangent transform compresses the fusion result to... The interval is then mapped to the affine transformation. The interval is used to obtain the final normalized dynamic trust weight. This weight comprehensively reflects the agent's structural position, cooperative stability, and historical reliability in the cooperative network. It can adaptively adjust with the dynamic evolution of the cooperative relationship graph and the periodic update of the reliability index, providing a reliable weight basis for subsequent decision weighting and peak representative decision generation.

[0087] Based on the edge weight distribution, a multi-hop neighborhood aggregation operation is performed, accumulating the positive and negative edge weights received by each node, and calculating the node centrality by combining the node's in-degree ratio, including:

[0088] Starting from each node, a breadth-first traversal is performed along the directed edges in the collaboration graph to record the shortest hop count to reach neighboring nodes. Based on the shortest hop count, the neighboring nodes are divided into different levels of neighborhood rings to construct a hierarchical neighborhood structure. A weight decay factor based on topological distance is set for each level of neighborhood ring.

[0089] For each node, traverse the incoming edges it receives, identify the positive and negative edge weights carried by the incoming edges, obtain the corresponding weight decay factor according to the neighborhood ring level to which the starting point of the incoming edge belongs, perform weighted calculations to obtain weighted positive and weighted negative values, accumulate all weighted positive values ​​to obtain the positive aggregation amount, and accumulate all weighted negative values ​​to obtain the negative aggregation amount.

[0090] The ratio of the number of incoming edges to the number of outgoing edges of each node is calculated to obtain the node's in-degree ratio. The net influence is obtained by subtracting the negative aggregation from the positive aggregation. The net influence is then weighted and merged with the node's in-degree ratio to generate node centrality.

[0091] After the collaborative relationship graph is constructed, it is necessary to accurately quantify the influence of each node in the graph in order to capture the differentiated contributions of neighboring nodes to the target node under different topological distances. To this end, a multi-hop neighborhood aggregation operation mechanism is introduced. Based on breadth-first traversal, the neighborhood structure of each node is expanded layer by layer, and combined with the topological decay characteristics of edge weights, fine-grained calculation of node centrality is realized.

[0092] Using a node in the collaboration relationship graph Starting from the node, perform a breadth-first traversal along the directed edges in the graph. During the traversal, record the nodes from which the graph originates. Starting from the shortest hop count to each reachable neighbor node, neighbor nodes with a shortest hop count of 1 form the first-level neighborhood cycle, neighbor nodes with a shortest hop count of 2 form the second-level neighborhood cycle, and so on, until all reachable nodes are traversed and covered, forming a node-based neighborhood cycle. A hierarchical neighborhood structure centered on [the element]. Let the upper limit of the maximum traversal jump be [the maximum number of traversals]. Then the neighborhood ring level number satisfy For each neighborhood ring, a weight decay factor is set based on its topological distance. The attenuation factor decreases monotonically with increasing level, specifically in the form of: ,in This is the decay rate control parameter, and its value is greater than 0. When... hour, This means that the edge weights of directly adjacent nodes are not attenuated; as the level increases, the attenuation factor decreases exponentially, and the influence of distant neighboring nodes on the target node is moderately compressed, thus reflecting the topological rationality of "strong influence of nearest neighbors and weak influence of distant neighbors".

[0093] After the hierarchical neighborhood structure is established, for nodes Traverse all its incoming edges, that is, all edges connected to the node For directed edges ending at node A, for each incoming edge, identify the type of edge weight it carries: if the edge represents a cooperative gain relationship between two nodes, it carries a positive edge weight; if the edge represents a conflict relationship between two nodes, it carries a negative edge weight. Let the starting point of an incoming edge be node B. The weight of the positive edge carried by this edge is (If it exists), the absolute value of the negative edge weight is (If it exists), node Relative to node The neighborhood ring level to which it belongs is The corresponding weight decay factor is The positive weights of the incoming edges are then decayed to obtain the weighted positive value of the edge. The negative weight of the incoming edge is then decayed to obtain the weighted negative value of the edge. .

[0094] For nodes The node is obtained by summing the weighted positive values ​​of all incoming edges. forward polymerization amount The calculation formula is ,in Represents all nodes The set of incoming edges from the endpoint, for each node Sum the weighted negative values ​​of all incoming edges to obtain the node. negative polymerization amount The calculation formula is Compared to a simple accumulation method that does not introduce topological decay, the positive and negative aggregation quantities here fully consider the natural decay effect of the distance between neighboring nodes on the propagation of influence, making the aggregation results more reflective of the true topological structure characteristics of the graph.

[0095] After calculating the positive and negative aggregation amounts, the in-degree ratio of each node is further calculated. Let the node... The number of incoming edges is The number of outgoing edges is Then the in-degree ratio of the node Defined as ,in To prevent extremely small smoothing constants with a denominator of zero, the in-degree / out-degree ratio is used. This reflects the ratio of a node's "degree of dependence" to its "degree of actively exerting influence" within the graph. When A larger value indicates that the node receives more cooperative relationships from other nodes, occupying a core position in the graph that is widely relied upon; when When the value is relatively small, it indicates that the node mainly exerts influence outwards and is relatively less dependent on itself.

[0096] Subtracting the negative aggregation from the positive aggregation yields the node. Net influence ,Right now Net influence comprehensively reflects the node's overall impact. The difference between the positive cooperative effect and the negative conflict effect received within the multi-hop neighborhood is the largest value. This indicates that the node is in a more favorable position in the cooperative relationship graph, receives stronger positive support, and is less affected by conflict interference.

[0097] Net impact Ratio of node in-degree to out-degree Perform weighted fusion to generate node centrality The fusion method is as follows: ,in The weighted average of net influence The fusion weight for the in-degree ratio, This is the adjustment index for the in-degree ratio, used to control the nonlinearity of the in-degree ratio's contribution to node centrality. and satisfy And all values ​​are greater than 0, which can be adjusted according to the specific focus of the multi-agent decision-making scenario: if more attention is paid to the net cooperative effect of nodes in the cooperative network, the value can be appropriately increased. If more attention is paid to the structural position of nodes in the graph, the size can be appropriately increased. .

[0098] The node centrality obtained through the above multi-hop neighborhood aggregation operation This approach integrates the positive and negative edge weights of multi-layer neighborhoods under topological distance decay, as well as the in-degree and out-degree structural features of nodes in the graph. This allows for a more comprehensive and accurate characterization of the overall influence and status of each agent node within the collaborative relationship graph. Subsequently, when calculating the collaborative contribution factor, this node centrality will be used as the core input, further fused with edge weight variance information. This will ultimately generate a collaborative contribution factor for dynamic trust weight calculation, ensuring that the multi-agent decision-making weighting process fully considers the true contribution capability and structural value of each agent in the collaborative network.

[0099] Based on the decision objective vector of the aforementioned decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weights to generate peak-representative decisions, including:

[0100] The decision target vectors of each agent are mapped to a unified decision feature space. The feature distance between the decision target vectors is calculated and a distance matrix is ​​constructed. The kernel function bandwidth parameter is determined based on the distance matrix. Local kernel functions are constructed with each decision target vector as the core and superimposed to generate a global kernel density estimation field.

[0101] The global kernel density estimation field is spatially gridded and sampled. Density estimates are calculated at grid nodes to form a density scalar field. Morphological watershed segmentation is performed on the density scalar field to identify the boundaries of the density gradient convergence basins. The maximum density point of each convergence basin is extracted as the peak of the decision peak. The decision peak region is extended along the peak of the decision peak to the boundary of the convergence basin.

[0102] The dynamic trust weights of agents within each decision peak region are extracted. Weighted centroid calculation is performed on the decision target vector within the decision peak region to obtain the centroid decision vector. The spatial variance and distribution skewness of the decision target vector within the decision peak region are calculated to construct a regional stability correction factor. The centroid decision vector is then corrected to obtain the peak representative decision of each decision peak region.

[0103] After transforming the decision proposals submitted by each agent into decision target vectors, these vectors need to be uniformly mapped to the same decision feature space to eliminate differences in the dimensions, units, and scales of decision representations among different agents. During the mapping process, each decision target vector undergoes standardization, ensuring that the mean of each feature dimension is zero and the variance is a unit value, thereby guaranteeing the fairness of subsequent distance calculations. After mapping, the feature distances between each pair of decision target vectors are calculated, specifically using Euclidean distance, filling the distance matrix with the pairwise distance values ​​between all agent decision target vectors. Among them The Line number Column elements Indicates the first The first agent and the second The Euclidean distance between the decision target vectors of each agent.

[0104] Based on distance matrix The kernel function bandwidth parameters are determined using an adaptive bandwidth estimation method. Specifically, calculate the vector of each decision objective to its nth... The distance to the nearest neighbor is used, and the median of these distance values ​​for all agents is taken as the global bandwidth parameter. ,in Take the total number of agents The integer value of the square root, i.e. This method of determining bandwidth based on the data's own distribution characteristics can maintain sufficient smoothness in sparse regions of decision points, while preserving local structural details in dense regions. It avoids the merging of multiple decision peaks due to excessive bandwidth or the appearance of spurious peaks due to insufficient bandwidth. Using each decision target vector as the core, a Gaussian kernel function is used to construct a local kernel density contribution. The local kernel functions of all agents are then superimposed to generate a global kernel density estimation field covering the entire decision feature space. Its expression is ,in For any sampling point in the decision feature space, For the first The decision objective vector of an agent. For the dimension of the decision feature space, This is the standard Gaussian kernel function.

[0105] After obtaining the global kernel density estimation field, uniform grid sampling is performed on the decision feature space. The grid resolution is adaptively set according to the range of the decision feature space and the number of agents, typically dividing each dimension into no fewer than 50 sampling nodes to ensure the accuracy of the density field. At each grid node, calculations are performed... The density estimates of all grid nodes are organized into a density scalar field. A morphological watershed segmentation algorithm is then applied to this density scalar field: the density scalar field is negatively evaluated as a topographic elevation map, with high-density areas corresponding to low-lying areas and low-density areas to high-lying areas. The watershed algorithm starts from local minima (i.e., local maxima of the original density field), simulates the water injection process, and identifies the watershed boundaries between different confluence basins. The grid node with the highest density within each confluence basin is extracted as the decision peak apex, and its corresponding feature space coordinates represent the center position of that decision peak. Extending outwards from the apex of each decision peak, and bounded by the watershed boundary, all grid nodes within the boundary and the decision target vectors falling within it are assigned to the corresponding decision peak regions, thus forming several sets of non-overlapping decision peak regions.

[0106] Within each decision peak region, extract the set of agent IDs belonging to that region, denoted as . ,in Number the decision peak region. Read. Dynamic trust weights for each agent Weighted centroid calculation is performed on the decision objective vectors of each agent within the region to obtain the centroid decision vector. The calculation method is as follows Weighted centroid calculation makes the influence of agents with high credibility and large collaborative contributions more prominent in peak representative decisions, while the influence of agents with large historical bias or weak collaboration is moderately suppressed, thereby improving the quality of peak representative decisions.

[0107] Relying solely on weighted centroids cannot fully reflect the distribution characteristics of decision schemes within the decision peak region. If the spatial distribution of the agent's decision target vector within a certain decision peak region is relatively dispersed, or exhibits a significant skewed distribution, the centroid decision vector will deviate from its true representative position within that region. Therefore, it is necessary to calculate the spatial variance of the decision target vector within the decision peak region. With distribution skewness Constructing a regional stability correction factor Spatial variance It reflects the dispersion of decision-making options within a region and is calculated as the mean of the squared weighted distances between all decision target vectors belonging to that region and the centroid decision vector. Distribution skewness The asymmetry of the decision target vector on both sides of the centroid is reflected by the ratio of the third-order central moment to the 3 / 2 power of the variance.

[0108] Regional stability correction factor Combining variance and skewness, the specific construction is as follows: ,in This is the variance penalty coefficient. , where is the skewness penalty coefficient, and both are positive hyperparameters that can be adjusted according to the decision distribution characteristics of the actual multi-agent system. When decision schemes within the region are highly concentrated and symmetrically distributed, and All are close to zero. The value approaches 1, with a very small correction margin; when decision-making options within the region are scattered or severely skewed, If the value is significantly less than 1, a large contraction correction is applied to the centroid decision vector, pulling it closer to the peak of the decision peak to reduce the interference of outlier decision schemes on the peak representative decision.

[0109] For the centroid decision vector When making corrections, compare them with the coordinates of the decision peak. Weighted interpolation is performed, and the corrected peak value represents the decision. The calculation method is as follows .when When the value approaches 1, the peak value indicates that the decision-making process is primarily based on the weighted centroid; when... When the density is low, the peak value represents the decision that converges towards the point with the highest density, ensuring that its representativeness is not distorted due to the uneven distribution of decision options within the region. Through the above process, a peak value representative decision is generated for each decision peak region, which serves as the input to the subsequent fusion optimization model and participates in the generation of global consensus decisions.

[0110] A fusion optimization model is constructed based on the representative decisions of each peak, with the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of the representative decisions of each peak, generating the final consensus decision, including:

[0111] Each peak represents a decision and is constructed into a peak decision matrix. Singular value decomposition is then performed, and the left singular vectors corresponding to the principal singular values ​​are spanned into a globally consistent constraint subspace.

[0112] Extract the cumulative dynamic trust weight of the agent within the decision peak region corresponding to each peak and the peak height of the kernel density. Multiply the cumulative dynamic trust weight by the peak height of the kernel density and apply a logarithmic transformation to obtain the peak influence measure.

[0113] The fusion coefficients of each peak representing the decision are initialized with equal weights. The first objective function is constructed as the projection modulus of the fused decision in the global consistency constraint subspace. The second objective function is constructed as the inner product of the fusion coefficients and the peak influence measure. The constraint condition is set that the sum of the fusion coefficients is a unit value and each component is non-negative, thus obtaining the bi-objective optimization problem.

[0114] A multi-objective evolutionary algorithm is used to solve the bi-objective optimization problem. The population is screened by fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, thus obtaining the Pareto optimal solution set.

[0115] In the Pareto optimal solution set, the solution with the largest sum of the normalized values ​​of the two objective functions is selected, and the corresponding fusion coefficients are extracted. A convex combination operation is then performed on each peak to represent the decision to generate the final consensus decision.

[0116] Each peak represents a decision. ( , Construct a peak decision matrix by piecing together columns representing the total number of decision peaks. Each column corresponds to a decision peak region, and the corrected peak value represents the decision vector. Perform singular value decomposition to obtain ,in It is a left singular matrix. It is a diagonal singular value matrix. This is a right singular matrix. Principal singular values. The corresponding left singular vector (Right now The first column captures all peaks representing the direction with the largest variance in the decision-making process, and represents the directional component with the greatest consensus potential in a multi-peak decision distribution. Zhang Cheng's one-dimensional global consistency constraint subspace The projection modulus of the subsequently fused decisions onto this subspace will serve as the core metric for measuring global consistency. In real-world scenarios, if the decision schemes of multiple agents are highly aligned in a certain dimension, the singular value in that direction will be significantly larger than in other directions. Therefore, defining the consistency subspace using the main left singular vector has good data adaptability and eliminates the need for manual specification of the consensus direction.

[0117] Extract the first Dynamic trust weights of all agents within a decision peak region ( Then sum them to obtain the cumulative dynamic trust weight value of the decision peak. Simultaneously, the coordinates of the peak of this decision peak region are extracted. The estimated field density value of the kernel density at that location This is the height of the kernel density peak. Multiplying the two and applying a logarithmic transform, we obtain the [missing information]. Peak influence measurement of each decision peak The introduction of logarithmic transformation effectively suppresses the excessive dominance of a few extremely high-weight or extremely high-density peaks in the fusion result, maintaining a dimensional balance in the influence of each decision peak. The peak influence measures of all decision peaks are then combined into a vector. This vector comprehensively reflects three aspects: the scale of each decision peak region, the credibility of the agent, and the degree of decision concentration, providing meaningful prior guidance for subsequent optimization of the fusion coefficient.

[0118] Initialize the fusion coefficient vector of each peak to represent the decision. ,make For all The consensus decision candidate vector is established, meaning that equal weight allocation serves as the starting point for iteration. The fused consensus decision candidate vector is represented as follows: The first objective function is constructed as the projection modulus of the fused decision onto the globally consistent constraint subspace, i.e. The larger the objective function, the more consistent the fusion decision-making and multi-peak consensus are, and the higher the global consistency. The second objective function is constructed as the inner product of the fusion coefficient vector and the peak influence measurement vector, i.e. The larger the objective function, the more the fusion coefficient tends to assign higher weights to the more influential decision peaks, and the higher the representativeness retention of the decision peaks. The constraints are set as follows: and In other words, the fusion coefficients constitute a simplex constraint, ensuring that the final consensus decision is a convex combination of decisions represented by each peak, with clear physical meaning and interpretable results. This constitutes a bi-objective optimization problem under simplex constraints, with a potential competition between the two objective functions: excessive pursuit of global consistency may ignore some highly influential decision peaks whose directions deviate from the principal singular vector, while excessive pursuit of representativeness preservation may cause the fusion result to deviate from the overall consensus direction.

[0119] A multi-objective evolutionary algorithm (such as the NSGA-II framework) is employed to solve the aforementioned bi-objective optimization problem. During population initialization, an initial set of fusion coefficient configurations is generated by uniform random sampling within the simplex constraint domain, and the population size is set to be sufficient to cover the Pareto front. In each iteration, calculations are performed on each individual in the population. and The function value is used to perform fast non-dominated sorting to divide the population into several non-dominated levels. The first level is the candidate set of Pareto fronts for the current iteration. Within the same non-dominated level, the crowding distance of each individual in the objective function space is calculated. The larger the crowding distance, the sparser the solution is on the Pareto front. Retaining individuals with larger crowding distances helps maintain the diversity of the Pareto front and avoids the solution set from being overly concentrated in a certain local region. Crossover and mutation operations are performed under simplex constraints. During mutation, bounded perturbations are applied to the fusion coefficient components and then normalized to ensure that the constraints are always satisfied. The fusion coefficient configuration is iteratively updated until the Pareto front of the population no longer changes significantly over several generations, or the preset maximum number of iterations is reached. At this point, the Pareto optimal solution set is output. Each element in the set corresponds to a set of fusion coefficient configurations and their corresponding... Target value pair.

[0120] In the Pareto optimal solution set The final solution is selected during the process. This involves selecting all solutions in the solution set. Value and The values ​​are then subjected to min-max normalization to obtain the normalized target value. and Calculate the overall score for each solution. Select the fusion coefficient vector corresponding to the solution with the highest overall score. This selection strategy is equivalent to finding the point closest to the ideal in the normalized target space. along The closest Pareto solution is found, achieving an equilibrium compromise between the two objectives without introducing artificial preference weights. Ultimately, with... Perform convex combination operations on each peak value to obtain the final consensus decision. This final consensus decision fully integrates the global consistency direction in the multi-peak decision distribution while retaining the representative information of each decision peak region, and can serve as a high-quality consensus output sent to the execution layer. In practical multi-agent task scenarios, if a certain decision peak region gathers a large number of highly trustworthy agents and has a significant kernel density peak, its corresponding... A higher value will naturally result in a larger fusion coefficient for the peak during the optimization process, thus tilting the final consensus decision toward this high-quality decision region. This reflects the method's dual sensitivity to agent reliability and decision concentration.

[0121] A multi-objective evolutionary algorithm is employed to solve the bi-objective optimization problem. The population is screened using fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, yielding a Pareto optimal solution set, including:

[0122] An initial population is generated, each individual is encoded as a fusion coefficient vector, the first objective function value and the second objective function value of each individual are calculated, a domination counter is established to record the number of times each individual is dominated, individuals with a counter of zero are added to the first non-dominated level, the domination list of individuals in the first non-dominated level is traversed and the counter of dominated individuals is decremented by one, and the process is repeated to form subsequent non-dominated levels.

[0123] The sum of the absolute values ​​of the differences between the individual and its neighbor in the two objective function values ​​within each non-dominated hierarchy is calculated as the crowding distance, and an infinite crowding distance is assigned to the individuals at the end of the hierarchy.

[0124] The crowding degree comparison selection operator is used to select parent individuals, and a uniform crossover operation is performed on the parent individuals to generate intermediate individuals and an adaptive mutation operation is performed. The mutated individuals are then processed into units to generate the offspring population.

[0125] The parent and offspring populations are merged and non-dominated sorting and crowding distance calculation are re-executed. Individuals at each non-dominated level are selected sequentially according to the elite strategy until the preset population size is reached. The last non-dominated level is truncated according to the crowding distance to form a new generation population.

[0126] Calculate the maximum movement distance of an individual in the first non-dominated level in the target space relative to the previous generation. When the maximum movement distance is lower than a preset threshold for multiple consecutive generations, convergence is determined, and all individuals in the first non-dominated level are output as the Pareto optimal solution set.

[0127] In solving the bi-objective optimization problem, a multi-objective evolutionary algorithm is used to perform a global search of the solution space formed by the fusion coefficient vector. In the initial stage, a batch of individuals is randomly generated, and each individual is encoded as a fusion coefficient vector. Each component of the vector corresponds to the fusion weight of each decision peak, and each component satisfies the nonnegativity and normalization constraints. For each individual in the population, substitute it into the first objective function. (Global Consistency) and the Second Objective Function The representativeness retention of the decision peaks is evaluated to obtain the coordinates of each individual in the target space. The initial population size is usually set to the total number of decision peaks. Several times that of the Pareto front, to ensure that the diversity of solutions covers the entire Pareto front.

[0128] Non-dominated sorting is a core step in screening population structure. A dominance counter is established for each individual, recording how many other individuals in the population are no worse than that individual on both objective functions and strictly better than that individual on at least one objective function; that is, the number of individuals that dominate that individual. An individual with a dominance counter of zero means that no individual in the current population can simultaneously dominate that individual. and These individuals surpass it in two dimensions, thus placing them into the first non-dominated level. Iterate through each individual in the first non-dominated level, finding the list of individuals it dominates, and decrement the dominance counter of each dominated individual by one. When the counter of a dominated individual drops to zero, it indicates that it has become a new non-dominated individual after being removed from the first non-dominated level, and it is added to the second non-dominated level. Repeat the above process to form subsequent non-dominated levels, such as the third and fourth, until all individuals in the population are assigned to a certain level.

[0129] Crowding distance is used to measure the sparsity of the distribution of individuals within the same non-dominated hierarchy in the target space, thereby maintaining solution diversity in subsequent choices. For individuals within each non-dominated hierarchy, according to... and The values ​​are sorted in ascending order. For individuals at the endpoints (i.e., those with the smallest or largest objective function value), an infinite crowding distance is assigned to ensure these extreme solutions have the highest priority during selection, preventing solutions at the ends of the Pareto front from being eliminated. For non-endpoint individuals within a hierarchy, the crowding distance is defined as the sum of the absolute values ​​of the differences in objective values ​​between adjacent individuals in two objective function directions. Suppose an individual is sorted in ascending order according to... The sorted adjacent individuals The absolute value of the difference is ,according to The sorted adjacent individuals The absolute value of the difference is Then the crowding distance for that individual is The greater the crowding distance, the lower the density of solutions around the individual, and preserving the individual helps maintain a uniform distribution of the Pareto front.

[0130] During the parent selection phase, a crowding comparison selection operator is used to select parent individuals from the current population. The comparison rules are: individuals with smaller level numbers are selected first; if two individuals belong to the same non-dominated level, the individual with a larger crowding distance is selected first. A tournament selection method is used, where several individuals are randomly selected each time for crowding comparison, and the winning individuals are added to the parent pool. After the parent pool is full, a uniform crossover operation is performed on the parent individuals: for each component of the fusion coefficient vector of the two parents, it is randomly inherited from both parents with equal probability to generate an intermediate individual. Uniform crossover can effectively mix the fusion strategies of different parents and avoid premature convergence to a local optimum.

[0131] After intermediate individuals are generated, an adaptive mutation operation is performed. The mutation probability is related to the current iteration number. A higher mutation probability is used in the early stages of evolution to enhance global exploration capabilities, while the mutation probability is reduced in the later stages to refine local searches. The mutation operation applies a Gaussian perturbation to randomly selected components of the fusion coefficient vector, with the perturbation amplitude gradually decreasing as the number of iterations increases. Mutated individuals may have negative values ​​or non-zero sums in their components; therefore, normalization is performed on the mutated individuals: all negative components are truncated to zero, and each component is divided by its sum to ensure that the fusion coefficient vector satisfies the probability distribution constraints, thus ensuring the generated consensus decision candidate vector. It is legal in a physical sense. After crossover and mutation, a progeny population is obtained.

[0132] The parent and offspring populations are merged to form a temporary population twice the size of the original population. The complete non-dominated sorting and crowding distance calculation process is then re-executed on this temporary population. Following an elitist strategy, individuals from each non-dominated level are added to the new generation population sequentially, starting from the first non-dominated level, until the remaining capacity is insufficient to accommodate the next complete level. For the last partially selected non-dominated level, individuals within that level are sorted from largest to smallest crowding distance, and the individuals with the largest crowding distances are truncated to fill the remaining population, forming a new generation population with its size restored to the preset value. This elitist retention mechanism ensures that high-quality non-dominated solutions found in previous generations are not lost due to random operations, while maintaining the uniformity of the solution set distribution through crowding distance truncation.

[0133] Convergence is determined based on the movement of individuals in the first non-dominated hierarchy within the target space. After each iteration, the Euclidean distances in the target space between all individuals in the current generation's first non-dominated hierarchy and their nearest counterparts in the previous generation's first non-dominated hierarchy are calculated, and the maximum value is taken as the maximum movement distance for this generation. .when For multiple consecutive generations, the convergence threshold was lower than the preset convergence threshold. When the algorithm has converged, it stops iterating. Convergence threshold. The normalization setting is determined based on the range of the objective function, typically between 0.1% and 1% of the diagonal length of the objective space, to balance convergence speed and solution accuracy. After convergence, all individuals in the first non-dominated level of the current generation are output as the Pareto optimal solution set. Each individual in the solution set corresponds to a fusion coefficient vector. This represents the optimal configuration scheme under different trade-offs between global consistency and the representativeness retention of the decision peak, used for subsequent comprehensive scoring to select the optimal fusion coefficient vector. use.

[0134] A second aspect of this invention provides an intelligent weighted consensus system for multi-agent decision-making, comprising:

[0135] The data acquisition unit is used to collect the decision-making schemes submitted by each agent and simultaneously obtain the historical contribution records and deviation correction records of each agent.

[0136] The credibility index calculation unit is used to calculate the long-term reliability of each agent based on the historical contribution record and the deviation correction record to obtain the credibility index.

[0137] The trust weight calculation unit is used to extract the decision target vector of the decision scheme, analyze the decision complementarity and conflict between the agents and construct a cooperative relationship graph, calculate the node centrality and edge weight distribution to generate a cooperative contribution factor, and perform nonlinear fusion with the trust index to obtain a dynamic trust weight.

[0138] The peak decision unit is used to identify high-density decision regions and form multiple decision peaks based on the decision target vector of the decision scheme through kernel density estimation, and to perform decision weighting in each decision peak region using the dynamic trust weight to generate peak representative decisions.

[0139] The consensus optimization unit is used to build a fusion optimization model based on the representative decisions of each peak. With the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks, it searches the Pareto optimal solution space by adaptively adjusting the fusion coefficients of the representative decisions of each peak to generate the final consensus decision.

[0140] The execution feedback unit is used to send the final consensus decision to the execution layer and monitor the execution effect, and update the credibility index of each agent in reverse according to the execution effect.

[0141] A third aspect of the present invention provides an electronic device, comprising:

[0142] processor;

[0143] Memory used to store processor-executable instructions;

[0144] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0145] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0146] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent weighted consensus method for multi-agent decision-making, characterized in that, include: Collect decision proposals submitted by each agent and simultaneously obtain historical contribution records and deviation correction records of each agent; Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a trust index; the decision target vector of the decision scheme is extracted, the decision complementarity and conflict between agents are analyzed and a cooperation relationship graph is constructed, the node centrality and edge weight distribution are calculated to generate a cooperation contribution factor, and the dynamic trust weight is obtained by nonlinear fusion with the trust index. Based on the decision target vector of the decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weight to generate peak representative decisions. A fusion optimization model is constructed based on the decisions of each peak representative. The dual optimization objectives are to maximize global consistency and maximize the retention of the representativeness of the decision peak. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of each peak representative decision to generate the final consensus decision. The final consensus decision is sent to the execution layer and the execution effect is monitored. The credibility index of each agent is updated in reverse based on the execution effect.

2. The method according to claim 1, characterized in that, Based on the historical contribution records and deviation correction records, the long-term reliability of each agent is calculated to obtain a reliability index, including: Extract the decision adoption frequency and decision execution success rate of each agent in each historical decision cycle from the historical contribution records, multiply them to obtain the contribution degree and construct a contribution degree sequence, apply a time decay function to the contribution degree sequence for weighting, and calculate the weighted historical contribution degree of each agent; Extract the decision deviation magnitude and deviation correction response speed of each agent from the deviation correction record, construct the deviation behavior feature vector, and calculate the deviation penalty coefficient of each agent by mapping the deviation behavior feature vector to the normalization space. The weighted historical contribution value and the deviation penalty coefficient are combined nonlinearly to generate a preliminary reliability score for each agent. The preliminary reliability score is then globally normalized and a stability constraint transformation is applied to obtain the credibility index for each agent.

3. The method according to claim 1, characterized in that, Extract the decision objective vector of the decision scheme, analyze the complementarity and conflict of decisions among the agents and construct a cooperative relationship graph, calculate node centrality and edge weight distribution to generate a cooperative contribution factor, and perform nonlinear fusion with the credibility index to obtain a dynamic trust weight, including: The decision schemes of each agent are mapped in the feature space to extract the decision target vector and the constraint boundary set. The proportion of the intersection between the projected correlation of the decision target vectors and the constraint boundary set between the agent pairs is calculated. When the projected correlation is positive and the intersection proportion is lower than the complementarity judgment threshold, it is marked as a complementary relationship. When the projected correlation is negative or the intersection proportion is higher than the conflict judgment threshold, it is marked as a conflict relationship. Using each agent as a node and complementary and conflicting relationships as connecting edges, a cooperative relationship graph is constructed. For complementary relationships, the positive edge weight is calculated using the cooperative gain degree of the decision objective vector, and for conflicting relationships, the negative edge weight is calculated using the overlapping conflict intensity of the constraint boundary, forming a signed edge weight distribution. Based on the edge weight distribution, perform multi-hop neighborhood aggregation operation, accumulate the positive edge weights and negative edge weights received by each node, and calculate the node centrality by combining the node in-degree ratio. The collaboration contribution factor is generated by weighted fusion of the variance features of node centrality and edge weight distribution; a nonlinear fusion mapping function is constructed, which takes the collaboration contribution factor and the trust index as two-dimensional inputs and achieves nonlinear coupling through hyperbolic tangent transformation to generate dynamic trust weights.

4. The method according to claim 3, characterized in that, Based on the edge weight distribution, a multi-hop neighborhood aggregation operation is performed, accumulating the positive and negative edge weights received by each node, and calculating the node centrality by combining the node's in-degree ratio, including: Starting from each node, a breadth-first traversal is performed along the directed edges in the collaboration graph to record the shortest hop count to reach neighboring nodes. Based on the shortest hop count, the neighboring nodes are divided into different levels of neighborhood rings to construct a hierarchical neighborhood structure. A weight decay factor based on topological distance is set for each level of neighborhood ring. For each node, traverse the incoming edges it receives, identify the positive and negative edge weights carried by the incoming edges, obtain the corresponding weight decay factor according to the neighborhood ring level to which the starting point of the incoming edge belongs, perform weighted calculations to obtain weighted positive and weighted negative values, accumulate all weighted positive values ​​to obtain the positive aggregation amount, and accumulate all weighted negative values ​​to obtain the negative aggregation amount. The ratio of the number of incoming edges to the number of outgoing edges of each node is calculated to obtain the node's in-degree ratio. The net influence is obtained by subtracting the negative aggregation from the positive aggregation. The net influence is then weighted and merged with the node's in-degree ratio to generate node centrality.

5. The method according to claim 1, characterized in that, Based on the decision objective vector of the aforementioned decision scheme, multiple decision peaks are formed by identifying high-density decision regions through kernel density estimation. Within each decision peak region, decision weighting is performed using the dynamic trust weights to generate peak-representative decisions, including: The decision target vectors of each agent are mapped to a unified decision feature space. The feature distance between the decision target vectors is calculated and a distance matrix is ​​constructed. The kernel function bandwidth parameter is determined based on the distance matrix. Local kernel functions are constructed with each decision target vector as the core and superimposed to generate a global kernel density estimation field. The global kernel density estimation field is spatially gridded and sampled. Density estimates are calculated at grid nodes to form a density scalar field. Morphological watershed segmentation is performed on the density scalar field to identify the boundaries of the density gradient convergence basins. The maximum density point of each convergence basin is extracted as the peak of the decision peak. The decision peak region is extended along the peak of the decision peak to the boundary of the convergence basin. The dynamic trust weights of agents within each decision peak region are extracted. Weighted centroid calculation is performed on the decision target vector within the decision peak region to obtain the centroid decision vector. The spatial variance and distribution skewness of the decision target vector within the decision peak region are calculated to construct a regional stability correction factor. The centroid decision vector is then corrected to obtain the peak representative decision of each decision peak region.

6. The method according to claim 1, characterized in that, A fusion optimization model is constructed based on the representative decisions of each peak, with the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks. The Pareto optimal solution space is searched by adaptively adjusting the fusion coefficients of the representative decisions of each peak, generating the final consensus decision, including: Each peak represents a decision and is constructed into a peak decision matrix. Singular value decomposition is then performed, and the left singular vectors corresponding to the principal singular values ​​are spanned into a globally consistent constraint subspace. Extract the cumulative dynamic trust weight of the agent within the decision peak region corresponding to each peak and the peak height of the kernel density. Multiply the cumulative dynamic trust weight by the peak height of the kernel density and apply a logarithmic transformation to obtain the peak influence measure. The fusion coefficients of each peak representing the decision are initialized with equal weights. The first objective function is constructed as the projection modulus of the fused decision in the global consistency constraint subspace. The second objective function is constructed as the inner product of the fusion coefficients and the peak influence measure. The constraint condition is set that the sum of the fusion coefficients is a unit value and each component is non-negative, thus obtaining the bi-objective optimization problem. A multi-objective evolutionary algorithm is used to solve the bi-objective optimization problem. The population is screened by fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, thus obtaining the Pareto optimal solution set. In the Pareto optimal solution set, the solution with the largest sum of the normalized values ​​of the two objective functions is selected, and the corresponding fusion coefficients are extracted. A convex combination operation is then performed on each peak to represent the decision to generate the final consensus decision.

7. The method according to claim 6, characterized in that, A multi-objective evolutionary algorithm is employed to solve the bi-objective optimization problem. The population is screened using fast non-dominated sorting and crowding distance calculation, and the fusion coefficient configuration is iteratively updated until convergence, yielding a Pareto optimal solution set, including: An initial population is generated, each individual is encoded as a fusion coefficient vector, the first objective function value and the second objective function value of each individual are calculated, a domination counter is established to record the number of times each individual is dominated, individuals with a counter of zero are added to the first non-dominated level, the domination list of individuals in the first non-dominated level is traversed and the counter of dominated individuals is decremented by one, and the process is repeated to form subsequent non-dominated levels. The sum of the absolute values ​​of the differences between the individual and its neighbor in the two objective function values ​​within each non-dominated hierarchy is calculated as the crowding distance, and an infinite crowding distance is assigned to the individuals at the end of the hierarchy. The crowding degree comparison selection operator is used to select parent individuals, and a uniform crossover operation is performed on the parent individuals to generate intermediate individuals and an adaptive mutation operation is performed. The mutated individuals are then processed into units to generate the offspring population. The parent and offspring populations are merged and non-dominated sorting and crowding distance calculation are re-executed. Individuals at each non-dominated level are selected sequentially according to the elite strategy until the preset population size is reached. The last non-dominated level is truncated according to the crowding distance to form a new generation population. Calculate the maximum movement distance of an individual in the first non-dominated level in the target space relative to the previous generation. When the maximum movement distance is lower than a preset threshold for multiple consecutive generations, convergence is determined, and all individuals in the first non-dominated level are output as the Pareto optimal solution set.

8. An intelligent weighted consensus system for multi-agent decision-making, used to implement the method as described in any one of claims 1-7, characterized in that, include: The data acquisition unit is used to collect the decision-making schemes submitted by each agent and simultaneously obtain the historical contribution records and deviation correction records of each agent. The credibility index calculation unit is used to calculate the long-term reliability of each agent based on the historical contribution record and the deviation correction record to obtain the credibility index. The trust weight calculation unit is used to extract the decision target vector of the decision scheme, analyze the decision complementarity and conflict between the agents and construct a cooperative relationship graph, calculate the node centrality and edge weight distribution to generate a cooperative contribution factor, and perform nonlinear fusion with the trust index to obtain a dynamic trust weight. The peak decision unit is used to identify high-density decision regions and form multiple decision peaks based on the decision target vector of the decision scheme through kernel density estimation, and to perform decision weighting in each decision peak region using the dynamic trust weight to generate peak representative decisions. The consensus optimization unit is used to build a fusion optimization model based on the representative decisions of each peak. With the dual optimization objectives of maximizing global consistency and maximizing the retention of representativeness of decision peaks, it searches the Pareto optimal solution space by adaptively adjusting the fusion coefficients of the representative decisions of each peak to generate the final consensus decision. The execution feedback unit is used to send the final consensus decision to the execution layer and monitor the execution effect, and update the credibility index of each agent in reverse according to the execution effect.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.