A large-scale complex group consensus decision-making method, a terminal device, and a medium
A rocket engine model selection method driven by reinforcement learning algorithms and Bayesian hierarchical models solves the problems of information quantification and consensus adjustment in complex group decision-making, and achieves a high-precision and efficient decision-making process.
Patent Information
- Application Number
- CN202511513895.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing technologies struggle to effectively handle complex multi-criteria decisions in rocket engine selection, neglecting uncertainties and resulting in low decision-making accuracy and inefficient consensus adjustments, which impacts rocket performance and mission success.
The algorithm uses reinforcement learning to adjust the probabilistic language evaluation matrix, and combines Bayesian hierarchical model and cloud model to achieve consensus and information quantification among agents. The reward function drives local adjustment, and the attribute weights and agent weights are used for decision optimization.
It improves the accuracy and efficiency of rocket engine model selection decisions, avoids redundant calculations and loss of key information, and enhances the accuracy and speed of large-scale complex group consensus decision-making.
Smart Images

Figure CN120996079B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of large-scale complex group decision-making, and particularly relates to a large-scale complex group consensus decision-making method, a terminal device and a medium. BACKGROUND
[0002] In the field of aerospace engineering, rocket engine type selection is a key decision-making link that determines the performance of a launch vehicle and the success of a mission. With the rapid development of aerospace technology, new types of engines are constantly emerging, and the decision-making environment is becoming increasingly complex. Traditional large-scale complex group consensus decision-making methods have been unable to meet the high-precision and high-reliability decision-making requirements.
[0003] The existing technology mainly has the following outstanding problems:
[0004] On the one hand, rocket engine selection needs to consider more than ten indicators such as thrust performance, specific impulse, reliability, cost, development cycle, and technology maturity. Traditional multi-criteria decision-making methods such as AHP and TOPSIS are difficult to effectively handle such a complex evaluation system. When evaluating new engines, intelligent agents often use probability language to express uncertainty, such as "high specific impulse is likely to be achieved" and "reliability may be affected". However, existing methods lack effective quantitative processing mechanisms. Moreover, the determination of indicator weights is mostly dependent on the subjective judgment of intelligent agents, which is easily influenced by personal preferences, and different intelligent agents have different perceptions of the importance of indicators, making it difficult to reach a consensus.
[0005] On the other hand, when dealing with probability language evaluations, existing technologies usually use expected values or simple weighted averages for quantification, resulting in the loss of a large amount of valuable information. For example, an intelligent agent's reliability evaluation of a certain type of engine "{high (0.6), medium (0.4)}" and "{medium (1.0)}" may result in the same value in traditional methods, but the former contains important uncertainty information. In high-risk decisions such as rocket engine selection, ignoring uncertainty can lead to disastrous consequences, and existing methods cannot meet the stringent requirements of aerospace engineering for decision-making accuracy.
[0006] On the other hand, aerospace project decision-making often involves intelligent agents from multiple professional fields such as aerodynamics, structure, propulsion, and control, with significant differences in opinions and a long consensus reaching process. Traditional consensus adjustment relies on meetings or email communication, which is inefficient, and a single decision often takes several weeks. The adjustment process lacks intelligent guidance, often resulting in "adjusting non-critical opinions" or "over-adjustment leading to intelligent agent resistance", which affects decision-making quality and team collaboration.
[0007] In summary, the existing technology has technical defects such as insufficient information expression, inaccurate quantification, and unintelligent consensus adjustment when dealing with complex aerospace engineering decision-making problems such as rocket engine type selection. These technical defects will affect the accuracy and efficiency of rocket engine type selection. SUMMARY
[0008] The technical problem solved by the present application is to provide a large-scale complex group consensus decision-making method, a terminal device and a medium, and to improve the accuracy of large-scale complex group consensus decision-making.
[0009] In a first aspect, the present application provides a large-scale complex group consensus decision-making method, which comprises the following steps:
[0010] Obtaining a probability linguistic evaluation matrix of a plurality of intelligent agents participating in decision-making on a plurality of candidate schemes in a plurality of scheme attributes; the candidate scheme is a rocket engine model, and the scheme attribute is a rocket engine performance evaluation index;
[0011] Adjusting the probability linguistic evaluation matrix based on a reinforcement learning algorithm to drive the plurality of intelligent agents to reach a consensus and obtain a new probability linguistic evaluation matrix; wherein the reward function of the reinforcement learning algorithm represents the change in the degree of hesitation and the degree of consensus of the intelligent agent before and after the action, and the action is designed to guide the intelligent agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other intelligent agents;
[0012] For the new probability linguistic evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated using a Bayesian hierarchical model; the attribute weight is used to measure the importance of the scheme attribute;
[0013] According to the degree of hesitation and the degree of consensus of each intelligent agent in the new probability linguistic decision matrix, the intelligent agent weight corresponding to the intelligent agent is determined; the intelligent agent weight is used to measure the objectivity of the intelligent agent;
[0014] The new probability linguistic evaluation matrix is converted into a numerical evaluation matrix through a cloud model; the cloud model realizes the faithful conversion of probability language to numerical value based on three parameters of expected value, entropy and hyper-entropy, the expected value is used to measure the evaluation level of the intelligent agent, the entropy is used to measure the size of the divergence of the intelligent agent, and the hyper-entropy is used to measure the evaluation stability of the intelligent agent;
[0015] The numerical evaluation matrix, the attribute weight and the intelligent agent weight are fused to obtain a scheme evaluation cloud model; the scheme evaluation cloud model is used to model the fuzziness and randomness in the evaluation process of the large-scale complex group in an integrated manner;
[0016] According to the scheme evaluation cloud model, the final scheme reaching the consensus of the large-scale complex group is determined from the plurality of candidate schemes.
[0017] Optionally, the expression of the reward function is ; wherein, represents the reward value corresponding to the action, represents the change in the degree of hesitation, represents the change in the degree of consensus, , represents the hesitation degree of the i-th candidate solution on the j-th solution attribute evaluated by the k-th agent at the t-th time step, represents the total number of candidate solutions, represents the total number of solution attributes, represents the total number of agents, represents the total number of time steps, represents the hesitation degree of the i-th candidate solution on the j-th solution attribute evaluated by the k-th agent at the t-th time step, , represents the maximum value of the probability of the i-th candidate solution on the j-th solution attribute evaluated by the k-th agent in the language terms, represents the cardinality of represents the set of probability language evaluations of the i-th candidate solution on the j-th solution attribute evaluated by the k-th agent, represents the consensus degree of the i-th candidate solution on the j-th solution attribute evaluated by the k-th agent with other agents at the t-th time step, The consensus degree calculation formula between the i-th agent and the j-th agent on the k-th scheme attribute is: , represents the probability linguistic evaluation vector of the i-th agent on the k-th scheme attribute of the j-th candidate scheme, represents the probability linguistic evaluation vector of the i-th agent on the k-th scheme attribute of the j-th candidate scheme, represents the Euclidean distance between the i-th agent and the j-th agent on the k-th scheme attribute of the j-th candidate scheme,
[0018] The value update rule of the reinforcement learning algorithm is: , wherein, represents the learning rate, represents the reward signal, which is a measure of consensus degree improvement or error reduction, represents the discount factor, represents the current state, which is used to determine the importance of future rewards relative to immediate rewards, represents the next state
[0019] The expression of the action of the reinforcement learning algorithm is: , represents the control parameter, , The smaller the value of is, the more the action adjustment tends to preserve the original information,
[0020] Optionally, for the new probability linguistic evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated using a Bayesian hierarchical model, including:
[0021] Based on the evaluation matrix of the new probabilistic language, the evaluation bias of the agent on the same solution attributes when evaluating different candidate solutions is calculated; the evaluation bias is used to measure the degree of importance the agent attaches to the solution attributes.
[0022] Based on the evaluation bias, the optimal and worst attributes are determined from multiple alternative attributes, and the advantage comparison matrix and the gap comparison matrix are calculated based on the probability degree algorithm. The advantage comparison matrix is used to measure the degree of preference of the optimal attribute relative to other alternative attributes, and the gap comparison matrix represents the degree of preference of other alternative attributes relative to the worst attribute.
[0023] By modeling the advantage comparison matrix and the gap comparison matrix using multinomial distribution, we obtain the first multinomial distribution expression and the second multinomial distribution expression;
[0024] For the first and second polynomial distribution expressions, a Bayesian hierarchical model is constructed, and the Bayesian hierarchical model is solved using the Markov chain Monte Carlo algorithm to obtain the attribute weights.
[0025] Optionally, the expression for evaluation bias is:
[0026]
[0027] in, Indicates the first The agent's evaluation in the 1st... Evaluation bias in the attributes of each solution Indicates the first The evaluation of the first agent The candidate solutions and the first When the candidate solution is at the th time Each scheme attribute Evaluation bias on and , Indicates the first The evaluation of the first agent The candidate solution is in the... The standardized probabilistic linguistic evaluation vectors for each scheme attribute are typically normalized to eliminate scale differences, based on the evaluation set of all agents. , Indicates the first The evaluation of the first agent The candidate solution is in the... Standardized probabilistic language evaluation vectors for each scheme attribute;
[0028] The expression for the optimal attribute is: ;
[0029] The expression for the worst-case property is: ;
[0030] The expression for the first polynomial distribution is: ;
[0031] The expression for the second polynomial distribution is: ;
[0032] For the new probabilistic language evaluation matrix, the attribute weights corresponding to each of the proposed solutions are calculated using a Bayesian hierarchical model, including:
[0033] Calculate the multinomial probability distribution of the worst-case attribute to characterize the other attributes relative to the worst-case attribute. The probability of the preference relationship; the expression for the probability of the preference relationship is: ,in, This represents a matrix comparing the preferences of other attributes for the worst-case attribute. , This represents the attribute weight vector to be determined. , , , For the elements of the preference matrix, reflecting the first Each solution attribute is compared with the worst attribute. The degree of preference, Denotes the polynomial coefficients, ensuring that the sum of probabilities equals 1;
[0034] Calculate the multinomial probability distribution of the optimal attribute to characterize the optimal attribute. The probability of preference relationships relative to other attributes is expressed as follows:
[0035]
[0036] in, This represents a matrix comparing the preferences of the optimal attribute for other attributes. , The elements of the preference matrix reflect the optimal attributes. For the The degree of preference for each option attribute;
[0037] To characterize the joint probability of the Bayesian hierarchical model, we fuse the preferences of multiple agents and construct the joint probability distribution of the hierarchical model, which connects input preferences and output weights. Its expression is:
[0038]
[0039] in, Indicates the overall attribute weight The prior distribution is determined by using the non-informative Dirichlet distribution to avoid subjective bias, and satisfies... , Dirichlet distribution with all parameters being 1, which is equivalent to a uniform distribution, the first attribute weight of the expert the conditional distribution of the integrated weight follows Dirichlet distribution , denotes the concentration parameter, for controlling the closeness to , the larger the value is, the more the expert weight concentrates on , needs to be modeled by a gamma distribution , , is the shape parameter, , are the optimal and the worst attribute preference matrices of the first expert respectively , the conditional distribution of its own weight ;
[0040] the posterior distribution is solved by Markov Chain Monte Carlo technique, and finally the attribute weight is output; wherein the integrated attribute weight of all experts is obtained by aggregating the posterior distribution of each expert weight, which is the final attribute weight used for decision-making.
[0041] Optionally, the expression of the agent weight is:
[0042]
[0043]
[0044]
[0045] wherein, denotes the agent weight corresponding to the first agent, denotes the hesitation degree of the first agent, denotes the consensus degree of the first agent, which is 0 when there is only one probability language term, and is 1 when there are language variables with equal possibility, and .
[0046] Optionally, the expression of the cloud model is:
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053] in, Indicates the first The evaluation of the first agent The candidate solution is in the... The cloud model transformation result for each scheme attribute is typically a triple (Ex, En, He), used to quantify the uncertainty of the evaluation. Expressing expectations The square of, Represents the square of the probability. The average value of entropy. This represents the average value of hyperentropy. This represents the expected value, which is the first rank of the agent's evaluation. The candidate solution is in the... Evaluation center values for each scheme attribute, The domain that represents the cloud model.
[0054] Optionally, the expression for the cloud model for evaluating the solution is: .
[0055] Optionally, based on the scheme evaluation cloud model, the final scheme for achieving consensus in large-scale complex groups is determined from multiple candidate schemes, including:
[0056] According to the cloud model Determine the positive ideal solution from the multiple solution attributes. and negative ideal solution ;in, , , , Take the absolutely ideal solution, which represents the state values of the scheme's attribute elements when they are theoretically at their best and worst.
[0057] Based on shape similarity and distance similarity, the similarity between each candidate solution and the positive ideal solution and the negative ideal solution is calculated; wherein, the expression for the solution similarity is: or , ;
[0058] Through calculation formula Candidate solutions were obtained. The closeness between the positive ideal solution and the negative ideal solution and the negative ideal solution The closeness reflects the degree to which the candidate solution is similar to the positive ideal solution and dissimilar to the negative ideal solution, and the smaller the closeness, the better the performance of the candidate solution;
[0059] The candidate solution with the largest closeness is determined as the final solution.
[0060] In a second aspect, the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above method when executing the computer program.
[0061] In a third aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the above method.
[0062] The present application has at least the following beneficial effects:
[0063] The reward function of the reinforcement learning algorithm is used to represent the change in the hesitation degree and the change in the consensus degree of the agent before and after performing the action, the action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents, and the local adjustment instead of global iteration is driven by the reward function, thereby avoiding redundant calculation, reducing the time consumption of the complex decision consensus process, improving the consensus reaching efficiency, and thus being beneficial to improving the precision of the large-scale complex group consensus decision; the cloud model realizes the faithful conversion of the probability language to the value based on the three parameters of the expected value, the entropy and the hyper entropy, realizes the synchronous digitization of the fuzziness and randomness in the probability language, avoids the loss of key information, and improves the precision of the large-scale complex group consensus decision. BRIEF DESCRIPTION OF DRAWINGS
[0064] The accompanying drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification, and are used to explain the technical solutions of the present application together with the embodiments of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0065] Figure 1 It is a flowchart of the large-scale complex group consensus decision method in one of the embodiments of the present application;
[0066] Figure 2 It is a schematic diagram of the change in the consensus degree and the hesitation degree of the agent consensus reaching in one of the embodiments of the present application;
[0067] Figure 3 It is a structural diagram of the terminal device in one of the embodiments of the present application. DETAILED DESCRIPTION
[0068] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0069] In view of the technical problem of low accuracy of a traditional large-scale complex group consensus decision method, the present application provides a large-scale complex group consensus decision method, a terminal device and a medium. A reward function of a reinforcement learning algorithm used by the method is used to represent the change in hesitation degree and consensus degree of an agent before and after an action is performed. The action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents. The reward function drives local adjustment instead of global iteration, avoids redundant calculation, reduces the time consumption of the complex decision consensus process, and improves the consensus efficiency, thereby helping to improve the accuracy of the large-scale complex group consensus decision.
[0070] Embodiment one
[0071] As shown in Figure 1 The large-scale complex group consensus decision method provided by the present application includes the following steps:
[0072] Step 11, obtaining a probability language evaluation matrix of a plurality of agents participating in decision-making on a plurality of candidate schemes in a plurality of scheme attributes.
[0073] In the embodiments of the present application, the candidate scheme is a rocket engine model. For example, the case of multiple candidate schemes is shown in Table 1.
[0074] Table 1
[0075]
[0076] The scheme attribute is a performance evaluation index of the rocket engine. In one possible implementation, the scheme attribute includes specific impulse, thrust, thrust-to-weight ratio, service life and stability.
[0077] Specifically, in the embodiments of the present application, the agent represents a decision-making subject participating in decision-making, which can be a human expert or an artificial intelligence system.
[0078] In one possible implementation, Likert scale 5-point scale is used to depict the evaluation information of the agent on the candidate scheme. The Likert scale 5-point scale is represented as: ={very disagree, disagree, hard to judge, agree, very agree} The evaluation value can be expressed as .
[0079] Step 12, adjusting the probabilistic language evaluation matrix based on the reinforcement learning algorithm, driving multiple agents to reach a consensus, and obtaining a new probabilistic language evaluation matrix.
[0080] In the embodiment of the application, the reward function of the reinforcement learning algorithm represents the change of the hesitation degree and the change of the consensus degree of the agent before and after the action is performed, and the action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents.
[0081] The process of adjusting the probabilistic language evaluation matrix based on the reinforcement learning algorithm, driving multiple agents to reach a consensus, and obtaining a new probabilistic language evaluation matrix will be described below, which specifically includes steps 12.1 to 12.5:
[0082] Step 12.1, initializing the state space, action space, learning rate, exploration rate and discount factor of the reinforcement learning algorithm.
[0083] Specifically, the state of the reinforcement learning algorithm is the average consensus degree of all agents in the current decision-making process. The action can be expressed as the score of all agents on a certain candidate scheme on a certain scheme attribute being changed to the score of another agent on the scheme on the attribute, which is expressed as , wherein is the learning object, is the adjusted attribute index, is the adjusted scheme index.
[0084] Step 12.2, determining the learning object from the agent with a consensus degree less than the average consensus degree of the large-scale complex group, and selecting the action with the maximum Q value in the current state.
[0085] Specifically, the expression of the adjustment action is: , wherein, represents the updated evaluation of the first agent on the first candidate scheme on the first scheme attribute at time step , , , , represents a control parameter, , The smaller the control parameter is, the more the action adjustment tends to retain the original information, represents all agents' evaluation of the first candidate scheme on the first scheme attribute, , The group average of the evaluation of the j-th solution attribute by the agents, calculated as a weighted average or an arithmetic average of the probability distribution, UCA denotes the agent, candidate solution, and solution attribute triplets.
[0086] Step 12.3, calculate the updated large-scale complex group average consensus degree, and calculate the immediate reward.
[0087] The expression of the reward function is ; wherein, represents the reward value corresponding to the action, represents the change in hesitation degree, represents the change in consensus degree, , represents the evaluation of the j-th agent on the i-th candidate solution on the i-th solution attribute at time step , , represents the total number of candidate solutions, , represents the total number of solution attributes, , represents the total number of agents, , represents the total time step, represents the evaluation of the j-th agent on the i-th candidate solution on the i-th solution attribute at time step , The calculation formula of the evaluation of the j-th agent on the i-th candidate solution on the i-th solution attribute is: , represents the maximum value of the probability in the linguistic term evaluated by the j-th agent on the i-th candidate solution on the i-th solution attribute, represents the minimum value of the probability in the linguistic term evaluated by the j-th agent on the i-th candidate solution on the i-th solution attribute, represents the cardinality of , represents the set of probability linguistic evaluations of the j-th agent on the i-th candidate solution on the i-th solution attribute, , represents the set of probability linguistic evaluations of the j-th agent on the i-th candidate solution on the i-th solution attribute, , represents the total number of , represents the total number of , , represents the set of probability linguistic evaluations of the j-th agent on the i-th candidate solution on the i-th solution attribute at time step the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step, the consensus degree between the agent and other agents on the mth scheme attribute of the nth candidate scheme at the tth time step,
[0088] Step 12.4, calculate the maximum Q value of the next state and update the Q value of the current state.
[0089] Specifically, the value update rule is: wherein, denotes the learning rate, denotes the reward signal, which is a measure of consensus degree improvement or error reduction, denotes the discount factor, denotes the current state, which is used to determine the importance of future rewards relative to immediate rewards, denotes the next state the maximum value of all possible actions in the next state.
[0090] Step 12.5, if the current state meets the preset termination condition, output the new probability language evaluation matrix; otherwise, repeatedly execute steps 12.3 to 12.4 until the current state meets the preset termination condition.
[0091] In one feasible implementation, the preset termination condition is that the number of updates is greater than or equal to a preset threshold.
[0092] In this embodiment of the invention, as the reinforcement learning algorithm continues to iterate, the changes in the consensus degree and hesitation degree of the agent consensus are as follows: Figure 2 As shown, Figure 2 The horizontal axis represents the number of iterations, the green line represents the degree of consensus, and the red line represents the degree of hesitation. In the early stages of iteration, the degree of consensus is low due to potentially significant disagreements among experts. Figure 2 It can be seen that as the reinforcement learning algorithm is adjusted, the agents' scores gradually converge, and the consensus increases accordingly. Furthermore, the consensus tends to stabilize with increasing iterations. Therefore, we can conclude that:
[0093] 1. The final average consensus score reached 0.921, indicating that the agent was able to reach a consensus through the iterative process of the reinforcement learning algorithm. Meanwhile, although the hesitation score fluctuated significantly, it generally showed a downward trend, with the average hesitation score decreasing from approximately 0.42 to approximately 0.41, indicating that the agent's hesitation score also decreased through the iterative process of the reinforcement learning algorithm.
[0094] 2. The consensus level exceeding the threshold was reached around the 160th iteration, indicating that the reinforcement learning algorithm provided by this invention can effectively promote consensus formation within a reasonable number of iterations. This result effectively verifies the effectiveness of adjusting the probabilistic language evaluation matrix based on the reinforcement learning algorithm to achieve consensus in large-scale complex groups.
[0095] Step 13: For the new probabilistic language evaluation matrix, use the Bayesian hierarchical model to calculate the attribute weights corresponding to each scheme attribute.
[0096] In this embodiment of the invention, attribute weights are used to measure the importance of the scheme attributes.
[0097] Step 14: Determine the agent weight corresponding to each agent based on the hesitation degree and consensus degree of each agent in the new probabilistic language decision matrix.
[0098] In this embodiment of the invention, agent weights are used to measure the objectivity of the agent.
[0099] Specifically, the expression for agent weights is:
[0100]
[0101]
[0102]
[0103] in, Indicates the first The agent weights corresponding to each agent. Indicates the first Hesitation level of an agent Indicates the first The consensus degree of an agent is 0 when there is exactly one probabilistic language term. If there are 1 linguistic variable and they have equal probability, then the degree of consensus is 1. .
[0104] Step 15: Convert the new probabilistic language evaluation matrix into a numerical evaluation matrix using the cloud model.
[0105] In this embodiment of the invention, the cloud model achieves a fidelity conversion from probabilistic language to numerical values based on three parameters: expected value, entropy, and hyperentropy. The expected value is used to measure the evaluation level of the agent, the entropy is used to measure the degree of disagreement of the agent, and the hyperentropy is used to measure the evaluation stability of the agent.
[0106] Specifically, the formalized expression of the cloud model is as follows: , representing the expected value, entropy, and hyperentropy, respectively. In one feasible implementation, the cloud model is expressed as:
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113] in, Indicates the first The evaluation of the first agent The candidate solution is in the... The cloud model transformation result for each scheme attribute is typically a triple (Ex, En, He), used to quantify the uncertainty of the evaluation. Expressing expectations The square of, Represents the square of the probability. The average value of entropy. This represents the average value of hyperentropy. This represents the expected value, which is the first value evaluated by the agent. The candidate solution is in the... Evaluation center values for each scheme attribute, This represents the domain of the cloud model.
[0114] Step 16, the numerical evaluation matrix, attribute weight and agent weight are fused to obtain a scheme evaluation cloud model.
[0115] In the embodiment of the application, the comprehensive evaluation cloud model is used for.
[0116] In an available implementation, the expression of the scheme evaluation cloud model is:
[0117] .
[0118] Step 17, according to the scheme evaluation cloud model, a final scheme for reaching a large-scale complex group consensus is determined from a plurality of candidate schemes.
[0119] Specifically, steps 17.1 to 17.4 are included:
[0120] Step 17.1, according to the cloud model , a positive ideal scheme is determined from a plurality of scheme attributes and a negative ideal scheme .
[0121] Wherein, , , , The absolute ideal solution is taken, which represents the state value when the scheme attribute element is theoretically optimal and worst.
[0122] Step 17.2, based on shape similarity and distance similarity, the scheme similarity between each candidate scheme and the positive ideal scheme and the negative ideal scheme is calculated.
[0123] Wherein, the expression of the scheme similarity is or ,
[0124] ,
[0125] ,
[0126] ;
[0127] Step 17.3, by calculating formula , the closeness of the candidate scheme to the positive ideal scheme and the negative ideal scheme is obtained.
[0128] The closeness reflects the degree to which the candidate scheme is similar to the positive ideal scheme and dissimilar to the negative ideal scheme, and the smaller the closeness, the better the performance of the candidate scheme.
[0129] Step 17.4: Select the candidate solution with the highest similarity as the final solution.
[0130] Example 2
[0131] This embodiment illustrates the process of calculating the attribute weights corresponding to each scheme attribute using a Bayesian hierarchical model for the new probabilistic language evaluation matrix in Embodiment 1.
[0132] Specifically, for the new probabilistic language evaluation matrix, the attribute weights corresponding to each solution attribute are calculated using a Bayesian hierarchical model, including:
[0133] Step 1: Based on the new probabilistic language evaluation matrix, calculate the evaluation bias of the agent on the same scheme attributes when evaluating different candidate schemes.
[0134] In this embodiment of the invention, the evaluation bias is used to measure the degree of importance that the agent attaches to the scheme attributes.
[0135] Specifically, the expression for evaluation bias is:
[0136]
[0137] in, Indicates the first The agent's evaluation in the 1st... Evaluation bias in the attributes of each solution Indicates the first The evaluation of the first agent The candidate solutions and the first When the candidate solution is at the th time Each scheme attribute Evaluation bias on and , Indicates the first The evaluation of the first agent The candidate solution is in the... The standardized probabilistic linguistic evaluation vectors for each scheme attribute are typically normalized to eliminate scale differences, based on the evaluation set of all agents. , Indicates the first The evaluation of the first agent The candidate solution is in the... Standardized probabilistic language evaluation vectors for each scheme's attributes.
[0138] Step II: Based on the evaluation bias, determine the optimal and worst attributes from multiple alternative attributes, and calculate the advantage comparison matrix and the gap comparison matrix respectively based on the probability degree algorithm.
[0139] In the embodiment of the present application, the advantage comparison matrix is used to measure the preference degree of the optimal attribute relative to other scheme attributes, and the gap comparison matrix represents the preference degree of other scheme attributes relative to the worst attribute.
[0140] Specifically, the expression of the optimal attribute is: ;
[0141] The expression of the worst attribute is: .
[0142] Step III, the advantage comparison matrix and the gap comparison matrix are modeled by using a polynomial distribution, to obtain a first polynomial distribution expression and a second polynomial distribution expression.
[0143] Specifically, the expression of the first polynomial distribution expression is: ;
[0144] The expression of the second polynomial distribution expression is: .
[0145] Step IV, a Bayesian hierarchical model is constructed for the first polynomial distribution expression and the second polynomial distribution expression, and the Bayesian hierarchical model is solved by using a Markov chain Monte Carlo algorithm to obtain the attribute weight.
[0146] The process of calculating the attribute weight corresponding to each scheme attribute by using the Bayesian hierarchical model for the new probability language evaluation matrix includes:
[0147] The polynomial probability distribution of the worst attribute is calculated to depict the preference relationship probability of other attributes relative to the worst attribute , and the expression of the preference relationship probability is , wherein represents the advantage comparison matrix of other attributes relative to the worst attribute, , represents the attribute weight vector to be solved, , , , is a preference matrix element, reflecting the preference degree of the th scheme attribute relative to the worst attribute , and represents a polynomial coefficient, ensuring that the probability sum is 1.
[0148] The polynomial probability distribution of the optimal attribute is calculated to depict the preference relationship probability of the optimal attribute relative to other attributes, and the expression of the preference relationship probability is:
[0149]
[0150] wherein denote the preference comparison matrix of the optimal attribute to other attributes, , is the element of the preference matrix, reflecting the preference degree of the optimal attribute to the attribute of the first scheme;
[0151] The Bayesian hierarchical model joint probability is depicted, the joint probability distribution of the hierarchical model is constructed by fusing the multi-agent preferences, and is used for connecting the input preferences and the output weights, and the expression is:
[0152]
[0153] wherein, denotes the prior distribution of the comprehensive attribute weight , an uninformative Dirichlet distribution is adopted to avoid subjective bias, and satisfies is a Dirichlet distribution, and when the parameters are all 1, it is equivalent to a uniform distribution, is the attribute weight of the first expert, and the conditional distribution of the comprehensive weight is subject to a Dirichlet distribution denotes a concentration parameter, , used for controlling the closeness of , the greater , the more the expert weight is concentrated on needs to be modeled through a gamma distribution is a shape parameter, , are the optimal attribute preference matrix and the worst attribute preference matrix of the first expert , respectively, and the conditional distribution of its own weight ;
[0154] The posterior distribution is solved through Markov chain Monte Carlo technology iteration, and finally the attribute weight is output; wherein the comprehensive attribute weight of all experts is obtained by aggregating the posterior distribution of each expert weight, and is the attribute weight finally used for decision-making.
[0155] Embodiment Three
[0156] This embodiment is an explanation and description of determining the final scheme reaching the consensus of a large-scale complex group from multiple candidate schemes according to the scheme evaluation cloud model in Embodiment One.
[0157] Specifically, according to the scheme evaluation cloud model, a final scheme for achieving consensus of a large-scale complex group is determined from a plurality of candidate schemes, comprising:
[0158] Step A, according to the cloud model , a positive ideal scheme is determined from a plurality of scheme attributes , and a negative ideal scheme .
[0159] Wherein, , , , Take the absolute ideal solution, which represents the state value of the scheme attribute element when it is theoretically optimal and worst;
[0160] Step B, based on shape similarity and distance similarity, the scheme similarity between each candidate scheme and the positive ideal scheme and the negative ideal scheme is calculated.
[0161] Wherein, the expression of the scheme similarity is or , .
[0162] Step C, the formula is calculated by the formula , the closeness of the candidate scheme to the positive ideal scheme and the negative ideal scheme is obtained , the closeness reflects the degree to which the candidate scheme is similar to the positive ideal scheme and dissimilar to the negative ideal scheme, and the greater the closeness, the better the performance of the candidate scheme.
[0163] Step D, the candidate scheme with the maximum closeness is determined as the final scheme.
[0164] The large-scale complex group consensus decision-making method provided by the application has at least the following advantages:
[0165] The reward function of the reinforcement learning algorithm is used to represent the change in the agent's hesitation and consensus before and after the action, and the action is designed to guide the agent whose consensus is less than the average consensus of the large-scale complex group to learn from other agents. Through the reward function, local adjustment is driven instead of global iteration, avoiding redundant calculation, reducing the time-consuming of complex decision consensus process, and improving the consensus efficiency, thereby improving the accuracy of large-scale complex group consensus decision-making; The cloud model realizes the faithful conversion of probability language to numerical value based on three parameters of expectation, entropy and hyperentropy, realizes the simultaneous digitization of fuzziness and randomness in probability language, avoids the loss of key information, and improves the accuracy of large-scale complex group consensus decision-making.
[0166] AsFigure 3 As shown, the embodiment of the present application provides a terminal device, such as Figure 3 As shown, the terminal device D10 of the embodiment comprises at least one processor D100 (only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps in any of the method embodiments described above when executing the computer program D102. Figure 3
[0167] Specifically, when the processor D100 executes the computer program D102, the probability linguistic evaluation matrix of a plurality of candidate schemes on a plurality of scheme attributes is obtained by a plurality of agents participating in decision-making; the probability linguistic evaluation matrix is adjusted based on a reinforcement learning algorithm to drive the plurality of agents to reach a consensus and obtain a new probability linguistic evaluation matrix; for the new probability linguistic evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated by using a Bayesian hierarchical model; the attribute weight is used to measure the importance of the scheme attribute; according to the hesitation degree and the consensus degree of each agent in the new probability linguistic decision matrix, the agent weight corresponding to the agent is determined; the new probability linguistic evaluation matrix is converted into a numerical evaluation matrix by using a cloud model; the numerical evaluation matrix, the attribute weight, and the agent weight are fused to obtain a scheme evaluation cloud model; and the final scheme reaching the large-scale complex group consensus is determined from the plurality of candidate schemes according to the scheme evaluation cloud model. Wherein, the reward function of the reinforcement learning algorithm is used to represent the change of the hesitation degree and the consensus degree of the agent before and after the action is performed, and the action is designed to guide the agent whose consensus degree is less than the average consensus degree of the large-scale complex group to learn from other agents, and the local adjustment instead of global iteration is driven by the reward function, which avoids redundant calculation, reduces the time consumption of the complex decision consensus process, improves the consensus reaching efficiency, and thus is conducive to improving the precision of the large-scale complex group consensus decision; the cloud model realizes the faithful conversion of the probability language to the value based on the three parameters of the expected value, the entropy, and the hyper entropy, realizes the synchronous digitization of the fuzziness and randomness in the probability language, avoids the loss of key information, and improves the precision of the large-scale complex group consensus decision.
[0168] The processor D100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0169] The memory D101 can be an internal storage unit of the terminal device D10 in some embodiments, such as a hard disk or a memory of the terminal device D10. The memory D101 can also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory D101 can include both an internal storage unit and an external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader, data, and other programs, such as program codes of the computer program, etc. The memory D101 can also be used to temporarily store data that has been output or will be output.
[0170] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above various method embodiments.
[0171] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is caused to implement the steps in the above various method embodiments.
[0172] Those skilled in the art should understand: the discussion of the above any embodiment is only exemplary, and is not intended to imply that the protection scope of the present application is limited to these examples; under the idea of the present application, the above embodiments or technical features in different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of one or more embodiments of the present application as described above. In order to be brief, they are not provided in details.
[0173] One or more embodiments of the present application are intended to cover all such alternatives, modifications, and variations as fall within the broad scope of the present application. Accordingly, any one or more embodiments of the present application are intended to encompass all such alternatives, modifications, and variations as fall within the broad scope of the present application.
Claims
1. A method for large-scale complex group consensus decision-making, characterized in that, The method comprises the following steps: obtaining a plurality of intelligent agents participating in decision-making to obtain a plurality of candidate schemes on a plurality of scheme attributes; the candidate scheme is a rocket engine model, and the scheme attribute is a rocket engine performance evaluation index; adjusting the probability language evaluation matrix based on a reinforcement learning algorithm to drive the plurality of intelligent agents to reach a consensus and obtain a new probability language evaluation matrix; for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model; the attribute weight is used to measure the importance of the scheme attribute; determining an intelligent agent weight corresponding to each intelligent agent in the new probability language decision matrix according to the hesitation degree and the consensus degree of the intelligent agent; the intelligent agent weight is used to measure the objectivity of the intelligent agent; converting the new probability language evaluation matrix into a numerical evaluation matrix through a cloud model; the cloud model realizes the faithful conversion of probability language to numerical value based on three parameters of expected value, entropy and hyper entropy; fusing the numerical evaluation matrix, the attribute weight and the intelligent agent weight to obtain a scheme evaluation cloud model; the scheme evaluation cloud model is used for integrated modeling of fuzziness and randomness in the large-scale complex group evaluation process; determining a final scheme reaching the large-scale complex group consensus from the plurality of candidate schemes according to the scheme evaluation cloud model.
2. The method of claim 1, wherein, The reward function of the reinforcement learning algorithm represents the change of the hesitation degree and the consensus degree of the intelligent agent before and after the action is performed, and the action is designed to guide the intelligent agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other intelligent agents; The expression for the reward function is: ;in, This represents the reward value corresponding to the action. This indicates the change in hesitation. This indicates the change in the degree of consensus. , Indicates at time step Time The evaluation of the first agent The candidate solution is in the... Hesitation level in the attributes of each option , This represents the total number of candidate solutions. , Indicates the total number of scheme attributes. , This represents the total number of intelligent agents. , Indicates the total time steps. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of hesitation in the attribute of the first option, the The evaluation of the first agent The candidate solution is in the... The formula for calculating the degree of hesitation on each option attribute is: , Indicates the first The evaluation of the first agent The candidate solution is in the... The maximum probability among the language terms of each scheme attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... The minimum probability among the language terms of each scheme attribute. express The base number, Indicates the first The evaluation of the first agent The candidate solution is in the... A set of probabilistic linguistic evaluations of the attributes of each solution. , Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the proposed solution and other intelligent agents in terms of its attributes. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the first scheme and other intelligent agents in terms of attributes, the degree of consensus between the second scheme and other intelligent agents. The evaluation of the first agent The candidate solution is in the... The formula for calculating the consensus degree between a scheme and other intelligent agents based on its attributes is as follows: , Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... The attributes of the first scheme are the same as the second. Euclidean distance between agents; the reinforcement learning algorithm The value update rule is: wherein, denotes a learning rate, denotes a reward signal, a measure of consensus improvement or error reduction, denotes a discount factor, denotes a current state, used to determine the importance of future rewards relative to immediate rewards, denotes a next state the maximum value over all possible actions ; The action of the reinforcement learning algorithm The expression is: ,in, Indicates at time step Time The evaluation of the first agent The candidate solution is in the... Evaluation of the updated attributes of each solution Indicates control parameters, , The smaller the value, the more the motion adjustment tends to retain the original information. This indicates the evaluation of all agents. The candidate solution is in the... The group average of the evaluation of j on each alternative attribute is calculated as a weighted average or arithmetic average of the probability distribution. UCA represents the triplet of agent, candidate alternative, and alternative attribute.
3. The method of claim 2, wherein, for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model, comprising: calculating the evaluation deviation of the intelligent agent when evaluating different candidate schemes on the same scheme attribute according to the new probability language evaluation matrix; the evaluation deviation is used to measure the importance of the intelligent agent to the scheme attribute; determining an optimal attribute and a worst attribute from the plurality of scheme attributes according to the evaluation deviation, and calculating an advantage comparison matrix and a gap comparison matrix based on a possibility algorithm; the advantage comparison matrix is used to measure the preference degree of the optimal attribute relative to other scheme attributes, and the gap comparison matrix represents the preference degree of other scheme attributes relative to the worst attribute; modeling the advantage comparison matrix and the gap comparison matrix by using a polynomial distribution to obtain a first polynomial distribution expression and a second polynomial distribution expression; constructing a Bayesian hierarchical model for the first polynomial distribution expression and the second polynomial distribution expression, and solving the Bayesian hierarchical model by using a Markov chain Monte Carlo algorithm to obtain the attribute weight.
4. The method of claim 3, wherein, The expression of the evaluation deviation is: wherein, represents the evaluation deviation of the i-th agent on the j-th scheme attribute, represents the evaluation deviation of the i-th agent on the j-th scheme attribute, and , represents the normalized probability linguistic evaluation vector of the i-th agent on the j-th scheme attribute, which is normalized to eliminate the scale difference, based on the evaluation set of all agents represents the normalized probability linguistic evaluation vector of the i-th agent on the j-th scheme attribute. The expression of the optimal attribute is: ; The expression of the worst attribute is: ; The expression expressed by the first polynomial distribution is: ; The expression expressed by the second polynomial distribution expression is: ; for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model, comprising: a polynomial probability distribution of the worst attribute is calculated, which characterizes the preference relationship probabilities of other attributes with respect to the worst attribute ; the expression of the preference relationship probability is , where represents a preference comparison matrix of other attributes with respect to the worst attribute, , represents an attribute weight vector to be solved, , , , is a preference matrix element, reflecting the preference degree of the first attribute of the scheme with respect to the worst attribute , represents a polynomial coefficient, which ensures that the probability sum is 1; A polynomial probability distribution of the optimal attribute is calculated, characterizing the optimal attribute The expression of the preference relation probability with respect to other attributes is: wherein, represents the preference comparison matrix of the optimal attribute to other attributes, , is the preference matrix element reflecting the preference degree of the optimal attribute to the first attribute of the scheme. characterizing the joint probability of the Bayesian hierarchical model, fusing the preferences of multiple intelligent agents, constructing the joint probability distribution of the hierarchical model, and using the joint probability distribution to link the input preference and the output weight, and the expression is: where, represents the integrated attribute weight is the prior distribution of the integrated attribute weight, which is modeled as a non-informative Dirichlet distribution to avoid subjective bias, satisfying , is the Dirichlet distribution, which is equivalent to the uniform distribution when all parameters are 1, is the attribute weight of the th expert is the conditional distribution of the integrated weight, which is modeled as a Dirichlet distribution , represents the concentration parameter, , which is used to control the closeness of to , The larger the , needs to be modeled by a gamma distribution , a, b are shape parameters, , are the best and worst attribute preference matrices of the th expert , is the conditional distribution of its own weight ; The attribute weight is finally output by iteratively solving the posterior distribution through Markov chain Monte Carlo technology; wherein the comprehensive attribute weight of all experts The attribute weight used for decision-making finally is obtained by aggregating the posterior distribution of each expert weight.
5. The method of claim 4, wherein, the expression of the intelligent agent weight is: wherein, represents the weight of the agent corresponding to the agent, represents the hesitancy of the agent, represents the consensus of the agent, the consensus is 0 when there is only one probability language term, and the consensus is 1 when there are language variables with equal probability, and at this time . 6. The method of claim 5, wherein, An expression of the cloud model is: wherein, represents the nth intelligent agent evaluation of the mth candidate solution on the nth scheme attribute, is a three-tuple (Ex, En, He), which is used to quantify the uncertainty of evaluation, represents the square of expectation, represents the domain of the cloud model. 7. The method of claim 6, wherein, The scheme evaluates the expression of the cloud model as: .
8. The method of claim 7, wherein The method for evaluating the cloud model according to the scheme to determine a final scheme for reaching a large-scale complex group consensus from the multiple candidate schemes includes: According to the cloud model , determining a positive ideal scheme from the multiple scheme attributes , and a negative ideal scheme ; wherein , , , Taking absolute ideal solution, representing the state value of the scheme attribute element when it is theoretically optimal and worst. Based on the shape similarity and the distance similarity, a scheme similarity between each of the candidate schemes and the positive ideal scheme and the negative ideal scheme is calculated; wherein, an expression of the scheme similarity is or, , ; The closeness between the candidate scheme and the positive ideal scheme and the negative ideal scheme is obtained by the calculation formula ; The candidate scheme with the maximum closeness is determined as the final scheme.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor implements the method of any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Simulation model credibility evaluation method based on hesitant cloud language term set and group decision
CN110717281A
Intelligent manufacturing capability maturity evaluation method, device and equipment
CN118586737A