Large-scale complex group consensus decision-making method, terminal equipment and medium

By driving agent consensus for rocket engine model selection using reinforcement learning algorithms and Bayesian hierarchical models, and combining numerical transformation of cloud models, the problems of information loss and low consensus efficiency in complex group decision-making are solved, achieving a high-precision and efficient decision-making process.

CN120996079AActive Publication Date: 2025-11-21NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202511513895.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle complex multi-criteria decisions in rocket engine model selection, and are unable to quantify probabilistic language evaluations, leading to information loss and inefficient consensus adjustment, thus affecting decision-making accuracy and efficiency.

Method used

The algorithm uses reinforcement learning to adjust the probabilistic language evaluation matrix, and combines Bayesian hierarchical model and cloud model. It drives agent consensus through reward function, calculates attributes and agent weights, realizes the fidelity conversion of probabilistic language to numerical values, and integrates multi-agent evaluation.

Benefits of technology

It improves the accuracy and efficiency of consensus decision-making in large-scale complex groups, avoids redundant calculations, realizes the faithful conversion and synchronous digitization of key information, and enhances the accuracy and speed of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996079A_ABST
    Figure CN120996079A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale complex group consensus decision-making method, terminal equipment and a medium. The method comprises the following steps: acquiring a probability language evaluation matrix of a plurality of agents participating in decision-making on a plurality of candidate schemes on a plurality of scheme attributes; adjusting the probabilistic language evaluation matrix based on a reinforcement learning algorithm, and driving a plurality of agents to reach a consensus to obtain a new probabilistic language evaluation matrix; for the new probabilistic language evaluation matrix, utilizing a Bayesian hierarchical model to calculate an attribute weight corresponding to each scheme attribute; according to the hesitation degree and the consensus degree of each agent in the new probability language decision matrix, determining an agent weight corresponding to the agent; converting the new probability language evaluation matrix into a numerical evaluation matrix through a cloud model; fusing the numerical evaluation matrix, the attribute weight and the agent weight to obtain a scheme evaluation cloud model; and determining a final scheme for achieving the large-scale complex group consensus from the plurality of candidate schemes according to the scheme evaluation cloud model. According to the invention, the precision of large-scale complex group consensus decision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of large-scale complex group decision-making, and particularly relates to a large-scale complex group consensus decision-making method, a terminal device and a medium. BACKGROUND

[0002] In the field of aerospace engineering, rocket engine type selection is a key decision-making link that determines the performance of a launch vehicle and the success of a mission. With the rapid development of aerospace technology, new types of engines are constantly emerging, and the decision-making environment is becoming increasingly complex. Traditional large-scale complex group consensus decision-making methods have been unable to meet the high-precision and high-reliability decision-making requirements.

[0003] The existing technology mainly has the following outstanding problems: On the one hand, rocket engine selection needs to consider more than ten indicators such as thrust performance, specific impulse, reliability, cost, development cycle, and technology maturity. Traditional multi-criteria decision-making methods such as AHP and TOPSIS are difficult to effectively handle such a complex evaluation system. When evaluating new engines, intelligent agents often use probability language to express uncertainty, such as "high specific impulse is likely to be achieved" and "reliability may be affected". However, existing methods lack effective quantitative processing mechanisms. Moreover, the determination of indicator weights is mostly dependent on subjective judgment by intelligent agents, which is easily influenced by personal preferences, and different intelligent agents have different perceptions of the importance of indicators, making it difficult to reach a consensus.

[0004] On the other hand, when dealing with probability language evaluations, existing technology usually uses expected values or simple weighted averages for quantification, resulting in the loss of a large amount of valuable information. For example, an intelligent agent's reliability evaluation of a certain type of engine "{high (0.6), medium (0.4)}" and "{medium (1.0)}" may result in the same value in traditional methods, but the former contains important uncertainty information. In high-risk decisions such as rocket engine selection, ignoring uncertainty can lead to disastrous consequences, and existing methods cannot meet the stringent requirements of aerospace engineering for decision-making accuracy.

[0005] On the other hand, aerospace project decision-making usually involves intelligent agents from multiple professional fields such as aerodynamics, structure, propulsion, and control, with significant differences in opinions and a long consensus reaching process. Traditional consensus adjustment relies on meetings or email communication, which is inefficient, and a single decision often takes several weeks. The adjustment process lacks intelligent guidance, often resulting in "adjusting non-critical opinions" or "over-adjustment leading to intelligent agent resistance", which affects decision-making quality and team collaboration.

[0006] In summary, the existing technology has technical defects such as insufficient information expression, inaccurate quantification, and unintelligent consensus adjustment when dealing with complex aerospace engineering decision-making problems such as rocket engine type selection. These technical defects will affect the accuracy and efficiency of rocket engine type selection. SUMMARY

[0007] The technical problem solved by the present application is to provide a large-scale complex group consensus decision-making method, a terminal device and a medium, and to improve the accuracy of large-scale complex group consensus decision-making.

[0008] In a first aspect, the present application provides a large-scale complex group consensus decision-making method, which comprises the following steps: Obtaining a probability linguistic evaluation matrix of a plurality of intelligent agents participating in decision-making on a plurality of candidate schemes in a plurality of scheme attributes; the candidate scheme is a rocket engine model, and the scheme attribute is a rocket engine performance evaluation index; Adjusting the probability linguistic evaluation matrix based on a reinforcement learning algorithm to drive the plurality of intelligent agents to reach a consensus and obtain a new probability linguistic evaluation matrix; wherein the reward function of the reinforcement learning algorithm represents the change in the degree of hesitation and the degree of consensus of the intelligent agent before and after the action, and the action is designed to guide the intelligent agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other intelligent agents; For the new probability linguistic evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated using a Bayesian hierarchical model; the attribute weight is used to measure the importance of the scheme attribute; According to the degree of hesitation and the degree of consensus of each intelligent agent in the new probability linguistic decision matrix, the intelligent agent weight corresponding to the intelligent agent is determined; the intelligent agent weight is used to measure the objectivity of the intelligent agent; The new probability linguistic evaluation matrix is converted into a numerical evaluation matrix through a cloud model; the cloud model realizes the faithful conversion of probability language to numerical value based on three parameters of expected value, entropy and hyper-entropy, the expected value is used to measure the evaluation level of the intelligent agent, the entropy is used to measure the size of the difference between the intelligent agents, and the hyper-entropy is used to measure the evaluation stability of the intelligent agent; The numerical evaluation matrix, the attribute weight and the intelligent agent weight are fused to obtain a scheme evaluation cloud model; the scheme evaluation cloud model is used to model the fuzziness and randomness in the evaluation process of the large-scale complex group; According to the scheme evaluation cloud model, the final scheme reaching the consensus of the large-scale complex group is determined from the plurality of candidate schemes.

[0009] Optionally, the expression of the reward function is ; wherein, represents the reward value corresponding to the action, represents the change in the degree of hesitation, represents the change in the degree of consensus, , represents the degree of hesitation of the intelligent agent in the time step evaluating the candidate scheme in the scheme attribute, .​​ , This represents the total number of candidate solutions. , Indicates the total number of scheme attributes. , This represents the total number of intelligent agents. , Indicates the total time steps. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of hesitation in the attribute of the first option, the The evaluation of the first agent The candidate solution is in the... The formula for calculating the degree of hesitation on each option attribute is: , Indicates the first The evaluation of the first agent The candidate solution is in the... The maximum probability among the language terms of each scheme attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... The minimum probability among the language terms of each scheme attribute. express The base number, Indicates the first The evaluation of the first agent The candidate solution is in the... A set of probabilistic linguistic evaluations of the attributes of each solution. , Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the proposed solution and other intelligent agents in terms of its attributes. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the first scheme and other intelligent agents in terms of attributes, the degree of consensus between the second scheme and other intelligent agents. The evaluation of the first agent The candidate solution is in the... The formula for calculating the consensus degree between a scheme and other intelligent agents based on its attributes is as follows: , Indicates the first The evaluation of the first agent The candidate solution is in the... a probability linguistic evaluation vector on the jth scheme attribute, represents the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute, a probability linguistic evaluation vector on the jth scheme attribute, represents the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute, a probability linguistic evaluation vector on the jth scheme attribute, represents the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute, a probability linguistic evaluation vector on the jth scheme attribute, represents the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute, a probability linguistic evaluation vector on the jth scheme attribute, represents the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute, the value update rule of the reinforcement learning algorithm is: wherein, denotes the learning rate, denotes the reward signal, which is a measure of consensus improvement or error reduction, denotes the discount factor, denotes the current state, which is used to determine the importance of future rewards relative to immediate rewards, denotes the next state the maximum value of all possible actions ; the expression of the action of the reinforcement learning algorithm is: wherein, denotes the updated evaluation of the jth agent's evaluation of the jth candidate scheme on the jth scheme attribute at time step denotes the control parameter, , The smaller the value of the control parameter, the more the action adjustment tends to preserve the original information. denotes the group average of all agents' evaluations of the jth candidate scheme on the jth scheme attribute, calculated as a weighted average or an arithmetic average of the probability distribution, and UCA represents the agent, candidate scheme, and scheme attribute triple.

[0010] Optionally, for the new probability linguistic evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated using a Bayesian hierarchical model, including: According to the new probability linguistic evaluation matrix, the evaluation deviation of the agent when evaluating different candidate schemes on the same scheme attribute is calculated; the evaluation deviation is used to measure the importance of the agent to the scheme attribute; ​​​​​​​According to the evaluation deviation, optimal attributes and worst attributes are determined from multiple scheme attributes, and a dominance comparison matrix and a gap comparison matrix are calculated based on a possibility degree algorithm respectively; the dominance comparison matrix is used for measuring the preference degree of the optimal attributes relative to other scheme attributes, and the gap comparison matrix represents the preference degree of other scheme attributes relative to the worst attributes; The dominance comparison matrix and the gap comparison matrix are modeled by using a polynomial distribution, to obtain a first polynomial distribution expression and a second polynomial distribution expression; A Bayesian hierarchical model is constructed for the first polynomial distribution expression and the second polynomial distribution expression, and a Markov chain Monte Carlo algorithm is used to solve the Bayesian hierarchical model, to obtain attribute weights.

[0011] Optionally, the expression of the evaluation deviation is: wherein, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, , represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, represents the evaluation deviation of the i th intelligent agent on the j th scheme attribute, The expression of the optimal attributes is: ; The expression of the worst attributes is: ; The expression of the first polynomial distribution expression is: ; The expression of the second polynomial distribution expression is: ; For the new probability language evaluation matrix, a Bayesian hierarchical model is used to calculate attribute weights corresponding to each scheme attribute, including: Calculate the multinomial probability distribution of the worst-case attribute to characterize the other attributes relative to the worst-case attribute. The probability of the preference relationship; the expression for the probability of the preference relationship is: ,in, This represents a matrix comparing the preferences of other attributes for the worst-case attribute. , This represents the attribute weight vector to be determined. , , , For the elements of the preference matrix, reflecting the first Each solution attribute is compared with the worst attribute. The degree of preference, Denotes the polynomial coefficients, ensuring that the sum of probabilities equals 1; Calculate the multinomial probability distribution of the optimal attribute to characterize the optimal attribute. The probability of preference relationships relative to other attributes is expressed as follows: in, This represents a matrix comparing the preferences of the optimal attribute for other attributes. , The elements of the preference matrix reflect the optimal attributes. For the The degree of preference for each option attribute; To characterize the joint probability of the Bayesian hierarchical model, we fuse the preferences of multiple agents and construct the joint probability distribution of the hierarchical model, which connects input preferences and output weights. Its expression is: in, Indicates the overall attribute weight The prior distribution is determined by using the non-informative Dirichlet distribution to avoid subjective bias, and satisfies... , The distribution is a Dirichlet distribution, and when all parameters are 1, it is equivalent to a uniform distribution. For the first Attribute weights of experts The conditional distribution of the overall weights follows a Dirichlet distribution. , Indicates concentration parameter, Used for control and The degree of closeness The larger the value, the more concentrated the expert weight. , Modeling is required using gamma distribution. , , For shape parameters, , The first The optimal and worst attribute preference matrix of each expert , its own weight Conditional distribution; The posterior distribution is iteratively solved using Markov chain Monte Carlo techniques, and the attribute weights are finally output; among them, the comprehensive attribute weights of all experts are included. The weights are obtained by aggregating the posterior distributions of the weights from each expert, and are the attribute weights ultimately used for decision-making.

[0012] Optionally, the expression for the agent weights is: in, Indicates the first The agent weights corresponding to each agent. Indicates the first Hesitation level of an agent Indicates the first The consensus degree of an agent is 0 when there is exactly one probabilistic language term. If there are 1 linguistic variable and they have equal probability, then the degree of consensus is 1. .

[0013] Optionally, the expression for the cloud model is: in, Indicates the first The evaluation of the first agent The candidate solution is in the... The cloud model transformation result for each scheme attribute is typically a triple (Ex, En, He), used to quantify the uncertainty of the evaluation. Expressing expectations The square of, Represents the square of the probability. The average value of entropy. This represents the average value of hyperentropy. This represents the expected value, which is the first rank of the agent's evaluation. The candidate solution is in the... Evaluation center values ​​for each scheme attribute, The domain that represents the cloud model.

[0014] Optionally, the expression for the cloud model for evaluating the solution is: .

[0015] Optionally, based on the scheme evaluation cloud model, the final scheme for achieving consensus in large-scale complex groups is determined from multiple candidate schemes, including: According to the cloud model Determine the positive ideal solution from the multiple solution attributes. and negative ideal solution ;in, , , , Take the absolutely ideal solution, which represents the state values ​​of the scheme's attribute elements when they are theoretically at their best and worst. Based on shape similarity and distance similarity, the similarity between each candidate solution and the positive ideal solution and the negative ideal solution is calculated; wherein, the expression for the solution similarity is: or , ; Through calculation formula Candidate solutions were obtained. Compared to the ideal solution and negative ideal solution Proximity between Proximity reflects the candidate solution The degree to which a candidate solution is similar to a positive ideal solution but not to a negative ideal solution; the smaller the similarity, the better the performance of the candidate solution. The candidate solution with the highest similarity will be selected as the final solution.

[0016] In a second aspect, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0017] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0018] The present invention has at least the following beneficial effects: The reward function of the reinforcement learning algorithm is used to represent the change of the hesitation degree and the change of the consensus degree of the agent before and after performing the action, and the action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents, and the local adjustment instead of global iteration is driven by the reward function, which avoids redundant calculation, reduces the time consumption of complex decision consensus process, and improves the consensus efficiency, thereby improving the accuracy of the large-scale complex group consensus decision; the cloud model realizes the faithful conversion of the probability language to the value based on the three parameters of the expected value, the entropy and the hyper entropy, realizes the synchronous digitization of the fuzziness and randomness in the probability language, avoids the loss of key information, and improves the accuracy of the large-scale complex group consensus decision. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification, and are used to explain the technical solutions of the present application together with the embodiments of the present application, and do not constitute a limitation on the technical solutions of the present application.

[0020] Figure 1 A flowchart of the large-scale complex group consensus decision method in one of the embodiments of the present application is shown in the figure. Figure 2 A change diagram of the consensus degree and the hesitation degree of the agent consensus in one of the embodiments of the present application is shown in the figure. Figure 3 A structure diagram of the terminal device in one of the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0022] The present application provides a large-scale complex group consensus decision-making method, a terminal device and a medium, which aims to solve the technical problem of low accuracy of traditional large-scale complex group consensus decision-making methods. The reward function of the reinforcement learning algorithm used in the present application is used to represent the change in the hesitation degree and the change in the consensus degree of the agent before and after the action is performed. The action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents. The reward function drives local adjustment instead of global iteration, avoiding redundant calculation, reducing the time-consuming of complex decision-making consensus process, and improving the efficiency of consensus achievement, thereby improving the accuracy of large-scale complex group consensus decision-making. The cloud model realizes the faithful conversion of probability language to numerical value based on the three parameters of expected value, entropy and hyper entropy, realizes the simultaneous mathematicalization of fuzziness and randomness in probability language, avoids the loss of key information, and improves the accuracy of large-scale complex group consensus decision-making.

[0023] Embodiment one As shown in Figure 1 The large-scale complex group consensus decision-making method provided by the present application comprises the following steps: Step 11, obtaining a probability language evaluation matrix of multiple agents participating in decision-making on multiple candidate schemes in multiple scheme attributes.

[0024] In the embodiment of the present application, the candidate scheme is a rocket engine model. For example, the multiple candidate schemes are shown in Table 1.

[0025] Table 1

[0026] The scheme attribute is a performance evaluation index of the rocket engine. In one possible implementation, the scheme attribute includes specific impulse, thrust, thrust-to-weight ratio, service life and stability.

[0027] Specifically, in the embodiment of the present application, the agent represents a decision-making subject participating in decision-making, which can be a human expert or an artificial intelligence system.

[0028] In one possible implementation, Likert scale 5-point scale is used to depict the evaluation information of the agent on the candidate scheme. Likert scale 5-point scale is represented as: ={very disagree, not very agree, difficult to judge, more agree, very agree}. The evaluation value can be represented as .

[0029] Step 12, adjusting the probability language evaluation matrix based on the reinforcement learning algorithm to drive multiple agents to reach a consensus and obtain a new probability language evaluation matrix.

[0030] In the embodiment of the present application, the reward function of the reinforcement learning algorithm represents the change of the hesitation degree and the change of the consensus degree of the agent before and after the action is performed, and the action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents.

[0031] The process of adjusting the probabilistic linguistic evaluation matrix based on the reinforcement learning algorithm to drive multiple agents to reach a consensus and obtain a new probabilistic linguistic evaluation matrix is described below, specifically including steps 12.1 to 12.5: Step 12.1, initialize the state space, action space, learning rate, exploration rate and discount factor of the reinforcement learning algorithm.

[0032] Specifically, the state of the reinforcement learning algorithm is the average consensus degree of all agents in the current decision-making process. The action can be represented as the score of all agents on a certain candidate scheme on a certain scheme attribute being changed to the score of another agent on the scheme on the attribute, which is expressed as , where is the learning object, is the adjusted attribute index, is the adjusted scheme index.

[0033] Step 12.2, determine the learning object from the agent with a consensus degree less than the average consensus degree of the large-scale complex group, and select the action with the maximum Q value in the current state.

[0034] Specifically, the expression of the adjustment action is: , where represents the updated evaluation of the th agent on the th candidate scheme on the th scheme attribute at time step , and represents a control parameter, , The smaller the control parameter is, the more the action adjustment tends to preserve the original information, represents the group average of the evaluation of all agents on the th candidate scheme on the th scheme attribute j, which is calculated as a weighted average or an arithmetic average of the probability distribution, and UCA represents a triple of agent, candidate scheme and scheme attribute.

[0035] Step 12.3, calculate the updated average consensus degree of the large-scale complex group and calculate the immediate reward.

[0036] The expression of the reward function is ; where represents the reward value corresponding to the action, represents the change in hesitation, represents the change in consensus, , represents the hesitation of the i-th agent on the j-th solution attribute of the k-th candidate solution at the t-th time step, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the hesitation of the i-th agent on the j-th solution attribute of the k-th candidate solution at the t-th time step, , represents the total number of candidate solutions, , represents the total number of solution attributes, , represents the total number of agents, , represents the total number of time steps, represents the hesitation of the i-th agent on the j-th solution attribute of the k-th candidate solution at the t-th time step, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the hesitation of the i-th agent on the j-th solution attribute of the k-th candidate solution at the t-th time step, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, , represents the maximum value of the probability of the i-th agent in the linguistic term on the j-th solution attribute of the k-th candidate solution, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the minimum value of the probability of the i-th agent in the linguistic term on the j-th solution attribute of the k-th candidate solution, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the cardinality of , represents the set of probability linguistic evaluations of the i-th agent on the j-th solution attribute of the k-th candidate solution, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, , represents the consensus of the i-th agent on the j-th solution attribute of the k-th candidate solution with other agents at the t-th time step, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the consensus of the i-th agent on the j-th solution attribute of the k-th candidate solution with other agents at the t-th time step, represents the i-th agent, represents the k-th candidate solution, represents the j-th solution attribute, represents the consensus of the i-th agent on the j-th solution attribute of the k-th candidate solution with other agents at the t-th time step, represents the i-th agent, The evaluation of the first agent The candidate solution is in the... The formula for calculating the consensus degree between a scheme and other intelligent agents based on its attributes is as follows: , Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... The attributes of the first scheme are the same as the second. Euclidean distance between agents.

[0037] Step 12.4: Calculate the maximum Q value of the next state and update the Q value of the current state.

[0038] Specifically, The value update rule is as follows: ,in, Indicates the learning rate. This indicates a reward signal, a measure of increased consensus or reduced error. Indicates the discount factor. This indicates the current state and is used to determine the importance of future rewards relative to immediate rewards. Indicate the next state All possible actions The largest value.

[0039] Step 12.5: If the current state meets the preset termination condition, output a new probabilistic language evaluation matrix; otherwise, repeat steps 12.3 to 12.4 until the current state meets the preset termination condition.

[0040] In one feasible implementation, the preset termination condition is that the number of updates is greater than or equal to a preset threshold.

[0041] In this embodiment of the invention, as the reinforcement learning algorithm continues to iterate, the changes in the consensus degree and hesitation degree of the agent consensus are as follows: Figure 2 As shown, Figure 2 The horizontal axis represents the number of iterations, the green line represents the degree of consensus, and the red line represents the degree of hesitation. In the early stages of iteration, the degree of consensus is low due to potentially significant disagreements among experts. Figure 2It can be seen that as the reinforcement learning algorithm is adjusted, the agents' scores gradually converge, and the consensus increases accordingly. Furthermore, the consensus tends to stabilize with increasing iterations. Therefore, we can conclude that: 1. The final average consensus score reached 0.921, indicating that the agent was able to reach a consensus through the iterative process of the reinforcement learning algorithm. Meanwhile, although the hesitation score fluctuated significantly, it generally showed a downward trend, with the average hesitation score decreasing from approximately 0.42 to approximately 0.41, indicating that the agent's hesitation score also decreased through the iterative process of the reinforcement learning algorithm.

[0042] 2. The consensus level exceeding the threshold was reached around the 160th iteration, indicating that the reinforcement learning algorithm provided by this invention can effectively promote consensus formation within a reasonable number of iterations. This result effectively verifies the effectiveness of adjusting the probabilistic language evaluation matrix based on the reinforcement learning algorithm to achieve consensus in large-scale complex groups.

[0043] Step 13: For the new probabilistic language evaluation matrix, use the Bayesian hierarchical model to calculate the attribute weights corresponding to each scheme attribute.

[0044] In this embodiment of the invention, attribute weights are used to measure the importance of the scheme attributes.

[0045] Step 14: Determine the agent weight corresponding to each agent based on the hesitation degree and consensus degree of each agent in the new probabilistic language decision matrix.

[0046] In this embodiment of the invention, agent weights are used to measure the objectivity of the agent.

[0047] Specifically, the expression for agent weights is: in, Indicates the first The agent weights corresponding to each agent. Indicates the first Hesitation level of an agent Indicates the first The consensus degree of an agent is 0 when there is exactly one probabilistic language term. If there are 1 linguistic variable and they have equal probability, then the degree of consensus is 1. .

[0048] Step 15: Convert the new probabilistic language evaluation matrix into a numerical evaluation matrix using the cloud model.

[0049] In this embodiment of the invention, the cloud model achieves a fidelity conversion from probabilistic language to numerical values ​​based on three parameters: expected value, entropy, and hyperentropy. The expected value is used to measure the evaluation level of the agent, the entropy is used to measure the degree of disagreement of the agent, and the hyperentropy is used to measure the evaluation stability of the agent.

[0050] Specifically, the formalized expression of the cloud model is as follows: , representing the expected value, entropy, and hyperentropy, respectively. In one feasible implementation, the cloud model is expressed as: in, Indicates the first The evaluation of the first agent The candidate solution is in the... The cloud model transformation result for each scheme attribute is typically a triple (Ex, En, He), used to quantify the uncertainty of the evaluation. Expressing expectations The square of, Represents the square of the probability. The average value of entropy. This represents the average value of hyperentropy. This represents the expected value, which is the first rank of the agent's evaluation. The candidate solution is in the... Evaluation center values ​​for each scheme attribute, This represents the domain of the cloud model.

[0051] Step 16: The numerical evaluation matrix, attribute weights, and agent weights are fused to obtain the scheme evaluation cloud model.

[0052] In this embodiment of the invention, the comprehensive evaluation cloud model is used.

[0053] In one feasible implementation, the expression for the scheme evaluation cloud model is: .

[0054] Step 17: Based on the scheme evaluation cloud model, determine the final scheme that achieves consensus among multiple candidate schemes for large-scale complex groups.

[0055] Specifically, this includes steps 17.1 to 17.4: Step 17.1, based on the cloud model determining the positive ideal solution and the negative ideal solution from multiple scheme attributes . .

[0056] wherein, , , , The absolute ideal solution is taken, and the values of the scheme attribute elements are taken when they are theoretically optimal and worst.

[0057] Step 17.2, based on shape similarity and distance similarity, calculating the scheme similarity between each candidate solution and the positive ideal solution and the negative ideal solution.

[0058] wherein, the expression of the scheme similarity is or , , , ; Step 17.3, obtaining the closeness degree between the candidate solution relative to the positive ideal solution and the negative ideal solution by calculating the formula .

[0059] The closeness degree reflects the degree to which the candidate solution is similar to the positive ideal solution and is not similar to the negative ideal solution, and the smaller the closeness degree is, the better the performance of the candidate solution is.

[0060] Step 17.4, determining the candidate solution with the largest closeness degree as the final solution.

[0061] Embodiment Two This embodiment is a description of the process of calculating the attribute weight corresponding to each scheme attribute by using the Bayesian hierarchical model for the new probabilistic language evaluation matrix in Embodiment One.

[0062] Specifically, for the new probabilistic language evaluation matrix, the attribute weight corresponding to each scheme attribute is calculated by using the Bayesian hierarchical model, including: Step I, calculating the evaluation deviation of the agent when evaluating different candidate solutions on the same scheme attribute according to the new probabilistic language evaluation matrix.

[0063] In the embodiment of the application, the evaluation deviation is used to measure the importance of the agent to the scheme attribute.

[0064] Specifically, the expression of the evaluation deviation is: ​ wherein, represents the evaluation deviation of the th agent on the th scheme attribute, represents the evaluation deviation of the th agent on the th candidate scheme and the th candidate scheme on the th scheme attribute, represents the evaluation deviation of the th agent on the th candidate scheme on the th scheme attribute, represents the normalized probability linguistic evaluation vector of the th agent on the th candidate scheme on the th scheme attribute, which is usually normalized to eliminate the scale difference based on the evaluation set of all agents represents the normalized probability linguistic evaluation vector of the th agent on the th candidate scheme on the th scheme attribute.

[0065] Step II, according to the evaluation deviation, determining the optimal attribute and the worst attribute from the multiple scheme attributes, and calculating the dominance comparison matrix and the gap comparison matrix based on the possibility algorithm respectively.

[0066] In the embodiments of the present application, the dominance comparison matrix is used to measure the preference degree of the optimal attribute relative to other scheme attributes, and the gap comparison matrix represents the preference degree of other scheme attributes relative to the worst attribute.

[0067] Specifically, the expression of the optimal attribute is: ; The expression of the worst attribute is: .

[0068] Step III, modeling the dominance comparison matrix and the gap comparison matrix by using the multinomial distribution, to obtain the first multinomial distribution expression and the second multinomial distribution expression.

[0069] Specifically, the expression of the first multinomial distribution expression is: ; The expression of the second multinomial distribution expression is: .

[0070] Step IV, constructing a Bayesian hierarchical model for the first multinomial distribution expression and the second multinomial distribution expression, and solving the Bayesian hierarchical model by using the Markov chain Monte Carlo algorithm to obtain the attribute weight.​

[0071] For the new probabilistic language evaluation matrix, the process of calculating the attribute weights corresponding to each alternative attribute using a Bayesian hierarchical model includes: Calculate the multinomial probability distribution of the worst-case attribute to characterize the other attributes relative to the worst-case attribute. The probability of preference relationships; the expression for the probability of preference relationships is: ,in, This represents a matrix comparing the preferences of other attributes for the worst-case attribute. , This represents the attribute weight vector to be determined. , , , For the elements of the preference matrix, reflecting the first Each solution attribute is compared with the worst attribute. The degree of preference, Denotes the polynomial coefficients, ensuring that the sum of probabilities equals 1; Calculate the multinomial probability distribution of the optimal attribute to characterize the optimal attribute. The probability of preference relationships relative to other attributes is expressed as follows: in, This represents a matrix comparing the preferences of the optimal attribute for other attributes. , The elements of the preference matrix reflect the optimal attributes. For the The degree of preference for each option attribute; To characterize the joint probability of the Bayesian hierarchical model, we fuse the preferences of multiple agents and construct the joint probability distribution of the hierarchical model, which connects input preferences and output weights. Its expression is: in, Indicates the overall attribute weight The prior distribution is determined by using the non-informative Dirichlet distribution to avoid subjective bias, and satisfies... , The distribution is a Dirichlet distribution, and when all parameters are 1, it is equivalent to a uniform distribution. For the first Attribute weights of experts The conditional distribution of the overall weights follows a Dirichlet distribution. , Indicates concentration parameter, Used for control and The degree of closeness, The larger the value, the more concentrated the expert weight. , Modeling is required using gamma distribution. , , For shape parameters, , The first The optimal and worst attribute preference matrix of each expert , its own weight Conditional distribution; The posterior distribution is iteratively solved using Markov chain Monte Carlo techniques, and the attribute weights are finally output; among them, the comprehensive attribute weights of all experts are included. The weights are obtained by aggregating the posterior distributions of the weights from each expert, and are the attribute weights ultimately used for decision-making.

[0072] Example 3 This embodiment is an explanation of the final solution that achieves a large-scale complex group consensus by selecting from multiple candidate solutions based on the scheme evaluation cloud model in Embodiment 1.

[0073] Specifically, based on the scheme evaluation cloud model, the final scheme for achieving consensus in large-scale complex groups is determined from multiple candidate schemes, including: Step A, based on the cloud model Determine the ideal solution from multiple solution attributes and negative ideal solution .

[0074] in, , , , Take the absolutely ideal solution, which represents the state values ​​of the scheme's attribute elements when they are theoretically at their best and worst. Step B: Based on shape similarity and distance similarity, calculate the similarity between each candidate solution and the positive ideal solution and the negative ideal solution.

[0075] The expression for scheme similarity is: or , .

[0076] Step C, calculate using the formula. Candidate solutions were obtained. Compared to the ideal solution and negative ideal solution Proximity between Proximity reflects the candidate solution The greater the closeness degree, the better the performance of the candidate solution is represented, similar to the positive ideal solution and dissimilar to the negative ideal solution.

[0077] Step D, the candidate solution with the greatest closeness degree is determined as the final solution.

[0078] The large-scale complex group consensus decision-making method provided by the application has at least the following advantages: The reward function of the reinforcement learning algorithm is used to represent the change in the hesitation degree and the change in the consensus degree of the agent before and after the action is performed, the action is designed to guide the agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other agents, the local adjustment instead of the global iteration is driven by the reward function, the redundant calculation is avoided, the time consumption of the complex decision consensus process is reduced, the consensus reaching efficiency is improved, and thus the precision of the large-scale complex group consensus decision-making is improved.

[0079] As shown in Figure 3 , the embodiment of the application provides a terminal device, as shown in Figure 3 , the terminal device D10 of the embodiment comprises at least one processor D100 (only one processor is shown in Figure 3 ), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps in any of the method embodiments described above when executing the computer program D102.

[0080] Specifically, the processor D100 executes the computer program D102 to obtain a plurality of candidate schemes on a plurality of scheme attributes by a plurality of agents participating in decision-making; adjust the probability language evaluation matrix based on a reinforcement learning algorithm to drive the plurality of agents to reach a consensus and obtain a new probability language evaluation matrix; calculate the attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model for the new probability language evaluation matrix; the attribute weight is used to measure the importance of the scheme attribute; determine the agent weight corresponding to each agent according to the hesitation degree and the consensus degree of the agent in the new probability language decision matrix; convert the new probability language evaluation matrix into a numerical evaluation matrix through a cloud model; fuse the numerical evaluation matrix, the attribute weight and the agent weight to obtain a scheme evaluation cloud model; determine the final scheme reaching the large-scale complex group consensus from the plurality of candidate schemes according to the scheme evaluation cloud model. The reward function of the reinforcement learning algorithm is used to represent the change of the hesitation degree and the consensus degree of the agent before and after the action is performed, and the action is designed to guide the agent whose consensus degree is less than the average consensus degree of the large-scale complex group to learn from other agents, and the local adjustment instead of global iteration is driven by the reward function, which avoids redundant calculation, reduces the time consumption of the complex decision-making consensus process, improves the consensus reaching efficiency, and thus is beneficial to improving the precision of the large-scale complex group consensus decision-making; the cloud model realizes the faithful conversion of the probability language to the value based on the three parameters of the expected value, the entropy and the hyper entropy, realizes the synchronous digitization of the fuzziness and randomness in the probability language, avoids the loss of key information, and improves the precision of the large-scale complex group consensus decision-making.

[0081] The processor D100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.

[0082] The storage D101 can be an internal storage unit of the terminal device D10, such as a hard disk or a memory of the terminal device D10, in some embodiments. The storage D101 can also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device D10, in other embodiments. Further, the storage D101 can include both an internal storage unit and an external storage device of the terminal device D10. The storage D101 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program, etc. The storage D101 can also be used to temporarily store data that has been output or is to be output.

[0083] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0084] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is caused to implement the steps in the above method embodiments.

[0085] Those skilled in the art should understand that the discussion of any embodiment herein is merely intended to be illustrative of the present application and is not intended to limit the scope of the present application. The technical features of the above embodiments or different embodiments can be combined, the steps can be implemented in any order, and there are many other changes to the different aspects of one or more embodiments of the present application as described above. For the sake of brevity, they are not provided in detail.

[0086] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations of one or more embodiments of the present application falling within the broadest scope of the present application. Therefore, any omission, modification, equivalent replacement, improvement, etc. made in the spirit and principle of one or more embodiments of the present application should be included in the scope of the present application.

Claims

1. A method for large-scale complex group consensus decision-making, characterized in that, The method comprises the following steps: obtaining a plurality of intelligent agents participating in decision-making to obtain a plurality of candidate schemes on a plurality of scheme attributes; the candidate scheme is a rocket engine model, and the scheme attribute is a rocket engine performance evaluation index; adjusting the probability language evaluation matrix based on a reinforcement learning algorithm to drive the plurality of intelligent agents to reach a consensus and obtain a new probability language evaluation matrix; for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model; the attribute weight is used to measure the importance of the scheme attribute; determining an intelligent agent weight corresponding to each intelligent agent in the new probability language decision matrix according to the hesitation degree and the consensus degree of the intelligent agent; the intelligent agent weight is used to measure the objectivity of the intelligent agent; converting the new probability language evaluation matrix into a numerical evaluation matrix through a cloud model; the cloud model realizes the faithful conversion of probability language to numerical value based on three parameters of expected value, entropy and hyper entropy; fusing the numerical evaluation matrix, the attribute weight and the intelligent agent weight to obtain a scheme evaluation cloud model; the scheme evaluation cloud model is used for integrated modeling of fuzziness and randomness in the large-scale complex group evaluation process; determining a final scheme reaching the large-scale complex group consensus from the plurality of candidate schemes according to the scheme evaluation cloud model.

2. The method of claim 1, wherein, The reward function of the reinforcement learning algorithm represents the change of the hesitation degree and the consensus degree of the intelligent agent before and after the action is performed, and the action is designed to guide the intelligent agent with a consensus degree less than the average consensus degree of the large-scale complex group to learn from other intelligent agents; The expression for the reward function is: ;in, This represents the reward value corresponding to the action. This indicates the change in hesitation. This indicates the change in the degree of consensus. , Indicates at time step Time The evaluation of the first agent The candidate solution is in the... Hesitation level in the attributes of each option , This represents the total number of candidate solutions. , Indicates the total number of scheme attributes. , This represents the total number of intelligent agents. , Indicates the total time steps. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of hesitation in the attribute of the first option, the The evaluation of the first agent The candidate solution is in the... The formula for calculating the degree of hesitation on each option attribute is: , Indicates the first The evaluation of the first agent The candidate solution is in the... The maximum probability among the language terms of each scheme attribute. Indicates the first The evaluation of the first agent The candidate solution is in the... The minimum probability among the language terms of each scheme attribute. express The base number, Indicates the first The evaluation of the first agent The candidate solution is in the... A set of probabilistic linguistic evaluations of the attributes of each solution. , Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the proposed solution and other intelligent agents in terms of attributes. Indicates at time step Time The evaluation of the first agent The candidate solution is in the... The degree of consensus between the first scheme and other intelligent agents in terms of attributes, the degree of consensus between the second scheme and other intelligent agents. The evaluation of the first agent The candidate solution is in the... The formula for calculating the consensus degree between a scheme and other intelligent agents based on its attributes is as follows: , Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute Indicates the first The evaluation of the first agent The candidate solution is in the... Probabilistic linguistic evaluation vectors for each solution attribute Indicates the first The evaluation of the first agent The candidate solution is in the... The attributes of the first scheme are the same as the second. Euclidean distance between agents; the reinforcement learning algorithm The value update rule is: wherein, denotes a learning rate, denotes a reward signal, a measure of consensus improvement or error reduction, denotes a discount factor, denotes a current state, used to determine the importance of future rewards relative to immediate rewards, denotes a next state the maximum value over all possible actions ; The action of the reinforcement learning algorithm The expression is: ,in, Indicates at time step Time The evaluation of the first agent The candidate solution is in the... Evaluation of the updated attributes of each solution Indicates control parameters, , The smaller the value, the more the motion adjustment tends to retain the original information. This indicates the evaluation of all agents. The candidate solution is in the... The group average of the evaluation of j on each alternative attribute is calculated as a weighted average or arithmetic average of the probability distribution. UCA represents the triplet of agent, candidate alternative, and alternative attribute.

3. The method of claim 2, wherein, for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model, comprising: calculating the evaluation deviation of the intelligent agent when evaluating different candidate schemes on the same scheme attribute according to the new probability language evaluation matrix; the evaluation deviation is used to measure the importance of the intelligent agent to the scheme attribute; determining an optimal attribute and a worst attribute from the plurality of scheme attributes according to the evaluation deviation, and calculating an advantage comparison matrix and a gap comparison matrix based on a possibility algorithm; the advantage comparison matrix is used to measure the preference degree of the optimal attribute relative to other scheme attributes, and the gap comparison matrix represents the preference degree of other scheme attributes relative to the worst attribute; modeling the advantage comparison matrix and the gap comparison matrix by using a polynomial distribution to obtain a first polynomial distribution expression and a second polynomial distribution expression; constructing a Bayesian hierarchical model for the first polynomial distribution expression and the second polynomial distribution expression, and solving the Bayesian hierarchical model by using a Markov chain Monte Carlo algorithm to obtain the attribute weight.

4. The method of claim 3, wherein, The expression of the evaluation deviation is: wherein, represents the evaluation deviation of the i-th agent on the j-th scheme attribute, represents the evaluation deviation of the i-th agent on the j-th scheme attribute, and , represents the normalized probability linguistic evaluation vector of the i-th agent on the j-th scheme attribute, which is normalized to eliminate the scale difference, based on the evaluation set of all agents represents the normalized probability linguistic evaluation vector of the i-th agent on the j-th scheme attribute.​​​​​​​​​​​​​​ The expression of the optimal attribute is: ; The expression of the worst attribute is: ; The expression expressed by the first polynomial distribution is: ; The expression expressed by the second polynomial distribution expression is: ; for the new probability language evaluation matrix, calculating an attribute weight corresponding to each scheme attribute by using a Bayesian hierarchical model, comprising: a polynomial probability distribution of the worst attribute is calculated, which characterizes the preference relationship probabilities of other attributes with respect to the worst attribute ; the expression of the preference relationship probability is , where represents a preference comparison matrix of other attributes with respect to the worst attribute, , represents an attribute weight vector to be solved, , , , is a preference matrix element, reflecting the preference degree of the first attribute of the scheme with respect to the worst attribute , represents a polynomial coefficient, which ensures that the probability sum is 1; A polynomial probability distribution of the optimal attribute is calculated, characterizing the optimal attribute The expression of the preference relation probability with respect to other attributes is: wherein, represents the preference comparison matrix of the optimal attribute to other attributes, , is the preference matrix element reflecting the preference degree of the optimal attribute to the first attribute of the scheme. characterizing the joint probability of the Bayesian hierarchical model, fusing the preferences of multiple intelligent agents, constructing the joint probability distribution of the hierarchical model, and using the joint probability distribution to link the input preference and the output weight, and the expression is: where, represents the integrated attribute weight is the prior distribution of the integrated attribute weight, which is modeled as a non-informative Dirichlet distribution to avoid subjective bias, satisfying , is the Dirichlet distribution, which is equivalent to the uniform distribution when all parameters are 1, is the attribute weight of the th expert is the conditional distribution of the integrated weight, which is modeled as a Dirichlet distribution , represents the concentration parameter, , which is used to control the closeness of to , The larger the , needs to be modeled by a gamma distribution , a, b are shape parameters, , are the best and worst attribute preference matrices of the th expert , is the conditional distribution of its own weight ; The attribute weight is finally output by iteratively solving the posterior distribution through Markov chain Monte Carlo technology; wherein the comprehensive attribute weight of all experts The attribute weight used for decision-making finally is obtained by aggregating the posterior distribution of each expert weight.

5. The method of claim 4, wherein, the expression of the intelligent agent weight is: wherein, represents the weight of the agent corresponding to the agent, represents the hesitancy of the agent, represents the consensus of the agent, the consensus is 0 when there is only one probability language term, and the consensus is 1 when there are language variables with equal probability, and at this time .​​​ 6. The method of claim 5, wherein, An expression of the cloud model is: wherein, represents the nth intelligent agent evaluation of the mth candidate scheme on the nth scheme attribute, is a three-tuple (Ex, En, He), which is used to quantify the uncertainty of evaluation, represents the square of expectation, represents the domain of the cloud model.​​​​​​​​​ 7. The method of claim 6, wherein, The scheme evaluates the expression of the cloud model as: .

8. The method of claim 7, wherein The method for evaluating the cloud model according to the scheme to determine a final scheme for reaching a large-scale complex group consensus from the multiple candidate schemes includes: According to the cloud model , determining a positive ideal scheme from the multiple scheme attributes , and a negative ideal scheme ; wherein , , , Taking absolute ideal solution, representing the state value of the scheme attribute element when it is theoretically optimal and worst. Based on the shape similarity and the distance similarity, a scheme similarity between each of the candidate schemes and the positive ideal scheme and the negative ideal scheme is calculated; wherein, an expression of the scheme similarity is or, , ; The closeness between the candidate scheme and the positive ideal scheme and the negative ideal scheme is obtained by the calculation formula ;​​​​ The candidate scheme with the maximum closeness is determined as the final scheme.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor implements the method of any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Simulation model credibility evaluation method based on hesitant cloud language term set and group decision

    CN110717281A

  • Intelligent manufacturing capability maturity evaluation method, device and equipment

    CN118586737A

  • Game behavior dynamic evolution and strategy deduction optimization method and system in space field

    CN119539090A

  • Aircraft cluster cooperative control method based on MADRL

    CN120540339A

  • First-aid medical evacuation decision making system and method based on multi-agent reinforcement learning

    WO2023221956A1

Cited By

  • Method and device for generating medical question and answer pairs

    CN121561117A