Multilayer group consensus decision-making method based on intuitionistic fuzzy set and reinforcement learning driving

Through a multi-level group consensus decision-making method combining intuitive fuzzy sets and reinforcement learning, expert weights are dynamically adjusted, which solves the problem of insufficient adaptability of the weight adjustment mechanism in multi-attribute group decision-making, improves the efficiency of consensus achievement and rationality of decision-making results, and is suitable for complex scenarios such as medical resource planning.

CN120450481APending Publication Date: 2025-08-08DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566248.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The dynamic adaptability of the weight adjustment mechanism in the existing multi-attribute group decisions is insufficient, resulting in inefficient consensus reach.

Method used

A multi-layer group consensus decision-making method based on intuitive fuzzy sets and reinforcement learning is adopted to model expert weights through the Markov decision-making process, and the weight vector is dynamically adjusted using the Q-learning algorithm. Combined with the multi-layer consensus feedback mechanism and parameterized adjustment factors, expert weights are optimized to improve group consensus efficiency.

Benefits of technology

It improves the efficiency of reaching group consensus, realizes real-time dynamic adjustment in a large-scale heterogeneous expert group, maintains the rationality and interpretability of decision-making results, and is suitable for complex decision-making scenarios such as medical resource planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450481A_ABST
    Figure CN120450481A_ABST
Patent Text Reader

Abstract

In order to solve the problem that an existing multi-attribute group decision weight adjustment mechanism is insufficient in dynamic adaptability, the invention provides a multilayer group consensus decision method based on intuitive fuzzy set and reinforcement learning driving, and the method comprises the following steps: initializing expert weights and collecting preference matrixes expressed by intuitive fuzzy numbers of experts; an expert weight determination problem is formalized into a Markov decision process, and reinforcement learning elements are defined; according to the current expert weight and the preference matrix, the individual consensus degree and the group consensus degree of multiple layers of experts are calculated, and when a threshold value is not reached, the weight is adjusted according to a weight dynamic adjustment mechanism, and low-consensus expert opinions are corrected in a parameterized mode; after expert weight adjustment and unreasonable expert opinion parameterization adjustment, recalculating the group consensus degree, if a threshold value is reached, outputting a final expert weight and preference matrix, otherwise, continuing iteration until the group consensus reaches the standard; and finally, fusing the adjusted weight and the preference matrix, and calculating a scheme score sequence by adopting an intuitionistic fuzzy operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of group decision-making, and specifically is a multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning for application in medical resource planning. Background Art

[0002] Multi-attribute group decision-making is a key research area in artificial intelligence and decision science. As a cross-disciplinary field between these two disciplines, its essence lies in the evaluation of a limited number of options by multiple experts based on multi-attribute criteria, followed by ranking and selecting the best options through information integration. Its core value lies in the construction of a scientific collective intelligence decision-making system. By integrating heterogeneous information from multiple sources and diverse group cognition, this theory can effectively address challenges in complex decision-making situations. The consensus-building mechanism, the innovative core of this theoretical system, utilizes dynamic feedback adjustment algorithms and preference coordination strategies to not only enhance group acceptance of decision options but also, more importantly, build a scientific and democratic decision-making ecosystem. This ensures the objectivity of decision-making at the data-driven level and the compatibility of multiple stakeholders' demands at the value integration level, ultimately resulting in optimized decision solutions that are both professional and feasible. This decision-making paradigm provides effective solutions to a variety of complex problems. Currently, group consensus-building methods have been applied in practical scenarios such as medical resource planning, medical service selection, and urban development potential prediction.

[0003] In the field of multi-attribute group decision-making, intuitionistic fuzzy sets (IFS) can effectively capture the cognitive ambiguity in expert evaluations through a ternary representation of membership, non-membership, and hesitation. Using IFS to express expert opinions has become a mainstream approach to multi-attribute group decision-making. Group consensus-building methods based on IFS offer a new paradigm for solving complex decision-making problems. The determination of expert weights is particularly important in the consensus-building process. On the one hand, weight distribution determines the influence of different expert opinions in the group consensus, directly affecting the objectivity of the decision outcome. On the other hand, dynamic decision-making scenarios require real-time capture of the evolving characteristics of expert opinions, and weight adjustment mechanisms are crucial for ensuring consensus convergence. Existing methods include Meng's objective weighting method based on fuzzy entropy, which quantifies expert credibility through hesitation and geometric distance; Pang Jifang's subjective-objective hybrid model, based on a multi-level consensus measurement framework (attribute-solution-expert); and Zi Yingping's trust transfer mechanism based on social network analysis, which employs quantum cognitive theory to construct a phase-angle interference model of trust networks. However, although these methods improve the rationality of weight distribution, they generally face the problem of insufficient dynamic adaptability. Summary of the Invention

[0004] In order to solve the problem of insufficient dynamic adaptability of the existing multi-attribute group decision-making weight adjustment mechanism, the present invention provides a multi-layer group consensus decision-making method driven by intuitionistic fuzzy sets and reinforcement learning. It quantifies expert consensus layer by layer from attribute level, scheme level to preference matrix level, designs a feedback adjustment mechanism according to the consensus degree at each level, and models the expert weight optimization as a Markov decision process. The weight vector is dynamically adjusted using the Q-learning algorithm, and the group consensus is maximized through the reward function. This solves the problems of local convergence and insufficient dynamic adaptability in traditional static weight adjustment, and improves the efficiency of reaching group consensus.

[0005] The technical solution adopted by the present invention to solve its technical problems is:

[0006] A multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning, the steps include:

[0007] S1: Initialize expert weights and collect the expert preference matrix expressed in intuitionistic fuzzy numbers;

[0008] Intuitionistic fuzzy sets are an extension of traditional fuzzy sets. They describe uncertainty more comprehensively by considering the membership, non-membership, and hesitation of elements simultaneously, making it possible to describe uncertainty from three perspectives at the same time. Their definitions are as follows:

[0009] An intuitionistic fuzzy set defined on a non-empty set can be expressed as:

[0010] E(x)={〈x,μ E (x),v E (x)>|x∈X}

[0011] Among them, μ E (x): X → [0, 1] and v E (x):X→[0,1] are respectively called the membership function and non-membership function, and the hesitation function is as follows:

[0012] π E (x) = 1 - μ E (x)-v E (x)

[0013] μ E (x), v E (x) and π E (x), together express the degree to which the element x in the non-empty set X belongs to the intuitionistic fuzzy set E, where π E (x)∈[0,1]. At the same time, for any x∈X, the following always holds: 0≤μ E (x)+v E (x)≤1.

[0014] In large group decision-making problems, the preference matrix of the experts expressed in intuitionistic fuzzy numbers is:

[0015] u decision experts E={e1,e2,...,e n}(n≥20) constitute a group to participate in decision-making, e u represents the u-th decision expert, 1≤u≤n; the group decision expert gives p feasible alternatives and m attribute sets for the decision event; the solution set is A={A1,A2,...,A P}, the attribute set is C={C1,C2,...,C m}, A i represents the i-th solution; C j represents the jth attribute; the decision expert weight is W={w1,w2,...,w n},w u Represents the expert weight of the u-th decision expert, and satisfies Indicates expert e u The decision preference matrix, element Indicates expert e u For Plan A i Medium attribute C j decision preference value.

[0016] In group decision-making problems, expert preference matrices are used to represent expert opinions. The purpose of decision-making is to unify expert opinions to a certain degree and obtain the final solution.

[0017] S2: Formalize the expert weight determination problem as a Markov decision process, define the state vector, action vector, immediate reward and weight distribution of the new state, which constitute the elements of reinforcement learning;

[0018] The Markov Decision Process (MDP) describes the task of reinforcement learning. The main idea of reinforcement learning is that the decision-making individual estimates the utility function (reward function) of different actions under a certain environmental state through trial and error through interaction with the environment, and selects those actions that obtain the maximum reward value to continuously improve the strategy. Reinforcement learning is different from supervised learning. Reinforcement learning does not have pre-provided training examples to use. Its supervisory information comes only from the environmental reward value obtained at the end of the final decision. In most decision-making problems, each stage or even each step of the decision may affect the final decision result. That is, the final reward value is affected by one or more actions currently taken. Its purpose is to solve a global optimal strategy that can maximize the final reward, which is consistent with the purpose of multi-stage decision-making problems.

[0019] The Markov decision process is described as follows: An MDP corresponds to a 5-tuple \(E = \langle S, A, P, R, \gamma \rangle\). Where: \(S\) is the state space, and each state \(s \in S\) represents the description of the environment perceived by the decision maker; \(A\) is the action space, and \(a \in A\) is the action that the decision maker can take in the current state; \(P: S \times A \times S \to [0, 1]\) specifies the state transition probability, indicating that if an action \(a \in A\) acts on the current state \(s\), the environment will transfer from the current state to another state with the transition probability \(p\); \(R: S \times A\) is the reward function; \(\gamma\) is the discount factor.

[0020] One of the learning algorithms for reinforcement learning is the Q-learning algorithm. It uses the state-action value function \(Q(s, a)\) as the evaluation function, considers each action in the entire action space during each iteration, and ensures the final convergence of the learning process. The basic iterative formula of the Q-learning algorithm is as follows:

[0021]

[0022] Where: \(\alpha\) is the learning rate, and \(R\) is the reward value brought by the transition from state \(S\) t to \(S\) t+1 As can be seen from the above formula, the optimal policy is to adopt the action \(A\) that maximizes the \(Q\) value in the current state \(S\).

[0023] Defining the state vector, action vector, immediate reward, and weight distribution of the new state in step S2 is specifically as follows:

[0024] The state vector is defined as the current expert weight distribution, expressed as:

[0025]

[0026] represents the weight of the \(u\)-th expert at the \(t\)-th iteration, satisfying

[0027] The action vector is defined as the weight adjustment amount, expressed as:

[0028]

[0029] represents the weight adjustment amount of the \(u\)-th expert at the \(t\)-th iteration;

[0030] The immediate reward is defined as the change in the group consensus degree, expressed as:

[0031] R t = \(\lambda \cdot (C\) t+1 - C\) t )

[0032] Where \(\lambda\) is the reward scaling factor, and \(C\) trepresents the group consensus at the tth iteration;

[0033] The new state of weight distribution is generated by weight adjustment and is expressed as:

[0034] S t+1 =S t +O t .

[0035] S3: Based on the current expert weights and preference matrix, calculate the multi-level expert individual consensus and group consensus at the attribute level, solution level, expert level, and group level, and compare them with the preset consensus threshold. If the threshold is reached, output the expert weight and preference matrix. Otherwise, perform S4 expert weight adjustment operation and S5 preference matrix adjustment operation;

[0036] The multi-level expert individual consensus and group consensus at the attribute level, solution level, expert level, and group level are calculated as follows:

[0037] a. Attribute level: It represents the distance between the preference value of the jth attribute of the i-th solution of the u-th decision expert and the corresponding group matrix. The consensus degree at the attribute level is:

[0038]

[0039] in, Indicates expert e u About A i Plan C j The attribute-level consensus of the attribute relative to the same attribute in the group matrix;

[0040] b. The consensus level at the solution level is:

[0041]

[0042] in, Indicates expert e u About A i The consensus degree of all the attributes of the scheme relative to the same scheme in the group matrix; m is the number of attributes;

[0043] c. The consensus among experts is:

[0044]

[0045] Among them, C u Indicates expert e u The preference matrix-level consensus of the preference matrix relative to the group matrix, i.e., the consensus of individual experts; p is the number of feasible alternatives;

[0046] d. The consensus at the group level is:

[0047]

[0048] Among them, C represents the average consensus of all experts as the group consensus; n is the number of experts.

[0049] S4: Expert weights are adjusted according to the dynamic weight adjustment mechanism;

[0050] Specifically:

[0051] First, the weight adjustment ΔW is generated by the following formula:

[0052] ΔW=ρ·v+c1·r1·(p best -W)+c2·r2·(g best -W)+ξ

[0053] Among them, ρ∈[0.8,0,95] is the momentum attenuation coefficient, which controls the influence intensity of the historical direction; v represents the historical velocity vector, and the update rule is v t =ρ·v t-1 +(1-ρ)·ΔW t-1 , initial v 0 =0;p best Indicates the weight corresponding to the highest historical consensus of individual experts; g best represents the weight corresponding to the highest historical consensus of the group; c1, c2∈[0.5, 1.0] are local / global learning factors that control the intensity of individual global optimization; r1, r2∈[0, 1] are uniformly distributed random numbers; ξ is a random perturbation;

[0054] Set weight adjustment constraints to make the weight adjustment range in a single iteration controllable:

[0055] |Δw u |≤k·w u ,k=0.2

[0056] Normalization is performed after weight adjustment:

[0057]

[0058] in, represents the weight of the u-th expert at the t-th iteration; represents the weight adjustment of the u-th expert at the t-th iteration;

[0059] Map the continuous action ΔW to a predefined discrete action set, select the closest action by Euclidean distance, and update its Q value:

[0060]

[0061] Among them, α = 0.8 is the learning rate, which decays exponentially to 0.1 with the number of iterations; γ = 0.9 is the discount factor, emphasizing long-term reward accumulation; s t represents the state at time t; Q(s t ,ΔW i ) represents the action-value function in the Q-learning algorithm; R t Indicates that the immediate reward is defined as the change in group consensus; s t+1 Represents the new state of weight distribution; A' represents the action space;

[0062] The termination condition is: C t ≥δ,C t It represents the group consensus at the tth iteration, and δ is the preset consensus threshold, which is set to 0.9.

[0063] S5: Based on the calculation results of individual expert consensus, a feedback adjustment mechanism is used to perform parameterized adjustments on unreasonable expert opinions whose attribute, solution, or expert-level consensus is lower than the preset consensus threshold;

[0064] Based on the multi-level consensus measurement results, the present invention designs a feedback adjustment mechanism based on group consensus. This mechanism fine-tunes unreasonable expert opinions through a three-stage closed-loop process of "identification-positioning-correction". When the consensus degree is less than the preset consensus threshold δ, it is considered that the experts’ opinions on the attributes under the current scheme are biased; when the scheme-level .... When the consensus degree C is less than the preset consensus threshold δ, it is considered that the experts’ opinions on this alternative are biased; when the preference matrix level consensus degree C u When it is less than the preset consensus threshold δ, all opinions of this expert are considered to be biased;

[0065] Through a progressive structure, we achieve targeted improvement of decision-making consensus. Based on the three-dimensional consensus measurement model of attributes, solutions, and experts, we systematically identify key nodes that deviate from group consensus. Through set operations, we locate specific correction targets layer by layer, achieving problem location from the expert group to the solution attribute level. Finally, we use parameterized adjustment factors to implement differentiated evaluation corrections, driving group consensus convergence while retaining the rationality of the original opinions of experts. The details are as follows:

[0066] First, construct the inconsistent expert set E D :

[0067] E D ={e u |C u <δ}

[0068] The expert e that is less than the preset consensus threshold δ u Add inconsistent expert set E D middle;

[0069] Then find the set of options X that need to be modified D :

[0070]

[0071] The inconsistent expert set E D Medium-level consensus The options that are less than the preset consensus threshold δ are put into the option set X D middle;

[0072] Find the property set mod that needs to be modified:

[0073]

[0074] The set of alternatives X D Medium attribute level consensus Attributes that are less than the preset consensus threshold δ are put into the attribute set mod;

[0075] Modify the corresponding evaluation of experts whose consensus is lower than the preset consensus threshold in the attribute set mod:

[0076] [μ u (x ij ),v u (x ij )]=(1-ε)[μ u (x ij ),v u (x ij )]+ε[μ g (x ij ),v g (x ij )]

[0077] Among them, [μ u (x ij ),v u (x ij )] indicates expert e u About A i Plan C j Evaluation of attributes, μ u (x ij ),v u (x ij ) are membership degree and non-membership degree respectively; [μ g (x ij ),v g (x ij )] represents the group matrix A i Plan C j The evaluation of the attribute; ε is the adjustment factor that controls the weight between expert evaluation and group evaluation.

[0078] S6: After adjusting the expert weights and the parameterization of unreasonable expert opinions, recalculate the group consensus. If the threshold is reached, output the final expert weight and preference matrix. Otherwise, return to step S3 to continue iteration.

[0079] S7: After the final expert weights are output, they are aggregated with the preference matrix after parameterized adjustment of unreasonable expert opinions. The final group matrix is calculated based on the intuitionistic fuzzy weighted geometric mean operator, and the total score of the solutions in the group matrix is calculated based on the intuitionistic fuzzy number scoring function. Finally, the solution ranking is obtained.

[0080] The intuitionistic fuzzy weighted geometric mean operator is specifically:

[0081] Suppose there is a set of intuitionistic fuzzy numbers in Represent the membership and non-membership of r respectively; the intuitionistic fuzzy weighted geometric mean operator is defined as:

[0082]

[0083] Where ω=(ω1,ω2,...,ω n ) T represents the weight vector and satisfies

[0084] The total score of the solutions in the group matrix is calculated according to the intuitionistic fuzzy number score function in step S7, specifically:

[0085] Let the intuitionistic fuzzy number t=(μ t ,v t ), where μ t ,v t Represent membership and non-membership respectively; the intuitionistic fuzzy number score function is expressed as:

[0086] s(t)=μ t -v t

[0087] Where s(t) is the score value of t, s(t)∈[-1,1];

[0088] Assume t1 and t2 are intuitionistic fuzzy numbers. According to the intuitionistic fuzzy number scoring function, the following ranking method is obtained:

[0089] If s(t1)>s(t2), then t1>t2; if s(t1)=s(t2), then t1=t2.

[0090] The score of each plan in the group matrix is calculated by the intuitionistic fuzzy number scoring function, and the plans are sorted according to the score, and the decision-making process ends.

[0091] This application also protects the application of the above-mentioned multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning in medical resource planning.

[0092] The beneficial effects of the present invention include:

[0093] (1) By modeling the expert weight distribution as a Markov decision process and using the Q-learning algorithm to achieve dynamic gradient optimization, the problems of local convergence and insufficient adaptability in traditional static weight adjustment are solved, and the efficiency of reaching group consensus is improved; (2) A multi-layer consensus measurement system of "attribute-solution-expert-group" is constructed, and a three-dimensional feedback mechanism is combined to accurately locate the nodes of disagreement, effectively balancing the preservation of individual preferences and the convergence of group consensus; (3) Through a weight adjustment mechanism including momentum decay coefficient, historical velocity vector, local / global learning factor and random perturbation, combined with the Q-learning algorithm, expert weights are updated, integrating global search and long-term strategy optimization capabilities; (4) Parameterized adjustment factors and weight constraint mechanisms are designed to avoid the monopoly of high-consensus expert opinions and enhance the rationality and interpretability of decision results; (5) It supports real-time interaction among large-scale heterogeneous expert groups, improves the sensitivity of solution ranking, and maintains linear time complexity when processing tens of thousands of attribute data, which has significant engineering application advantages.

[0094] In summary, the present invention breaks through the limitations of traditional methods in efficiency, accuracy and scalability through reinforcement learning-driven dynamic weight optimization and multi-layer consensus quantization mechanism, and can cope with the large number of expert groups more quickly and effectively, which has important practical significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0096] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0097] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0098] The present invention proposes a multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning for medical resource planning. Aiming at the core problems of expert weight solidification and low consensus efficiency in complex multi-attribute group decision-making, a dynamic intelligent decision-making framework is proposed through the deep integration of reinforcement learning and intuitionistic fuzzy set theory. Its characteristics are: based on the Markov decision process, the expert weight distribution is modeled as a dynamic optimization problem, and the Q-learning algorithm is used to realize the adaptive gradient adjustment of weights, breaking the limitations of traditional static weight distribution; constructing a "attribute-scheme-expert-group" multi-layer consensus measurement system, accurately identifying the nodes of disagreement through a three-dimensional feedback mechanism, and combining parameterized adjustment factors to achieve a dynamic balance between individual preference retention and group consensus convergence; adopting a hybrid search strategy of particle swarm optimization and reinforcement learning, taking into account global exploration and local optimization capabilities, effectively avoiding local optimality and improving convergence efficiency; through the weight constraint mechanism and dynamic normalization processing, ensuring that the influence of experts is strictly matched with their professional contributions, while supporting large-scale heterogeneous groups to interact in real time. In addition to medical resource planning scenarios, this framework also provides a group consensus solution that is efficient, robust, and explainable for complex scenarios such as environmental governance and emergency decision-making, significantly improving the scientificity and practicality of multi-attribute decision-making.

[0099] The scheme is as follows: First, initialize the expert weights and collect the preference matrix expressed in intuitionistic fuzzy numbers; construct a Markov decision model, map the weight distribution to the state space, define the weight adjustment amount as the action space, design the reward function based on the consensus gain, and explore the optimal weight strategy in combination with the Q-learning algorithm; then, calculate the consensus at the attribute level, solution level, expert level and group level layer by layer. If the threshold is not reached, the three-dimensional feedback mechanism is triggered, and the low-consensus nodes are located through set operations, and the individual preferences are corrected using parameterized adjustment factors; introduce particle swarm optimization to accelerate the weight search process, and ensure the weight feasible domain constraints through dynamic normalization processing; cyclically execute the closed-loop process of consensus evaluation-weight optimization-feedback adjustment until the group consensus reaches the standard; finally, assemble the corrected expert preference matrix, generate the group decision matrix based on the intuitionistic fuzzy weighted geometric operator, and calculate the scheme ranking through the intuitionistic fuzzy number scoring function.

[0100] To verify its effectiveness, this method was applied to a real case of optimizing a medical resource planning scheme, verifying the consensus efficiency of group decision-making and the rationality of the scheme.

[0101] Example 1: A tertiary hospital faces an imbalance in the allocation of medical resources (average daily patient volume exceeds 8,000, bed occupancy reaches 120%, and key departments are operating at overload). Despite recent improvements to information technology and flexible scheduling, long outpatient wait times, congestion in examination departments, and a shortage of emergency resources remain unresolved. The hospital is currently considering optimizing resource allocation from four perspectives, each addressing four attribute dimensions: average patient wait time, equipment utilization, labor cost coefficient, and patient satisfaction. A committee of 20 in-hospital experts participated in the decision-making process, including four hospital leaders, five department heads, seven key clinical staff, three operations managers, and one information technology expert. The experts identified four medical resource allocation measures, forming the solution set A = {A1, A2, A3, A4}, and four attribute dimensions, forming the attribute set C = {C1, C2, C3, C4}.

[0102] Solution A1: Intelligent triage guidance system

[0103] Deploy AI systems to analyze the characteristics of waiting people (condition, priority, historical waiting data) in real time, dynamically allocate patients to idle consulting rooms or examination departments, and simultaneously provide waiting time predictions and route navigation.

[0104] Option A2: Cross-department equipment sharing pool

[0105] Establish a collaborative scheduling mechanism for large equipment such as CT / MRI in multiple departments such as radiology, cardiology, and orthopedics, break down departmental barriers, and dynamically allocate equipment usage time periods according to real-time demand.

[0106] Plan A3: Dynamic grouping of emergency medical staff

[0107] Based on historical emergency flow data (such as seasonal fluctuations and diurnal peaks), peak hours are predicted, and the medical team (including doctors, nurses, and technicians) is reorganized across departments in advance to form a flexible emergency response team.

[0108] Plan A4: Outpatient-Inpatient Resource Linkage System

[0109] Connect the outpatient electronic medical record, surgery scheduling and inpatient bed systems, predict hospitalization needs (such as the bed occupancy period of elective surgery patients) through algorithms, and dynamically coordinate outpatient examinations and inpatient bed allocation.

[0110] The following is a detailed introduction to the decision-making process:

[0111] (1) First, the experts are assigned initial weights. The experts express their preference evaluations for the alternatives and their attributes using intuitionistic fuzzy sets. The initial weights of all experts and the attribute weights under each alternative are averaged, and the sum of the expert weights and attribute weights is 1.

[0112] The decision-making experts’ initial preference evaluations of the four identified options and the experts’ initial preference evaluations of the option attributes:

[0113] Table 1 Initial expert preference matrix

[0114]

[0115]

[0116] (2) Define reinforcement learning elements according to step S2.

[0117] (3) Calculate the group consensus according to the description of step S3 and compare it with the preset consensus threshold δ=0.9.

[0118] (4) Obtain the determined expert weight through step S2-6.

[0119] Table 2 Expert weight table

[0120] Expert Number Weight Expert Number Weight <![CDATA[e1]]> 0.034 <![CDATA[e 11 ]]> 0.053 <![CDATA[e2]]> 0.034 <![CDATA[e 12 ]]> 0.047 <![CDATA[e3]]> 0.034 <![CDATA[e 13 ]]> 0.045 <![CDATA[e4]]> 0.043 <![CDATA[e 14 ]]> 0.045 <![CDATA[e5]]> 0.055 <![CDATA[e 15 ]]> 0.028 <![CDATA[e6]]> 0.034 <![CDATA[e 16 ]]> 0.035 <![CDATA[e7]]> 0.035 <![CDATA[e 17 ]]> 0.035 <![CDATA[e8]]> 0.034 <![CDATA[e 18 ]]> 0.045 <![CDATA[e9]]> 0.045 <![CDATA[e 19 ]]> 0.063 <![CDATA[e 10 ]]> 0.045 <![CDATA[e 20 ]]> 0.045

[0121] The majority of the resulting weight vectors are concentrated in a relatively small range (0.034, 0.045, and 0.053), indicating that most decision makers have a relatively balanced influence on the overall decision. However, a few decision makers have larger weights, such as 0.055 and 0.063. These larger weights indicate that the expert's opinions on the various options were closer to the group's, representing the decision makers whose opinions were more important in the group decision.

[0122] (5) Decision results

[0123] The target consensus (threshold = 0.90) was reached through 2 rounds of iterations.

[0124] Table 3. Number of iterations

[0125] Iteration rounds Group consensus initial 0.8477 Round 1 0.8860 Round 2 0.9050

[0126] (6) Rank the solutions calculated in step S7 and compare them with the results of other literature.

[0127] Table 4. The number of iterations of different methods and the ranking of solutions

[0128] Number of iterations Scheme ranking References[1] 4 <![CDATA[A4>A1>A2>A3]]> References[2] 6 <h2 style=";text-align:left;direction:ltr"><![CDATA[A1>A4>A3>A2]]><h2 style=";text-align:left;direction:ltr"> References[3] 3 <![CDATA[A1>A4>A3>A2]]> The present invention 2 <![CDATA[A1>A4>A3>A2]]>

[0129] Table 5 Solution scores of different methods

[0130] Solution score References[1] <![CDATA[A1=0.815、A2=0.746、A3=0.691、A4=0.932]]> References[2] <![CDATA[A1=0.7044、A2=0.4009、A3=0.4071、A4=0.6385]]> References[3] <![CDATA[A1=0.2657、A2=0.1298、A3=0.1872、A4=0.2094]]> The present invention <![CDATA[A1=1.0309、A2=0.242、A3=0.574、A4=0.6671]]>

[0131] Table 6 Principles of different methods

[0132] References[3] References[2] References[1] The present invention Weight basis Closeness coefficient Network centrality Hesitation Expert Similarity Whether to build a network no yes yes no Weight adjustment mechanism Static adjustment Dynamic Adjustment Static adjustment Gradient Adaptation Typical application scenarios Small fixed experts Large-scale dynamic groups Social network related groups Large-scale dynamic groups

[0133] Reference 1: Xu Xuanhua, Huang Li. Method and application of expert opinion and trust information fusion in large-scale emergency decision-making based on complex networks[J]. Data Analysis and Knowledge Discovery, 2022, 6(Z1): 348-363.

[0134] Reference 2: Liang Xia, Guo Jie, Liu Peide, et al. Dynamic large group emergency decision-making method based on group consensus[J]. Systems Science and Mathematics, 2025, 45(03):853-869.

[0135] Reference 3: Chen Xiaohong, Zhang Weiwei, Xu Xuanhua. Large group decision-making method based on hesitation and consistency in social network environment[J]. Systems Engineering - Theory and Practice, 2020, 40(05): 1178-1192.

[0136] While the four methods consistently ranked the solutions, the greater variance in scores for the reinforcement learning method suggests it more accurately reflects the strengths and weaknesses of the solutions. By dynamically adjusting expert weights, this method can better account for inter-expert differences, improve scoring sensitivity, and enhance decision accuracy.

[0137] In the field of group decision-making, although existing methods adjust expert weights through differentiation strategies to reach consensus, they generally have mechanism defects: the proximity coefficient method in reference [3] relies on static numerical proximity to assign weights. Although it can output stable rankings, it cannot capture the dynamic changes of opinions, resulting in a high number of iterations and insufficient consensus stability. Under the same ranking, the score variance reaches 15%; reference [2] further introduces expert similarity networks to construct graph structure relationships. Although it can mine group association characteristics, the complexity of its network topology calculation (6 iterations) and the risk of over-dominance of key nodes make it difficult to adapt to large-scale dynamic decision-making needs; reference [1] uses hesitation as the only adjustment basis. Although it focuses on uncertainty measurement, it ignores semantic direction differences. For example, (0.7, 0.2) and (0.2, 0.7) have the same hesitation but opposite tendencies, resulting in high-hesitation experts being incorrectly marginalized. In contrast, the dynamic weight adjustment method proposed in this invention breaks through the traditional limitations through three innovations: first, a joint similarity model integrating membership, non-membership and hesitation is constructed to quantify opinion uncertainty and capture semantic direction differences, thus avoiding the semantic misjudgment in the literature [1]; second, a real-time dynamic weight update mechanism is adopted to replace the static adjustment in the literature [3], reducing the number of iterations from 3 to 2, improving efficiency by 33%. At the same time, the score variance is compressed to within 5% through the gradient adaptive strategy, significantly enhancing the consensus stability; finally, the complex process of relying on subjective network construction in the literature [2] is abandoned, and the weight is directly calculated based on the similarity distance of individual opinions, reducing the algorithm complexity by 40%, making it more suitable for large-scale dynamic scenarios such as real-time risk assessment. Experiments show that under the same consensus threshold, the method of this invention is significantly superior to the comparison method in terms of ranking consistency and score difference.

[0138] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning, its characteristic steps include: S1: Initialize expert weights and collect the expert preference matrix expressed in intuitionistic fuzzy numbers; S2: Formalize the expert weight determination problem as a Markov decision process, define the state vector, action vector, immediate reward and weight distribution of the new state, which constitute the elements of reinforcement learning; S3: Based on the current expert weights and preference matrix, calculate the multi-level expert individual consensus and group consensus at the attribute level, solution level, expert level, and group level, and compare them with the preset consensus threshold. If the threshold is reached, output the expert weight and preference matrix. Otherwise, perform S4 expert weight adjustment operation and S5 preference matrix adjustment operation; S4: Expert weights are adjusted according to the dynamic weight adjustment mechanism; S5: Based on the calculation results of individual expert consensus, a feedback adjustment mechanism is used to perform parameterized adjustments on unreasonable expert opinions whose attribute, solution, or expert-level consensus is lower than the preset consensus threshold; S6: After adjusting the expert weights and the parameterization of unreasonable expert opinions, recalculate the group consensus. If the threshold is reached, output the final expert weight and preference matrix. Otherwise, return to step S3 to continue iteration. S7: After the final expert weights are output, they are aggregated with the preference matrix after parameterized adjustment of unreasonable expert opinions. The final group matrix is calculated based on the intuitionistic fuzzy weighted geometric mean operator, and the total score of the solutions in the group matrix is calculated based on the intuitionistic fuzzy number scoring function. Finally, the solution ranking is obtained.

2. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 1 is characterized in that: The preference matrix of the experts expressed in intuitionistic fuzzy numbers in step S1 is specifically: u decision experts E = {e1, e2, ..., e n }(n≥20) constitute a group to participate in decision-making, e u represents the u-th decision expert, 1≤u≤n; the group decision expert gives p feasible alternatives and m attribute sets for the decision event; the solution set is A={A1,A2,...,A P }, the attribute set is C={C1,C2,...,C m }, A i represents the i-th solution; C j represents the jth attribute; the decision expert weight is W={w1,w2,...,w n },w u Represents the expert weight of the u-th decision expert, and satisfies Indicates expert e u The decision preference matrix, element Indicates expert e u For Plan A i Medium attribute C j decision preference value.

3. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 2 is characterized in that: Step S2 defines the state vector, action vector, immediate reward, and new state of weight distribution, specifically: The state vector is defined as the current expert weight distribution, expressed as: Represents the weight of the u-th expert at the t-th iteration, satisfying The action vector is defined as the weight adjustment amount, expressed as: represents the weight adjustment of the u-th expert at the t-th iteration; The immediate reward is defined as the change in group consensus, expressed as: R t =λ·(C t+1 -C t ) Where λ is the reward scaling factor, C t represents the group consensus at the tth iteration; The new state of weight distribution is generated by weight adjustment and is expressed as: S t+1 =S t +O t 。 4. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 3 is characterized in that: The calculation of the multi-level expert individual consensus and group consensus at the attribute level, solution level, expert level, and group level in step S3 is specifically as follows: a. Attribute level: It represents the distance between the preference value of the jth attribute of the i-th solution of the u-th decision expert and the corresponding group matrix. The consensus degree at the attribute level is: in, Indicates expert e u About A i Plan C j The attribute-level consensus of the attribute relative to the same attribute in the group matrix; b. The consensus level at the solution level is: in, Indicates expert e u About A i The consensus degree of all the attributes of the scheme relative to the same scheme in the group matrix; m is the number of attributes; c. The consensus among experts is: Among them, C u Indicates expert e u The preference matrix-level consensus of the preference matrix relative to the group matrix, i.e., the consensus of individual experts; p is the number of feasible alternatives; d. The consensus at the group level is: Among them, C represents the average consensus of all experts as the group consensus; n is the number of experts.

5. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 4 is characterized in that: Step S4 is to adjust the expert weights according to the dynamic weight adjustment mechanism, specifically: First, the weight adjustment ΔW is generated by the following formula: ΔW=ρ v+c1 r1 (p best -W)+c2·r2·(g best -W)+ξ Among them, ρ∈[0.8,0,95] is the momentum attenuation coefficient, which controls the influence intensity of the historical direction; v represents the historical velocity vector, and the update rule is v t =ρ·v t-1 +(1-ρ)·ΔW t-1 , initial v 0 =0;p best Indicates the weight corresponding to the highest historical consensus of individual experts; g best represents the weight corresponding to the highest historical consensus of the group; c1, c2∈[0.5, 1.0] are local / global learning factors that control the intensity of individual global optimization; r1, r2∈[0, 1] are uniformly distributed random numbers; ξ is a random perturbation; Set weight adjustment constraints to make the weight adjustment range in a single iteration controllable: |Δw u |≤k·w u ,k=0.2 Normalization is performed after weight adjustment: in, represents the weight of the u-th expert at the t-th iteration; represents the weight adjustment of the u-th expert at the t-th iteration; Map the continuous action ΔW to a predefined discrete action set, select the closest action by Euclidean distance, and update its Q value: Among them, α = 0.8 is the learning rate, which decays exponentially to 0.1 with the number of iterations; γ = 0.9 is the discount factor, emphasizing long-term reward accumulation; s t represents the state at time t; Q(s t ,ΔW i ) represents the action-value function in the Q-learning algorithm; R t Indicates that the immediate reward is defined as the change in group consensus; s t+1 Represents the new state of weight distribution; A' represents the action space; The termination condition is: C t ≥δ,C t It represents the group consensus at the tth iteration, and δ is the preset consensus threshold.

6. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 5 is characterized in that: Step S5 uses the feedback adjustment mechanism to perform parameterized adjustments on unreasonable expert opinions whose attribute, solution or expert level consensus is lower than the preset consensus threshold. Specifically, when the attribute level consensus is lower than the preset consensus threshold, When the consensus degree is less than the preset consensus threshold δ, it is considered that the experts’ opinions on the attributes under the current scheme are biased; when the scheme-level .... When the consensus degree C is less than the preset consensus threshold δ, it is considered that the experts’ opinions on this alternative are biased; when the preference matrix level consensus degree C u When it is less than the preset consensus threshold δ, all opinions of this expert are considered to be biased; First, construct the inconsistent expert set E D : AND D ={and u |C u <δ} The expert e that is less than the preset consensus threshold δ u Add inconsistent expert set E D middle; Then find the set of options X that need to be modified D : The inconsistent expert set E D Medium-level consensus The options that are less than the preset consensus threshold δ are put into the option set X D middle; Find the property set mod that needs to be modified: The set of alternatives X D Medium attribute level consensus Attributes that are less than the preset consensus threshold δ are put into the attribute set mod; Modify the corresponding evaluation of experts whose consensus is lower than the preset consensus threshold in the attribute set mod: [m u (x ij ),v u (x ij )]=(1-e)[μ u (x ij ),v u (x ij )]+e[m g (x ij ),v g (x ij )] Among them, [μ u (x ij ),v u (x ij )] indicates expert e u About A i Plan C j Evaluation of attributes, μ u (x ij ),v u (x ij ) are membership degree and non-membership degree respectively; [μ g (x ij ),v g (x ij )] represents the group matrix A i Plan C j The evaluation of the attribute; ε is the adjustment factor that controls the weight between expert evaluation and group evaluation.

7. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 6 is characterized in that: The intuitionistic fuzzy weighted geometric mean operator in step S7 is specifically: Suppose there is a set of intuitionistic fuzzy numbers in Represent the membership and non-membership of r respectively; the intuitionistic fuzzy weighted geometric mean operator is defined as: Where ω=(ω1,ω2,...,ω n ) T Represents the weight vector and satisfies ω r ∈[0,1], 8. The multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning according to claim 7 is characterized in that: The total score of the solutions in the group matrix is calculated according to the intuitionistic fuzzy number score function in step S7, specifically: Let the intuitionistic fuzzy number t=(μ t ,v t ), where μ t ,v t Represent membership and non-membership respectively; the intuitionistic fuzzy number score function is expressed as: s(t)=μ t -v t Where s(t) is the score value of t, s(t)∈[-1,1]; Assume t1 and t2 are intuitionistic fuzzy numbers. According to the intuitionistic fuzzy number scoring function, the following ranking method is obtained: If s(t1)>s(t2), then t1>t2; if s(t1)=s(t2), then t1=t2. The score of each plan in the group matrix is calculated by the intuitionistic fuzzy number scoring function, and the plans are sorted according to the score, and the decision-making process ends.

9. Application of the multi-layer group consensus decision-making method based on intuitionistic fuzzy sets and reinforcement learning as described in any one of claims 1 to 8 in medical resource planning.