An Emergency Decision-making Method for Emergencies Based on Deep Reinforcement Learning
Through the deep reinforcement learning method, combined with TF-IDF theme mining and improved K-means clustering algorithm, the problem of difficulty in unified quantification of decision information in emergency decision-making is solved, effective coordination between experts and public opinions and participation of multiple opinions is achieved, and the quality of emergency decision-making is improved.
Patent Information
- Application Number
- CN202510664062.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When faced with dynamic and complex emergencies, existing emergency decision-making methods are difficult to unify and quantify decision-making information, resulting in insufficient decision-making quality, especially in the coordination of experts and public opinions.
Using a deep reinforcement learning method, the K-means clustering algorithm based on TF-IDF topic mining and improved, combined with sentiment analysis and weight derivation, the initial overall opinions of large groups are constructed, and the weight is adjusted using reinforcement learning to reach group consensus and calculate the optimal emergency decision-making plan.
It improves the quality of emergency decisions, and by balancing the weight of experts and the public, it promotes effective interaction between the two parties in the decision-making process and participation of multiple perspectives, ensuring the rationality and effectiveness of decision-making results.
Smart Images

Figure CN120181626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of public safety management, and in particular to an emergency decision-making method for emergencies based on deep reinforcement learning. Background Art
[0002] With the rapid development of society and the acceleration of globalization, the frequency and complexity of emergencies are increasing, posing severe challenges to public safety and economic development. These emergencies include, but are not limited to, extreme weather disasters, public health incidents, environmental pollution accidents, and other major public emergencies. How to make scientific and efficient emergency decisions and mitigate the impact of emergencies on society has become a research priority in both academia and practice.
[0003] Currently, many scholars are researching emergency decision-making methods for unconventional emergencies. Management scholars have developed models for emergency management decision-making using methods such as mathematical statistics and probability analysis, optimization analysis, data chain analysis, and scenario modeling. Furthermore, the rapid development of artificial intelligence (AI) technology has provided new approaches to solving emergency management decision-making problems. For example, swarm intelligence methods, fuzzy theory, artificial neural networks, support vector machines, and genetic algorithms have been effectively applied in decision-making related to disaster prediction and assessment.
[0004] Traditional emergency decision-making methods have long relied primarily on empiricism and rule-driven models. These approaches typically address emergencies by building expert rule bases, relying on historical data, and employing fixed decision-making frameworks. These expert rule bases are often built on past experience and are designed to address specific, historically relevant event types. However, with the increasing diversity, complexity, and frequency of emergencies, traditional emergency decision-making methods face increasing challenges, particularly in dynamic, complex, and highly uncertain environments, where their limitations are increasingly exposed. Summary of the Invention
[0005] The present invention provides an emergency decision-making method for sudden incidents based on deep reinforcement learning, the purpose of which is to improve the quality of decision-making by solving the problem that decision-making information is difficult to quantify uniformly during the emergency decision-making process.
[0006] To achieve the above objectives, the present invention provides an emergency decision-making method for emergencies based on deep reinforcement learning, comprising:
[0007] Step 1: Obtain the public's comments and multiple alternative plans for dealing with emergencies when they occur;
[0008] Step 2: Use the TF-IDF topic mining algorithm to mine the review text to obtain decision attributes and attribute weights, and analyze the review text to obtain the initial linear uncertain preference value as the initial opinion of the public;
[0009] Step 3: Input the decision attributes, attribute weights, and initial opinions of the public into the improved K-means clustering algorithm to cluster the initial opinions of the public and obtain the initial opinions of the expert subgroups. The initial opinions of the expert subgroups are then input into the weight derivation algorithm to derive the weights and obtain the initial weights of the expert subgroups and the initial weights of the public.
[0010] Step 4: Calculate the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts based on the initial weights of the expert subgroup, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroup. Based on the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts, and with the reinforcement learning goal of minimizing the time pressure of emergency response and social losses, input the initial weights of the expert subgroup and the initial weights of the public into the deep reinforcement learning network for adjustment to obtain the weights of the expert subgroup and the weights of the public.
[0011] Step 5: Calculate the group consensus measure using the weights of the expert subgroup and the public. When the group consensus measure is greater than the soft consensus threshold and the expert subgroup and the public reach a group consensus, adjust the initial overall opinion of the decision-making group to obtain the biased opinion of the decision-making group.
[0012] Step 6: Calculate the expected value of each alternative plan for dealing with emergencies using the expected function constructed based on the biased opinions of the decision-making group, and take the alternative plan with the highest expected value as the optimal emergency decision plan.
[0013] More specifically, step 2 includes:
[0014] Use the TF-IDF topic mining algorithm to mine frequently appearing feature words in the review text and determine the decision attributes;
[0015] Calculate the attribute weight by obtaining the frequency of frequently appearing feature words and the proportion of frequently appearing feature words to all words in the review text;
[0016] Sentiment analysis is performed on the review text to obtain the sentiment analysis results, which are then converted into initial linear uncertain preference values as the initial opinions of the public.
[0017] Furthermore, the calculation expression of attribute weight is:
[0018]
[0019] in, represents the attribute weight, Indicates frequently appearing feature words frequency, Indicates frequently appearing feature words The proportion of all words in the review text, Characteristic words quantity, Indicates the number of words in the review text.
[0020] Furthermore, the conversion formula for converting sentiment analysis results into initial linear uncertain preference values is:
[0021]
[0022]
[0023] in, represents the initial linear uncertain preference value, that is, the initial opinion of the public, Represents a conversion operation, Indicates the The sentiment value of the group sample, represents the minimum sentiment value, Indicates the maximum sentiment value, Indicates the public's response to the review text About The positive sentiment value of each decision attribute, Indicates the public's response to the review text About The negative sentiment value of the decision attribute, Indicates the public's response to the review text About The hesitation sentiment value of a decision attribute.
[0024] Furthermore, before step 3, the k-means clustering algorithm is also improved. The improvement process is as follows:
[0025] The similarity definition is given based on the distance between uncertain variables;
[0026] Calculate the similarity value between the expert preference opinions based on the similarity definition;
[0027] The expert preference opinions are sorted in descending order of similarity values, and the sorted expert preference opinions are used as the initial cluster center of each subgroup in turn. One expert preference opinion corresponds to the initial cluster center of one subgroup.
[0028] Furthermore, the initial opinions of the expert subgroups are input into the weight derivation algorithm to derive the weights, and the initial weights of the expert subgroups and the initial weights of the public are obtained, including:
[0029] By formula Calculate the opinion similarity of each expert subgroup, where Indicates the The similarity of opinions of expert subgroups, Indicates the Experts in the expert subgroup Opinions, Indicates the In the expert subgroup, except for the experts Other than opinions, represents the number of experts in the expert subgroup;
[0030] The comprehensive similarity of each expert subgroup is obtained by the opinion similarity of each expert subgroup and the number of experts in the expert subgroup. The expression is:
[0031]
[0032] in, Indicates the The comprehensive similarity of expert subgroups, represents the T-norm function, represents the number of experts in the expert subgroup, represents the number of experts in the decision-making group;
[0033] Through the The comprehensive similarity of expert subgroups Normalization is performed to obtain the initial weights of the expert subgroup and the initial weights of the general public, which are expressed as follows:
[0034]
[0035]
[0036] in, represents the initial weight of the expert subgroup, Represents the initial weight of the public.
[0037] Furthermore, the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups are used to calculate the initial overall opinions of the decision-making group, the reliability of the public, and the acceptability of the experts, including:
[0038] The initial overall opinion of the decision-making group is calculated based on the initial weights of the expert subgroup, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroup. The calculation expression is:
[0039]
[0040] in, represents the initial overall opinion of the decision-making group, represents the initial opinions of the expert subgroup;
[0041] The reliability of the public is determined based on the initial weight of the expert subgroup, the initial weight of the public, and the similarity of opinions between the public and the expert subgroup. The expression is:
[0042]
[0043] in, Indicates the similarity of opinions between the general public and expert subgroups;
[0044] The acceptability of the expert subgroup is calculated based on the initial overall opinion of the decision-making group and the initial opinion of the expert subgroup. The calculation expression is:
[0045]
[0046] in, represents the acceptability of the expert subgroup, represents the distance function.
[0047] Furthermore, the expression for calculating the group consensus measure by the weight of the expert subgroup and the weight of the public is:
[0048]
[0049] in, represents the group consensus measure, represents the weight of the expert subgroup, Indicates the weight of the public. Indicates the similarity of public opinions.
[0050] Furthermore, by adjusting the initial overall opinion of the decision-making group, the expression of the tendency opinion of the decision-making group is obtained as follows:
[0051]
[0052] in, Indicates the tendency opinions of a large decision-making group, expressed as linear uncertain preference values, Represents the feedback mechanism parameters.
[0053] More specifically, the expression of the expected function is:
[0054]
[0055] in, Representing variables The expected value of Indicates alternatives The reliability, Representing variables The inverse uncertainty distribution of represents the minimum value of the linear uncertain preference value, Indicates the maximum value of the linear uncertain preference value.
[0056] The above solution of the present invention has the following beneficial effects:
[0057] The present invention mines and analyzes the obtained comment texts of the public when an emergency occurs, obtains decision attributes, attribute weights, and the initial opinions of the public, inputs the improved K-means clustering algorithm to cluster the initial opinions of the public, obtains the initial opinions of the expert subgroups, and derives the weights of the expert subgroups and the initial opinions of the public to obtain the initial weights of the expert subgroups and the public; calculates the initial overall opinions of the decision-making group, the reliability of the public, and the acceptability of the experts based on the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups; takes the time pressure of the emergency response and the minimization of social losses as the reinforcement learning goals, inputs the initial weights of the expert subgroups and the public into the deep reinforcement learning network for adjustment, and obtains the weights of the expert subgroups and the public; The weighted method calculates the group consensus measure. When the group consensus measure is greater than the soft consensus threshold and the expert subgroup reaches a group consensus with the public, the initial overall opinion of the large decision-making group is adjusted to obtain the tendency opinion of the large decision-making group, which is used to construct an expectation function to calculate the expected value of each alternative plan for responding to emergencies, and the alternative plan with the highest expected value is used as the optimal emergency decision-making plan. Compared with the existing technology, the present invention mines and analyzes the comment texts of the public to obtain decision attributes, attribute weights, and the initial opinions of the public, which can solve the problem that decision information is difficult to be quantified uniformly in the emergency decision-making process. In the emergency decision-making process, the weights of the public and the expert subgroup are balanced through the training weight mechanism, which promotes the effective interaction between the two parties in the decision-making process. It can not only coordinate the opinions of experts and the public, but also ensure the effective participation of multiple views in decision-making, thereby improving the quality of decision-making.
[0058] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To make the technical problems, technical solutions, and advantages to be solved by the present invention more clear, the following is a detailed description with reference to the accompanying drawings and specific embodiments. It is obvious that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0061] In the description of the present invention, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and should not be understood as indicating or implying relative importance.
[0062] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0063] In response to existing problems, the present invention provides an emergency decision-making method for sudden incidents based on deep reinforcement learning.
[0064] like Figure 1 As shown, an embodiment of the present invention provides an emergency decision-making method for emergencies based on deep reinforcement learning, comprising:
[0065] Step 1: Obtain the public's comments and multiple alternative plans for dealing with emergencies when they occur;
[0066] Step 2: Use the TF-IDF topic mining algorithm to mine the review text to obtain decision attributes and attribute weights, and analyze the review text to obtain the initial linear uncertain preference value as the initial opinion of the public;
[0067] Step 3: Input the decision attributes, attribute weights, and initial opinions of the public into the improved K-means clustering algorithm to cluster the initial opinions of the public and obtain the initial opinions of the expert subgroups. The initial opinions of the expert subgroups are then input into the weight derivation algorithm to derive the weights and obtain the initial weights of the expert subgroups and the initial weights of the public.
[0068] Step 4: Calculate the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts based on the initial weights of the expert subgroup, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroup. Based on the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts, and with the reinforcement learning goal of minimizing the time pressure of emergency response and social losses, input the initial weights of the expert subgroup and the initial weights of the public into the deep reinforcement learning network for adjustment to obtain the weights of the expert subgroup and the weights of the public.
[0069] Step 5: Calculate the group consensus measure using the weights of the expert subgroup and the public. When the group consensus measure is greater than the soft consensus threshold and the expert subgroup and the public reach a group consensus, adjust the initial overall opinion of the decision-making group to obtain the biased opinion of the decision-making group.
[0070] Step 6: Calculate the expected value of each alternative plan for dealing with emergencies using the expected function constructed based on the biased opinions of the decision-making group, and take the alternative plan with the highest expected value as the optimal emergency decision plan.
[0071] In an embodiment of the present invention, a radiation environmental pollution incident is considered an emergency, and commentary texts related to the radiation environmental pollution incident are crawled from public websites as commentary texts of the public when the emergency occurs, and a variety of emergency plans for the radiation environmental pollution incident are obtained from the public websites as alternative solutions for dealing with the emergency.
[0072] Specifically, step 2 includes:
[0073] Use the TF-IDF topic mining algorithm to mine frequently appearing feature words in the review text and determine the decision attributes;
[0074] Calculate the attribute weight by obtaining the frequency of frequently appearing feature words and the proportion of frequently appearing feature words to all words in the review text;
[0075] Sentiment analysis is performed on the review text to obtain the sentiment analysis results, which are then converted into initial linear uncertain preference values as the initial opinions of the public.
[0076] In order to identify decision attributes and estimate corresponding weights in the process of large-scale emergency decision-making, the embodiment of the present invention adopts the TF-IDF algorithm for mining. The main idea of TF-IDF is that only words that appear frequently in multiple texts are more representative. The importance of sentiment words can be obtained through word frequency calculation. Let TF be the original word frequency, that is, the number of times a word appears in the comment text, and IDF represents the logarithm of the proportion of feature words appearing in the corpus to the entire comment text. An importance measurement result, namely the attribute weight, can be calculated by TF and IDF. The calculation expression is:
[0077]
[0078] in, represents the attribute weight, Indicates frequently appearing feature words The frequency of the word, Indicates frequently appearing feature words The ratio of all words in the review text, that is, the inverse document frequency, indicates that only words that appear frequently in multiple review texts are more representative. Characteristic words quantity, Indicates the number of words in the review text.
[0079] It should be noted that the purpose of sentiment analysis is to analyze commentary texts with opinions, including emotions, attitudes, and emotions related to emergencies, thereby achieving the classification of text sentiment tendencies. The embodiment of the present invention focuses on positive emotions in the text. After the sentiment analysis is completed, the values of positive sentences, negative sentences, and hesitant sentences of the associated alternatives based on its criteria are calculated. If the public has more positive emotions, they will have a higher level of satisfaction with the alternative. Therefore, public satisfaction can be expressed as the percentage of positive values associated with the corresponding attributes of each option. Therefore, the conversion formula for converting the sentiment analysis results into initial linear uncertain preference values is:
[0080]
[0081]
[0082] in, represents the initial linear uncertain preference value, that is, the initial opinion of the public, Represents a conversion operation, Indicates the The sentiment value of the group sample, represents the minimum sentiment value, Indicates the maximum sentiment value, Indicates the public's response to the review text About The positive sentiment value of each decision attribute, Indicates the public's response to the review text About The negative sentiment value of the decision attribute, Indicates the public's response to the review text About The hesitation sentiment value of a decision attribute.
[0083] Specifically, before step 3, the k-means clustering algorithm is also improved. The improvement process is as follows:
[0084] The similarity definition is given based on the distance between uncertain variables;
[0085] Calculate the similarity value between the expert preference opinions based on the similarity definition;
[0086] The expert preference opinions are sorted in descending order of similarity values, and the sorted expert preference opinions are used as the initial cluster center of each subgroup in turn. One expert preference opinion corresponds to the initial cluster center of one subgroup.
[0087] Specifically, the similarity definition is given based on the distance between uncertain variables, including:
[0088] Define two independent uncertain variables that follow a linear uncertainty distribution The distance expression between them is:
[0089]
[0090] in, Indicates uncertain variables The minimum value of Indicates uncertain variables The maximum value of Indicates uncertain variables The minimum value of Indicates uncertain variables The maximum value of Representing an event reliability;
[0091] The similarity is defined as follows based on the distance between uncertain variables:
[0092]
[0093] in, Indicates uncertain variables and The distance between.
[0094] As can be seen from the above formula, the higher the similarity, the closer the connection between decision-making objects. In the emergency decision-making process of large groups, decision-making knowledge is usually diverse. Therefore, it is necessary to apply cluster analysis to the dimensionality reduction process of large decision-making groups, divide the decision-making objects into subgroups with similar opinions, and improve the accuracy and efficiency of decision-making.
[0095] The embodiment of the present invention uses an improved k-means clustering algorithm to cluster the initial opinions of the public. The process is as follows:
[0096] By formula Calculate the similarity between expert preference opinions ;
[0097] Select the preference opinion with the greatest similarity As the initial cluster center of the first expert subgroup, and let , expressed as ;
[0098] choose As a second subgroup of experts The initial cluster center of and The similarity between them should be minimum, that is , expressed as ;
[0099] Next, select As the third expert subgroup The initial cluster center of and 、 The average similarity should be the smallest, that is , expressed as ;
[0100] Repeat the above steps K-1 times until ,and ,in, . The initial cluster center of the K-1 expert subgroups is , the initial cluster center set is ;
[0101] Calculate the similarity between individual preference relationship and cluster center, and transform individual matrix Merge to the most similar clusters In, that is .
[0102] Specifically, the initial opinions of the expert subgroups are input into the weight derivation algorithm to derive the weights, and the initial weights of the expert subgroups and the initial weights of the public are obtained, including:
[0103] By formula Calculate the opinion similarity of each expert subgroup, where Indicates the The similarity of opinions of expert subgroups, Indicates the Experts in the expert subgroup Opinions, Indicates the In the expert subgroup, except for the experts Other than opinions, represents the number of experts in the expert subgroup;
[0104] The comprehensive similarity of each expert subgroup is obtained by the opinion similarity of each expert subgroup and the number of experts in the expert subgroup. The expression is:
[0105]
[0106] in, Indicates the The comprehensive similarity of expert subgroups, represents the T-norm function, , represents the number of expert subgroups, and M represents the number of experts in the large decision-making group;
[0107] Through the The comprehensive similarity of expert subgroups Normalization is performed to obtain the initial weights of the expert subgroup and the initial weights of the general public, which are expressed as follows:
[0108]
[0109]
[0110] in, represents the initial weight of the expert subgroup, represents the initial weight of the public, .
[0111] Specifically, the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups are used to calculate the initial overall opinions of the decision-making group, the reliability of the public, and the acceptability of the experts, including:
[0112] The initial overall opinion of the decision-making group is calculated based on the initial weights of the expert subgroup, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroup. The calculation expression is:
[0113]
[0114] in, represents the initial overall opinion of the decision-making group, represents the initial opinions of the expert subgroup;
[0115] The reliability of the public is determined based on the initial weight of the expert subgroup, the initial weight of the public, and the similarity of opinions between the public and the expert subgroup. The expression is:
[0116]
[0117] in, Indicates the similarity of opinions between the general public and expert subgroups;
[0118] The acceptability of the expert subgroup is calculated based on the initial overall opinion of the decision-making group and the initial opinion of the expert subgroup. The calculation expression is:
[0119]
[0120] in, Indicates the acceptability of the expert subgroup.
[0121] Specifically, based on the initial overall opinions of the decision-making group, the reliability of the public, and the acceptability of the experts, and with the time pressure of emergency response and the minimization of social losses as the reinforcement learning goals, the initial weights of the expert subgroup and the initial weights of the public are input into the deep reinforcement learning network for adjustment. The weights of the expert subgroup and the public are obtained, including:
[0122] The minimization of time constraints and social losses is considered as a Markov decision process. The Markov decision process consists of three parts: state space, action space, and reward function. The three parts of the Markov decision process are constructed as follows:
[0123] State space: Assuming that there are N decision makers (public and expert subgroups) in a large decision-making group to discuss emergency decision-making problems, the state space is expressed as , Represents decision makers The opinion range is [0,1];
[0124] Action space: The degree to which individuals adjust their opinions constitutes the action space , For decision makers the degree of adjustment to the opinion;
[0125] Reward function: The reward function is used to provide feedback for opinion correction behavior, thereby taking better actions. The reward in the opinion fusion environment should include the following conditions:
[0126] 1) Since experts have professional knowledge, it is generally recognized that their opinions are more reliable than those of the general public;
[0127] 2) Each state must ensure a high level of reliability;
[0128] 3) Sufficient acceptability between experts and the public should also be considered, that is, the change in collective opinion should be as small as possible to control the weight change of each expert subgroup, that is, the speed of weight adjustment;
[0129] Therefore, the reward function is calculated as follows:
[0130]
[0131] in, Indicates the total number of steps in each training round;
[0132] The adjusted weight is expressed as , represents the weight of the expert subgroup, Indicates the weight of the public.
[0133] In this embodiment of the present invention, in order to ensure the rationality of the decision-making results, it is necessary to reach a group consensus to ensure that the selection of the optimal solution is determined by the decision-making group. Therefore, the group consensus measure is calculated by the weight of the expert subgroup and the weight of the public, and the expression is:
[0134]
[0135] in, represents the group consensus measure, represents the weight of the expert subgroup, Indicates the weight of the public. represents the similarity of opinions of expert subgroups, Indicates the similarity of public opinions.
[0136] The consensus threshold given in the embodiment of the present invention is , when the consensus measure When the expert subgroup and the general public reach a group consensus.
[0137] In order to make all expert subgroups reach a consensus, it is assumed that is the initial opinion of the decision-making group, where , the objective function is constructed as:
[0138]
[0139] Assumptions Expert subgroup The weight vector of and , then the consensus feedback planning equation is constructed as follows:
[0140] Model 1
[0141]
[0142]
[0143] Assumptions Expert subgroup The initial opinion is expressed as the initial linear uncertain preference value, where ;
[0144] Model 1 can be equivalently transformed into Model 2
[0145]
[0146]
[0147] This model 2 takes into account both minimal information loss and trust interaction, and on this basis reaches a group consensus to ensure the rationality of the decision-making results.
[0148] Specifically, the initial overall opinion of the decision-making group is adjusted to obtain the expression of the tendency opinion of the decision-making group:
[0149]
[0150] in, Indicates the tendency opinions of a large decision-making group, expressed as linear uncertain preference values, represents the feedback mechanism parameters, and .
[0151] Specifically, step 6 includes:
[0152] The expectation function is constructed based on the tendency opinions of the decision-making group, and the expression is:
[0153]
[0154] in, Representing variables The expected value of Indicates alternatives The reliability, Representing variables The inverse uncertainty distribution of represents the minimum value of the linear uncertain preference value, represents the maximum value of linear uncertain preference value;
[0155] Finally, by sorting the expected values under different alternative plans, the alternative plan with the highest expected value is taken as the optimal emergency decision-making plan to assist in the emergency response to emergencies.
[0156] The embodiment of the present invention mines and analyzes the obtained comment texts of the public when an emergency occurs, obtains decision attributes, attribute weights, and the initial opinions of the public, inputs the improved K-means clustering algorithm to cluster the initial opinions of the public, obtains the initial opinions of the expert subgroups, and derives the weights of the initial opinions of the expert subgroups to obtain the initial weights of the expert subgroups and the public; calculates the initial overall opinions of the decision-making group, the reliability of the public, and the acceptability of the experts based on the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups; takes the time pressure of the emergency response and the minimization of social losses as the reinforcement learning goals, inputs the initial weights of the expert subgroups and the public into the deep reinforcement learning network for adjustment, and obtains the weights of the expert subgroups and the public; through the weights of the expert subgroups and the public The group consensus measure is calculated. When the group consensus measure is greater than the soft consensus threshold and the expert subgroup and the public reach a group consensus, the initial overall opinion of the large decision-making group is adjusted to obtain the tendency opinion of the large decision-making group for constructing an expectation function to calculate the expected value of each alternative plan for responding to emergencies, and the alternative plan with the highest expected value is used as the optimal emergency decision plan. Compared with the existing technology, the embodiment of the present invention mines and analyzes the comment text of the public to obtain decision attributes, attribute weights, and the initial opinions of the public, which can solve the problem that decision information is difficult to be uniformly quantified in the emergency decision-making process. In the emergency decision-making process, the weights of the public and the expert subgroup are balanced through the training weight mechanism, which promotes the effective interaction between the two parties in the decision-making process. It can not only coordinate the opinions of experts and the public, but also ensure the effective participation of multiple views in decision-making, thereby improving the quality of decision-making.
[0157] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for emergency decision-making based on deep reinforcement learning, characterized in that: include: Step 1: Obtain the public's comments and multiple alternative plans for dealing with emergencies when they occur; Step 2: Mining the review text using the TF-IDF topic mining algorithm to obtain decision attributes and attribute weights, and performing sentiment analysis on the review text to obtain sentiment analysis results. The sentiment analysis results are converted into initial linear uncertain preference values as the initial opinions of the public. The conversion formula for converting the sentiment analysis results into initial linear uncertain preference values is: in, represents the initial linear uncertain preference value, that is, the initial opinion of the public, Represents a conversion operation, Indicates the The sentiment value of the group sample, represents the minimum sentiment value, Indicates the maximum sentiment value, Indicates the public's response to the review text About The positive sentiment value of each decision attribute, Indicates the public's response to the review text About The negative sentiment value of the decision attribute, Indicates the public's response to the review text About The hesitation sentiment value of each decision attribute; Step 3: Improve the k-means clustering algorithm. The improvement process is as follows: The similarity definition is given based on the distance between uncertain variables; Calculating similarity values between expert preference opinions based on the similarity definition; Sort the expert preference opinions in descending order of similarity values, and use the sorted expert preference opinions as the initial cluster center of each subgroup. One expert preference opinion corresponds to the initial cluster center of one subgroup. Inputting the decision attributes, the attribute weights, and the initial opinions of the public into an improved K-means clustering algorithm to cluster the initial opinions of the public to obtain the initial opinions of the expert subgroups, and inputting the initial opinions of the expert subgroups into a weight derivation algorithm to perform weight derivation to obtain the initial weights of the expert subgroups and the initial weights of the public; Step 4: Calculate the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts based on the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups. Based on the initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts, and with the time pressure of the emergency response and the minimization of social losses as the reinforcement learning goals, input the initial weights of the expert subgroups and the initial weights of the public into the deep reinforcement learning network for adjustment to obtain the weights of the expert subgroups and the weights of the public. Step 5: Calculate the group consensus measure using the weights of the expert subgroup and the weights of the public. When the group consensus measure is greater than the soft consensus threshold and the expert subgroup and the public reach a group consensus, adjust the initial overall opinion of the decision-making group to obtain the tendency opinion of the decision-making group. Step 6: Calculate the expected value of each alternative plan for dealing with the emergency using the expected function constructed based on the tendency opinions of the large decision-making group, and take the alternative plan with the highest expected value as the optimal emergency decision plan.
2. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 1, characterized in that: The step 2 includes: Using the TF-IDF topic mining algorithm to mine frequently appearing feature words in the review text, and determine the decision attributes; Calculating attribute weights by obtaining the frequencies of frequently appearing feature words and the proportion of frequently appearing feature words to all words in the review text; A sentiment analysis is performed on the review text to obtain a sentiment analysis result, and the sentiment analysis result is converted into an initial linear uncertain preference value as an initial opinion of the public.
3. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 2, characterized in that: The calculation expression of the attribute weight is: in, represents the attribute weight, Indicates frequently appearing feature words frequency, Indicates frequently appearing feature words The proportion of all words in the review text, Characteristic words quantity, Indicates the number of words in the review text.
4. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 1, characterized in that: The initial opinions of the expert subgroup are input into the weight derivation algorithm to perform weight derivation, thereby obtaining the initial weights of the expert subgroup and the initial weights of the public, including: By formula Calculate the opinion similarity of each expert subgroup, where Indicates the The similarity of opinions of expert subgroups, Indicates the Experts in the expert subgroup Opinions, Indicates the Experts in the expert subgroup Opinions, represents the number of experts in the expert subgroup; The comprehensive similarity of each expert subgroup is obtained by the opinion similarity of each expert subgroup and the number of experts in the expert subgroup. The expression is: in, Indicates the The comprehensive similarity of expert subgroups, represents the T-norm function, represents the number of expert subgroups, represents the number of experts in the decision-making group; Through the The comprehensive similarity of expert subgroups Normalization is performed to obtain the initial weights of the expert subgroup and the initial weights of the general public, which are expressed as follows: in, represents the initial weight of the expert subgroup, Represents the initial weight of the public.
5. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 4, characterized in that: The initial overall opinion of the decision-making group, the reliability of the public, and the acceptability of the experts are calculated based on the initial weights of the expert subgroups, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroups, including: The initial overall opinion of the decision-making group is calculated based on the initial weights of the expert subgroup, the initial weights of the public, the initial opinions of the public, and the initial opinions of the expert subgroup. The calculation expression is: in, represents the initial overall opinion of the decision-making group, represents the initial opinions of the expert subgroup; The reliability of the public is determined based on the initial weight of the expert subgroup, the initial weight of the public, and the similarity of opinions between the public and the expert subgroup. The expression is: in, Indicates the similarity of opinions between the general public and expert subgroups; The acceptability of the expert subgroup is calculated based on the initial overall opinion of the decision-making group and the initial opinion of the expert subgroup. The calculation expression is: in, represents the acceptability of the expert subgroup, represents the distance function.
6. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 5, characterized in that: The expression for calculating the group consensus measure by the weight of the expert subgroup and the weight of the public is: in, represents the group consensus measure, represents the weight of the expert subgroup, Indicates the weight of the public. Indicates the similarity of public opinions.
7. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 6, characterized in that: The initial overall opinion of the decision-making group is adjusted to obtain the expression of the tendency opinion of the decision-making group: in, Indicates the tendency opinions of a large decision-making group, expressed as linear uncertain preference values, Represents the feedback mechanism parameters.
8. The emergency decision-making method for emergencies based on deep reinforcement learning according to claim 7, characterized in that: The expression of the expected function is: in, Representing variables The expected value of Indicates alternatives The reliability, Representing variables The inverse uncertainty distribution of represents the minimum value of the linear uncertain preference value, Indicates the maximum value of the linear uncertain preference value.
Citation Information
Patent Citations
Large-group satellite emergency scheme decision-making method in social network environment
CN114841498A
Apparatus and a method for the generation of provider data
US12014428B1