New energy material safety evaluation method and system based on multi-agent reinforcement learning, and electronic device

By constructing a data set of new energy materials through multi-agent reinforcement learning methods, and utilizing multi-agent action selection and reward functions to calculate the centroid of clusters, the difficult problem of safety prediction of new energy materials is solved, and the correlation analysis between hot spot evolution and combustion and explosion is realized, providing a more accurate safety assessment.

WO2025194706A1PCT designated stage Publication Date: 2025-09-25TAIYUAN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
PCT/CN2024/117199
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2024-09-05
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing technologies lack a systematic understanding of the evolution of hot spots and the main influencing factors in new energy batteries. Machine learning methods for data structure and data features need to be improved, making it difficult to effectively predict the safety of new energy materials.

Method used

A multi-agent reinforcement learning method is used to construct a dataset of new energy materials. Through the multi-agent action selection strategy and reward function, the sample silhouette coefficient, cluster silhouette coefficient and total silhouette coefficient are calculated, the good cluster subset is determined, and the cluster centroid is calculated to realize the safety assessment of new energy materials.

Benefits of technology

It realizes the safety assessment of new energy materials, can predict the relationship between hot spot evolution and combustion and explosion, and provides a more accurate safety assessment model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024117199_25092025_PF_FP_ABST
    Figure CN2024117199_25092025_PF_FP_ABST
Patent Text Reader

Abstract

A new energy material safety evaluation method and system based on multi-agent reinforcement learning, and an electronic device, relating to the field of data mining and analysis for the safety of new energy materials against heavy object impact and thermal stimulation. The new energy material safety evaluation method based on multi-agent reinforcement learning comprises: using high-throughput computing to acquire data sets; providing a safety evaluation threshold on the basis of a hotspot formation and evolution process; labeling each sample in an initial data set with a safety label on the basis of the threshold to obtain a training data set; establishing a mapping relationship between the training data set and multi-agents; using an agent action selection policy to select an action having the maximum reward value to obtain a current cluster distribution result of the multi-agents; calculating a sample silhouette coefficient, a cluster silhouette coefficient, and an overall silhouette coefficient; using a reward function to calculate reward values of the multi-agents, and obtaining cluster centroids when the number of iterations satisfies a preset threshold; and determining a new energy material safety evaluation result on the basis of the cluster centroids.
Need to check novelty before this filing date? Find Prior Art

Description

A new energy material safety assessment method, system and electronic equipment based on multi-agent reinforcement learning

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 22, 2024, with application number 202410329954.0 and invention name “A method, system and electronic equipment for safety assessment of new energy materials based on multi-agent reinforcement learning”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the technical field of data mining and analysis of the safety of new energy materials against heavy object impact and thermal stimulation, and in particular to a new energy material safety assessment method, system and electronic equipment based on multi-agent reinforcement learning. Background Art

[0003] The evolution of various defects and hotspots in energetic materials is known to be extremely complex, and their relationship to energetic materials lacks quantitative understanding and theoretical description. With the development of computational science and advancements in data mining, modeling and high-throughput computation based on hotspot evolution processes have become new approaches to understanding the safety of energetic materials. Simultaneously, the latest machine learning methods are expected to establish the relationship between these fundamental evolutionary processes and material safety, thereby developing predictive safety assessment models.

[0004] Existing solutions mainly discuss the causes of accidents such as battery shell cracking and chemical leakage from the perspective of factors affecting battery aging, such as battery capacity, power, internal resistance, capacitance, and number of cycles, or discuss battery safety assessment methods from the perspective of the correlation between factors such as cycle time, average discharge current / voltage, and battery temperature and battery health; there are also literatures that conduct research on the selection and improvement of electrolytes to improve power density and ensure battery safety. For example, [1] Hu Jiangtao, Zheng Jiaxin, Pan Feng. Research progress on the correlation between structure and performance of lithium iron phosphate cathode materials for lithium batteries [J]. Acta Physico-Chimica Sinica, 2019, 35(04): 22-31. DOI: 10.3866 / PKU.WHXB201805102.

[0005] In recent years, there have been related studies on the development of data-driven new energy battery safety assessment methods using high-throughput computing and machine learning technologies. For example, in the literature [2] Yuan Jun, Artificial Intelligence Prediction of Lithium-ion Battery Health Status and Remaining Service Life, University of Electronic Science and Technology of China [D], 2023, a time series-based recurrent neural network method was proposed to predict the battery's SOH and available capacity, and a two-layer integrated prediction model (Stacking Regression, SR) was established. Based on the same health factor, the SR model was used to perform health status assessment and life prediction on batteries under different training data, and the influence of discharge strategy and experimental temperature on the SR model was explored. In the literature [3] Mu Qiuqian, Research on Data-driven Lithium-ion Battery Remaining Service Life Prediction Method, Chang'an University, 2021, a GRU-CNN hybrid prediction model based on orthogonal parameter optimization was proposed. The optimized hybrid model was used to predict the remaining service life of the battery, and the battery health status and remaining service life prediction results of the target vehicle in the future at a practical significance were obtained.

[0006] Existing research on the evolution of hotspots and the main contributing factors in new energy batteries lacks a systematic understanding and the necessary foundational data. Existing machine learning methods targeting the data structure and characteristics of these hotspots need further improvement. Therefore, a method based on data distribution analysis is urgently needed to predict the safety of new energy materials based on data.

[0007] Summary of the Invention

[0008] In order to solve the above problems, the purpose of this application is to provide a new energy material safety assessment method, system and electronic equipment based on multi-agent reinforcement learning.

[0009] To achieve the above objectives, this application provides the following technical solutions:

[0010] A new energy material safety assessment method based on multi-agent reinforcement learning, the new energy material safety assessment method comprising:

[0011] Construct a data set for various new energy materials; the data set includes the material property factors and external loads of the new energy materials, as well as the corresponding physical quantity field evolution, time domain data, and safety annotation data; the physical quantity field evolution is determined based on the formation and evolution process of hotspots; the time domain data is obtained by numerical calculation based on the material property factors and external loads of the new energy materials; and the safety annotation data is determined based on the results of physical quantity and time domain data;

[0012] Using the datasets of various new energy materials as agents, a multi-agent action selection strategy is applied to select the action with the maximum reward value to obtain the current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an agent action correction operation;

[0013] Calculate the sample silhouette coefficient, cluster silhouette coefficient, and total silhouette coefficient based on the current cluster distribution results of the multi-agent;

[0014] Determining the agents belonging to the good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient;

[0015] Applying a reward function to calculate a reward value for the multi-agent, and updating a current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to a good cluster subset and agents not belonging to a good cluster subset;

[0016] When the number of iterations meets the preset number threshold, the final cluster distribution result of the multi-agent is obtained, and the cluster centroid is calculated based on the final cluster distribution result of the multi-agent;

[0017] The safety assessment result of the new energy material is determined according to the centroid of the cluster.

[0018] A new energy material safety assessment system based on multi-agent reinforcement learning, applying the above-mentioned new energy material safety assessment method based on multi-agent reinforcement learning, the new energy material safety assessment system includes:

[0019] A construction module is used to construct a data set for various new energy materials; the data set includes material property factors and external loads of the new energy materials, as well as corresponding physical quantity field evolution, time domain data, and safety annotation data; the physical quantity field evolution is determined based on the formation and evolution process of hot spots; the time domain data is obtained by numerical calculation based on the material property factors and external loads of the new energy materials; and the safety annotation data is determined based on the results of physical quantity and time domain data;

[0020] A selection module is configured to use the various new energy materials as agents and apply a multi-agent action selection strategy to select an action with the maximum reward value, thereby obtaining a current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an agent action correction operation;

[0021] A first calculation module is used to calculate the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient according to the current cluster distribution result of the multi-agent;

[0022] A classification module, configured to determine agents belonging to a good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient;

[0023] a second calculation module, configured to apply a reward function to calculate a reward value for the multi-agent, and update a current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to a good cluster subset and agents not belonging to a good cluster subset;

[0024] A cluster centroid calculation module is used to obtain the final cluster distribution result of the multi-agent when the number of iterations meets a preset number threshold, and calculate the cluster centroid based on the final cluster distribution result of the multi-agent;

[0025] The safety assessment module is used to determine the safety assessment result of the new energy material according to the centroid of the cluster.

[0026] An electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the above-mentioned new energy material safety assessment method based on multi-agent reinforcement learning.

[0027] Optionally, the memory is a readable storage medium.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] This application first uses high-throughput computing to obtain a data set, gives a safety assessment threshold based on the hotspot formation and evolution process, and annotates each sample in the initial data set with a safety label according to the threshold to obtain a training data set; establishes a mapping relationship between the training data set and multiple agents, applies the agent action selection strategy to select the action with the largest reward value, and obtains the current cluster distribution result of the multiple agents; calculates the sample silhouette coefficient, cluster silhouette coefficient and total silhouette coefficient, applies the reward function, and calculates the reward value of the multiple agents. When the number of iterations meets the preset number threshold, the final cluster distribution result of the multiple agents is obtained, and the cluster centroid is obtained based on this; and determines the safety assessment result of the new energy material based on the cluster centroid.

[0030] Figures in the specification

[0031] The present application will be further described below with reference to the accompanying drawings:

[0032] FIG1 is a flowchart of the good cluster subset merging of the present application;

[0033] Figure 2 is a flow chart of the agent action selection strategy of this application;

[0034] Figure 3 is a flow chart of the new energy material safety assessment method based on multi-agent reinforcement learning in this application. DETAILED DESCRIPTION

[0035] The following is a detailed description of the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments; based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work shall fall within the scope of protection of this application.

[0036] The purpose of this application is to provide a new type of new energy material safety assessment and analysis method, system and electronic equipment based on multi-agent reinforcement learning, which can be used for data-driven new energy material safety assessment.

[0037] While existing technologies primarily discuss factors influencing battery safety from the perspective of battery aging, this application focuses on assessing and predicting the safety of lithium iron phosphate cathode materials from the perspective of the more destructive safety issues of combustion and explosion. High-throughput computing is used to generate a data set, and threshold criteria for combustion and explosion occurrence are derived based on the evolution of hotspots. This application can be used to analyze data related to hotspot evolution and combustion and explosion in new energy batteries, representing a novel safety assessment technology for new energy materials.

[0038] The main problem addressed by this application is to apply multi-agent reinforcement learning to cluster distribution analysis, transform the cluster distribution problem into a sequential decision-making problem, and construct a new energy material safety assessment method.

[0039] Explanation of terms:

[0040] Multi-agent reinforcement learning: Multi-agent reinforcement learning is a branch of reinforcement learning that involves multiple agents interacting and collaborating to solve complex tasks. In multi-agent reinforcement learning, each agent has its own observations, actions, and rewards, and needs to learn how to maximize its cumulative reward by interacting with other agents.

[0041] Silhouette coefficient: The silhouette coefficient is a clustering quality evaluation metric that objectively reflects the clarity of each cluster's outline. If the silhouette coefficient is close to 1, it means that the distance between the data point and its own cluster is much smaller than the distance to other clusters, indicating that the clustering effect is very good. If the silhouette coefficient is close to 0, it means that the distance between the data point and its own cluster is roughly equal to the distance to other clusters, indicating that the data point may be on or very close to the cluster boundary. If the silhouette coefficient is close to -1, it means that the distance between the data point and its own cluster is much smaller than the distance to other clusters, indicating that the clustering effect may not be very good.

[0042] Good cluster subset: A good cluster subset is a structure defined in this application. It uses the sample silhouette coefficient as a judgment condition and can accurately estimate the reward value obtained by the agent from the cluster environment after taking an action, thereby overcoming the non-stationarity problem of multi-agent reinforcement learning.

[0043] Agent Action Correction: The purpose of agent action correction is to introduce corrective actions to the agents during the final iteration, aiming to maximize the objective function and achieve the best clustering results. The implementation of agent action correction is as follows: When the method enters the final iteration, each agent calculates the sample silhouette coefficient of its own clusters and selects the action with the largest sample silhouette coefficient as the output.

[0044] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0045] Example 1

[0046] As shown in FIG3 , the present application provides a new energy material safety assessment and analysis method based on multi-agent reinforcement learning, the method comprising:

[0047] Step S1: Construct a dataset for multiple new energy materials. The dataset includes material property factors and external loads of the new energy materials, as well as corresponding physical quantity field evolution, time-domain data, and safety annotation data. The physical quantity field evolution is determined based on the formation and evolution process of hotspots. The time-domain data is obtained through numerical calculations based on the material property factors and external loads of the new energy materials. The safety annotation data is determined based on the physical quantity and time-domain data. Material property factors include sample particle size and doping data; external load data includes external load temperature and impact velocity data; and time-domain data includes time-evolution data of sample average temperature, stress, strain, energy, and other characteristics. Specifically, the dataset is data related to lithium iron phosphate cathode materials. The dataset primarily consists of experimental data and simulation data. The experimental data primarily includes data on sample internal particle size and impurity phase content; constitutive relationship data for the electrode material; and data on combustion and explosion temperature conditions. The simulation data is primarily obtained using high-throughput computing. By serializing the changes in particle size, impurity phase content, and external load data, data on the structural evolution, combustion, and explosion processes of lithium-ion battery materials are obtained. By integrating internal structural factors and external load factors, concurrent high-throughput computing is achieved; the data sampling method is optimized, the high-throughput computing results are stored in a specified data format, and a high-quality new energy battery dataset is constructed.

[0048] Step S2: Taking the data sets of various new energy materials as the agent, apply the multi-agent action selection strategy to select the action with the largest reward value to obtain the current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an agent action correction operation.

[0049] Specifically, a mapping relationship between data sets and multiple agents is established, and each piece of new energy material data will be assigned an agent to obtain the corresponding relationship between new energy material data and multiple agents.

[0050] From the first iteration to the penultimate iteration, the agent's action selection strategy adopts a greedy strategy, and the agent has a probability of 1-∈ to select the action with the largest reward value, and a probability of ∈ to select an action randomly; when the final iteration comes, the agent takes an action correction operation, and each agent calculates the sample silhouette coefficient when it belongs to different clusters, and selects the action with the largest sample silhouette coefficient as the result output.

[0051] As shown in Figure 2, the agent's action selection strategy consists of a greedy strategy and an agent action correction. In the first iteration, the greedy strategy randomly selects an action. In the second to T-1 iterations, the greedy strategy first generates a random number, ran_num. If ran_num is less than or equal to the set greed coefficient, the agent randomly selects an action. If ran_num is greater than the greed coefficient, the agent selects the action with the highest reward.

[0052] If the agent action correction is the Tth iteration, the agent will first select the action with the largest reward value, then it will traverse all agents in turn and calculate the sample silhouette coefficient of this agent when it belongs to different clusters, and select the action with the largest sample silhouette coefficient as the final action of the agent.

[0053] Furthermore, the action selection strategy is:

[0054] Among them, π i,t is the action taken by the i-th agent at the t-th iteration, ∈ is the greed coefficient; is the lth action selected by the i-th agent at the t-th iteration; k is the total number of clusters; r i,t-1 is the reward value of the i-th agent at the last (t-1) iteration.

[0055] Step S3: Calculate the sample silhouette coefficient, cluster silhouette coefficient and total silhouette coefficient based on the current cluster distribution results of the multi-agent.

[0056] Specifically, the calculation formula of the sample silhouette coefficient is:

[0057] The calculation formula of cluster silhouette coefficient is:

[0058] The calculation formula of the total silhouette coefficient is:

[0059] According to the preset conditions, determine whether the agent belongs to the good cluster subset. The formula for determining the good cluster subset is: Ψ l,t (s i )=θ i,t ,θ i,t >0ands i ∈c l ;

[0060] Among them, φ l,t is the cluster silhouette coefficient obtained at the tth iteration; π t is the action set; S is the target data set; n l is the number of sample points in the lth cluster; π t (s i ) is s i The action selected; θ i,t is the sample silhouette coefficient obtained at the tth iteration; p i,t is the inter-cluster dissimilarity of sample points; b i,t is the intra-cluster dissimilarity of sample points; Indicates s i The number of sample points in the cluster; π t (s j )=π t (s i ) means s j With s i In the same cluster; d(s i ,s j ) is sample s i With sample s j The Euclidean distance between them; k is the total number of clusters; l≠π t (s i ) represents the selected cluster and s i Belong to different clusters; Φ t is the total silhouette coefficient obtained at the tth iteration.

[0061] The specific meaning of the formula is as follows: The first formula is used to calculate the β quantile of the t-th iteration, which is used to determine the selection range of the good cluster subset. The second formula is used to determine whether the agent is a candidate for the good cluster subset (all agents with a sample silhouette coefficient greater than 0 are candidates for the good cluster subset). The third formula is used to determine whether the agent belongs to the good cluster subset. If an agent is a candidate for the good cluster subset and its ranking is in the top β according to the size of the sample silhouette coefficient, then the agent belongs to the good cluster subset.

[0062] Step S4: Determine the agents that belong to the good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient.

[0063] S4 specifically includes:

[0064] Step S41: According to preset conditions, the samples in the current cluster distribution result of the multi-agent are divided to obtain multiple good cluster subsets.

[0065] The preset conditions include a first preset condition, a second preset condition and a third preset condition.

[0066] The first preset condition is: using quantiles to determine the number of samples in the good cluster subset.

[0067] Where β is the quantile obtained at the t-th iteration; T is the maximum number of iterations, and α is the good cluster subset size adjustment factor.

[0068] The second preset condition is: determining the candidate samples of the good cluster subset according to the candidate set. l,t (s i )=θ i,t ,θ i,t >0ands i ∈c l ;

[0069] Among them, l,t is the candidate set; θ i,t is the sample silhouette coefficient obtained at the tth iteration; c l is the lth cluster; s i is the i-th sample.

[0070] The third preset condition is: for the candidate set Ψ obtained in the tth iteration l,t Arranged in descending order, the candidate samples in the first β constitute the good cluster subset Ω of the tth iteration l,t .

[0071] Among them, Ω l,t is the good cluster subset of the tth iteration; θ i,t is the sample silhouette coefficient obtained at the tth iteration; s i is the i-th sample; l,t is the candidate set; β is the quantile obtained in the tth iteration; c l is the lth cluster.

[0072] Step S42: merging the plurality of good cluster subsets according to the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient, and determining the agents belonging to the good cluster subsets.

[0073] In practical applications, the sample silhouette coefficient, cluster silhouette coefficient, and total silhouette coefficient are calculated. It is then determined whether the agent belongs to a good cluster subset. Multiple good cluster subsets are merged to obtain agents that belong to the good cluster subset.

[0074] S42 specifically includes:

[0075] Step S421: Obtain the number of good cluster subsets and a list of the number of agents in each good cluster subset.

[0076] Step S422: According to the list of the number of agents, determine the good cluster subset with the largest number of agents as the merging subject, and determine multiple good cluster subsets other than the merging subject as merging objects.

[0077] Step S423: Calculate the cluster silhouette coefficient of the good cluster subset corresponding to the current merge object and the good cluster subset corresponding to the merge subject as the same cluster, and obtain the merged cluster silhouette coefficient.

[0078] Step S424: When the pre-merger cluster silhouette coefficient corresponding to the merged object is smaller than the post-merger cluster silhouette coefficient, the current good cluster subset corresponding to the merged object is merged with the good cluster subset corresponding to the merged subject to obtain a merged good cluster subset;

[0079] Step S425: Update the merged entity according to the merged good cluster subset, and update the current good cluster subset corresponding to the merged object according to the next good cluster subset corresponding to the merged object, return to step S423 to continue execution, and use the merged good cluster subset as the current cluster distribution result of the multi-agent.

[0080] As shown in Figure 1, the specific steps for merging good cluster subsets include: first, initialization parameter settings are performed; second, the method obtains the number of GCSs and the number of agents in each GCS; third, the GCS with the most agents is defined as the merging subject, and the other GCSs are defined as merging objects; fourth, all merging objects are traversed in turn, and the merging subject and the merging object are merged (the actions of the agents in the merging object are modified to the actions corresponding to the merging subject agents, so that they belong to the same good cluster subset); fifth, determine whether the cluster silhouette coefficient of the current merging subject has changed. If the cluster silhouette coefficient increases, the merger is valid, and the action of the merging object is updated to make it the same as the merging subject. If the cluster silhouette coefficient remains unchanged or decreases, the merger fails, and the merging object is reset to its original action.

[0081] Step S5: Apply the reward function to calculate the reward value of the multi-agent, and update the current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to the good cluster subset and agents not belonging to the good cluster subset.

[0082] The calculation formula for the reward value is:

[0083] When the agent belongs to the good cluster subset, the current action will be given a reward value of "1", while other actions will be given a reward value of "0". The reward function is:

[0084] When the agent does not belong to the good cluster subset, all actions will inherit the reward value of the previous iteration. The reward function is:

[0085] in, is the reward value of the tth iteration; agent i is the i-th agent; Ω l,t is the lth good cluster subset at the tth iteration; is the lth action selected by the i-th agent at the t-th iteration; e t is the cluster distribution environment of the current agent in the tth iteration; Ω anyone,t is the set of candidates for the t-th good cluster subset; e t-1 is the cluster distribution environment of the current agent at the t-1th iteration; is the lth action selected by the i-th agent at the t-1th iteration; Ω l1,t is the l1th good cluster subset at the tth iteration.

[0086] Step S6: When the number of iterations meets the preset threshold, the final multi-agent cluster distribution is obtained, and the cluster centroid is calculated based on the final multi-agent cluster distribution. The centroid is calculated by averaging the data of the same dimension for agents belonging to the same cluster. The average value of all dimensions represents the location of the centroid.

[0087] Step S7: Determine the safety assessment result of the new energy material according to the cluster centroid.

[0088] The cluster centroid is compared to the safety assessment threshold. If the cluster centroid is greater than the safety threshold, all multi-agents in that cluster are deemed unsafe, and the safety assessment result is obtained based on this. For the test data, the Euclidean distance between the cluster centroid and each cluster centroid is first calculated. The cluster corresponding to the closest centroid is the cluster label for that data. The safety assessment result of the test sample is then obtained based on the multi-agent safety assessment result of that cluster.

[0089] In practical applications, high-throughput computing is used to obtain corresponding data for a new energy material to be safety assessed. The new energy material to be safety assessed belongs to a type of new energy material included in a new energy material data set. For example, the new energy material to be safety assessed is a lithium iron phosphate cathode material. The corresponding data refers to material attribute factors and external loads. After the corresponding data of the new energy material to be safety assessed is processed through steps S2 to S7, a safety assessment result for the new energy material to be safety assessed is obtained.

[0090] This application has the following advantages:

[0091] 1. This application constructs a new energy material safety assessment and analysis method based on multi-agent reinforcement learning, realizing the safety assessment of new energy materials.

[0092] 2. This application transforms the cluster distribution problem into a sequential decision problem, establishes a mapping relationship between the training data set and multi-agents, and constructs a multi-agent cluster distribution model.

[0093] 3. This application defines a good cluster subset structure defined by the sample silhouette coefficient. This structure helps to accurately evaluate the rewards obtained by multi-agents.

[0094] 4. To accelerate convergence, this application designs a new agent action correction operation. Multi-agent reinforcement learning requires extensive exploration of the environment to find the optimal set of actions, and agent action correction operations can optimize this learning process.

[0095] Example 2

[0096] In order to execute the method corresponding to the above embodiment 1 and achieve the corresponding functions and technical effects, a new energy material safety assessment system based on multi-agent reinforcement learning is provided below. The new energy material safety assessment system includes:

[0097] A construction module is used to construct a data set for various new energy materials; the data set includes the material property factors and external loads of the new energy materials and the corresponding physical quantity field evolution, time domain data, and safety annotation data; the physical quantity field evolution is determined based on the hotspot formation and evolution process; the time domain data is obtained by applying numerical calculations based on the material property factors and external loads of the new energy materials; and the safety annotation data is determined based on the results of physical quantity and time domain data.

[0098] A selection module is used to use various new energy materials as intelligent agents, apply a multi-agent action selection strategy to select the action with the largest reward value, and obtain the current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an intelligent agent action correction operation.

[0099] The first calculation module is used to calculate the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient according to the current cluster distribution results of the multi-agent.

[0100] The classification module is used to determine the intelligent agents belonging to the good cluster subset according to the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient.

[0101] The second calculation module is used to apply the reward function to calculate the reward value of the multi-agent and update the current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to the good cluster subset and agents not belonging to the good cluster subset.

[0102] The cluster centroid calculation module is used to obtain the final cluster distribution result of the multi-agent when the number of iterations meets the preset number threshold, and calculate the cluster centroid based on the final cluster distribution result of the multi-agent.

[0103] The safety assessment module is used to determine the safety assessment result of the new energy material according to the centroid of the cluster.

[0104] Example 3

[0105] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the new energy material safety assessment method based on multi-agent reinforcement learning of embodiment one.

[0106] Optionally, the above-mentioned electronic device may be a server.

[0107] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the data analysis method based on multi-agent reinforcement learning in embodiment one.

[0108] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0109] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A new energy material safety assessment method based on multi-agent reinforcement learning, characterized in that: The new energy material safety assessment method includes: Construct a data set for various new energy materials; the data set includes the material property factors and external loads of the new energy materials, as well as the corresponding physical quantity field evolution, time domain data, and safety annotation data; the physical quantity field evolution is determined based on the formation and evolution process of hotspots; the time domain data is obtained by numerical calculation based on the material property factors and external loads of the new energy materials; and the safety annotation data is determined based on the results of physical quantity and time domain data; Using the datasets of various new energy materials as agents, a multi-agent action selection strategy is applied to select the action with the maximum reward value to obtain the current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an agent action correction operation; Calculate the sample silhouette coefficient, cluster silhouette coefficient, and total silhouette coefficient based on the current cluster distribution results of the multi-agent; Determining the agents belonging to the good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient; Applying a reward function to calculate a reward value for the multi-agent, and updating a current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to a good cluster subset and agents not belonging to a good cluster subset; When the number of iterations meets the preset number threshold, the final cluster distribution result of the multi-agent is obtained, and the cluster centroid is calculated based on the final cluster distribution result of the multi-agent; The safety assessment result of the new energy material is determined according to the centroid of the cluster.

2. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The calculation formula of the sample silhouette coefficient is: The calculation formula of the cluster silhouette coefficient is: The calculation formula of the total silhouette coefficient is: Among them, φ l,t is the cluster silhouette coefficient obtained at the tth iteration; π t is the action set; S is the target data set; n l is the number of sample points in the lth cluster; π t (s i ) is s i The action selected; θ i,t is the sample silhouette coefficient obtained at the tth iteration; p i,t is the inter-cluster dissimilarity of sample points; b i,t is the intra-cluster dissimilarity of sample points; Represents sample s i The number of sample points in the cluster; π t (s j )=π t (s i ) represents sample s j With sample s i In the same cluster; d(s i ,s j ) is sample s i With sample s j The Euclidean distance between them; k is the total number of clusters; l≠π t (s i ) represents the selected cluster and s i Belong to different clusters; Φ t is the total silhouette coefficient obtained at the tth iteration.

3. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 1 is characterized in that: Determining the agents belonging to the good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient, specifically includes: According to preset conditions, the samples in the current cluster distribution results of the multi-agent are divided to obtain multiple good cluster subsets; The plurality of good cluster subsets are merged according to the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient to determine the intelligent agents belonging to the good cluster subsets.

4. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 3 is characterized in that: The preset conditions include a first preset condition, a second preset condition and a third preset condition; The first preset condition is: using quantiles to determine the number of samples in the good cluster subset; Where β is the quantile obtained at the tth iteration; T is the maximum number of iterations, and α is the good cluster subset size adjustment factor; The second preset condition is: determining candidate samples of the good cluster subset based on the candidate set; P l,t (s i )=θ i,t ,i i,t >0ands i ∈c l ; Among them, l,t is the candidate set; θ i,t is the sample silhouette coefficient obtained at the tth iteration; c l is the lth cluster; s i is the i-th sample; The third preset condition is: for the candidate set Ψ obtained in the tth iteration l,t Arranged in descending order, the candidate samples in the first β constitute the good cluster subset Ω of the tth iteration l,t ; Among them, Ω l,t is the good cluster subset of the tth iteration; θ i,t is the sample obtained at the tth iteration The silhouette coefficient; s i is the i-th sample; l,t is the candidate set; β is the quantile obtained in the tth iteration; c l is the lth cluster.

5. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 3 is characterized in that: Merging the plurality of good cluster subsets according to the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient to determine agents belonging to the good cluster subsets specifically includes: Obtaining the number of good cluster subsets and a list of the number of agents in each good cluster subset; According to the list of the number of agents, determining the good cluster subset with the largest number of agents as the merging subject, and determining a plurality of the good cluster subsets other than the merging subject as merging objects; Calculating the cluster silhouette coefficient of the good cluster subset corresponding to the current merge object and the good cluster subset corresponding to the merge subject as the same cluster, to obtain the merged cluster silhouette coefficient; When the cluster silhouette coefficient before merging corresponding to the merged object is less than the cluster silhouette coefficient after merging, the current good cluster subset corresponding to the merged object is merged with the good cluster subset corresponding to the merged subject to obtain a merged good cluster subset; according to the merged good cluster subset, the merged subject is updated, and according to the next good cluster subset corresponding to the merged object, the current good cluster subset corresponding to the merged object is updated, and the step of "calculating the cluster silhouette coefficient of the current good cluster subset corresponding to the merged object and the good cluster subset corresponding to the merged subject as the same cluster to obtain the merged cluster silhouette coefficient" is returned to continue execution, and the merged good cluster subset is used as an intelligent body belonging to the good cluster subset.

6. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 1 is characterized in that: When the agent belongs to the good cluster subset, the reward function is: When the agent does not belong to the good cluster subset, the reward function is: in, is the reward value of the tth iteration; agent i is the i-th agent; Ω l,t For the tth The lth good cluster subset of the iteration; is the lth action selected by the i-th agent at the t-th iteration; e t is the cluster distribution environment of the current agent in the tth iteration; Ω anyone,t is the set of candidates for the t-th good cluster subset; e t-1 is the cluster distribution environment of the current agent at the t-1th iteration; is the lth action selected by the i-th agent at the t-1th iteration; is the l1th good cluster subset at the tth iteration.

7. The new energy material safety assessment method based on multi-agent reinforcement learning according to claim 1 is characterized in that: The action selection strategy is: Among them, π i,t is the action taken by the i-th agent at the t-th iteration, ∈ is the greed coefficient; is the lth action selected by the i-th agent at the t-th iteration; k is the total number of clusters; r i,t-1 is the reward value of the i-th agent in the last iteration, which is t-1.

8. A new energy material safety assessment system based on multi-agent reinforcement learning, characterized in that: The new energy material safety assessment system includes: A construction module is used to construct a data set for various new energy materials; the data set includes material property factors and external loads of the new energy materials, as well as corresponding physical quantity field evolution, time domain data, and safety annotation data; the physical quantity field evolution is determined based on the formation and evolution process of hot spots; the time domain data is obtained by numerical calculation based on the material property factors and external loads of the new energy materials; and the safety annotation data is determined based on the results of physical quantity and time domain data; A selection module is configured to use the various new energy materials as agents and apply a multi-agent action selection strategy to select an action with the maximum reward value, thereby obtaining a current cluster distribution result of the multi-agent; the multi-agent action selection strategy includes a greedy strategy and an agent action correction operation; A first calculation module is used to calculate the sample silhouette coefficient, the cluster silhouette coefficient and the total silhouette coefficient according to the current cluster distribution result of the multi-agent; A classification module, configured to determine agents belonging to a good cluster subset based on the sample silhouette coefficient, the cluster silhouette coefficient, and the total silhouette coefficient; a second calculation module, configured to apply a reward function to calculate a reward value for the multi-agent, and update a current cluster distribution result of the multi-agent according to the reward value; the multi-agent includes agents belonging to a good cluster subset and agents not belonging to a good cluster subset; A cluster centroid calculation module is used to obtain the final cluster distribution result of the multi-agent when the number of iterations meets a preset number threshold, and calculate the cluster centroid based on the final cluster distribution result of the multi-agent; The safety assessment module is used to determine the safety assessment result of the new energy material according to the centroid of the cluster.

9. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the new energy material safety assessment method based on multi-agent reinforcement learning according to any one of claims 1 to 7.

10. The electronic device according to claim 9, characterized in that: The memory is a readable storage medium.

Citation Information

Patent Citations

  • Reinforcement learning multi-agent communication and decision method

    CN108921298A

  • Cooperative multi-agent reinforcement learning method based on self-adaptive reward allocation

    CN113780576A

  • Multi-agent learning method based on average field, storage medium and electronic equipment

    CN113887708A

  • Power distribution network multi-region cooperative reactive power optimization method based on multi-agent reinforcement learning

    CN115483703A

  • New energy material safety assessment method and system based on multi-agent reinforcement learning, and electronic equipment

    CN118155773A

Cited By

  • Scientific and technological intelligence agent training method and system based on reinforcement learning

    CN121328607A

  • Time sequence knowledge graph reasoning method and system based on fuzzy clustering and reinforcement learning

    CN121413785A

  • Streaming computing task scheduling method based on space-time perception and multi-agent reinforcement learning

    CN121479359A

  • Mine safety interlocking control method, device and equipment based on reinforcement learning and medium

    CN122172603A