Large language model standardization processing method for reservoir group scheduling rule
Through large language model and decision tree algorithm, standardized processing and decision-making complexity evaluation of reservoir group scheduling rules is solved, and the defects of natural language description of traditional reservoir scheduling rules are achieved, and more efficient automated processing and application are achieved.
Patent Information
- Application Number
- CN202510103134.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional reservoir scheduling rules are described in natural language, and there are problems of different formats, different terms, and lack of clear structure, making it difficult to achieve automated processing and application.
A large language model is used to analyze the decision-making characteristics and standardization of reservoir group scheduling rules, introduce a decision tree algorithm, extract the reservoir group scheduling rules, and propose a method to evaluate the decision-making complexity of reservoir group scheduling rules.
Through standardized processing, the standardization and digitization of reservoir group scheduling rules are improved, the defects of natural language description are reduced, and the automated processing capabilities and application efficiency of scheduling rules are improved.
Smart Images

Figure CN120067248A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of energy dispatching, and in particular relates to a large language model standardization processing method for reservoir group dispatching rules. Background Art
[0002] In recent years, with the development of artificial intelligence and big data technology, machine learning, deep learning and other methods have been applied to the field of reservoir operation. In particular, the powerful ability of large language model (LLM) based on machine learning technology in the field of natural language processing has provided a new way for text processing and decision feature analysis of reservoir operation rules. With its powerful semantic understanding and generation capabilities, the large language model can effectively convert the originally complex and fuzzy text description of the operation rules into structured logical expressions (such as knowledge graphs), thereby improving the standardization and operability of the operation rules. In addition, by logically presenting each link, condition judgment and decision path in the operation decision process, it provides a clear decision view for the dispatchers, which helps to quickly identify and optimize the bottlenecks and deficiencies in the decision process. In the limited literature on the use of large language models to study the operation rules, the focus is mainly on the knowledge graphing of the form of the operation rules. However, the knowledge graph can represent multi-dimensional entities and relationships, which is suitable for complex and dynamic decision problems. Its complex structure makes the reasoning and derivation process less transparent, and in actual operation, it requires high user understanding and participation. In addition, the construction of knowledge graphs usually requires a lot of manual intervention and domain knowledge, and the dynamic updating and expansion of graphs is also a challenge.
[0003] In summary, the technical problem to be solved by the present invention is that the traditional reservoir dispatching rules are described in natural language, which have problems such as inconsistent formats, different terminologies, and lack of clear structure, making it difficult to realize automatic processing and application. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a large language model standardization processing method for reservoir group dispatching rules. The decision characteristics and standardization of reservoir group dispatching rules are analyzed through a large language model, and a decision tree algorithm is introduced to extract reservoir group dispatching rules, and the decision complexity of evaluating reservoir group dispatching rules is proposed. This method can be applied to reservoir group dispatching and different river ecological protection requirements and objectives, has good adaptability and strong practicality, and can effectively reduce the problems of different formats, different terms, lack of clear structure, etc. in the reservoir dispatching rules that describe dispatching principles, objectives and constraints in natural language, and enhance the standardization and digitization of dispatching rules.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for standardizing a large language model for reservoir group scheduling rules, the steps are as follows: Step 1, construct a DT scheduling rule model based on a large language model, and obtain the text description recognition of the scheduling rule and process it into a logical form.
[0006] Step 1.1, perform local deployment of the large language model and construct the DT scheduling rules for the reservoir group: realize the local deployment of the large language model (Qian wen Large Language Model, QwenLLM) based on the Windows PowerShell platform and the open-source Ollama platform; Step 1.2, prepare the hardware and software environment. In terms of hardware, deploying the large language model requires high-performance computer equipment, including high-performance CPUs, GPUs, and large-capacity storage spaces. In terms of software, the computer equipment system version is Windows 11; Windows PowerShell is a command-line shell and scripting language integrated tool developed by Microsoft, mainly used for automated system management tasks, and version 5.0 (version 5.1.22621.4391) is adopted; Ollama is an open-source framework designed for conveniently deploying and running large language models on local machines, and the version adopted is 0.4.1.
[0007] Step 1.3, Model Selection and Deployment: Select the latest version Qwen 2.5 of QwenLLM for local deployment. Qwen 2.5 is the latest member of the QwenLLM series, showing high capabilities in the fields of programming and mathematics, and providing a series of base language models with scales ranging from 50 million to 72 billion parameters. Qwen 2.5 supports a context of up to 128K tokens and can generate text of up to 8K tokens. In addition, Qwen 2.5 has strong capabilities in instruction following, long text generation, structured data understanding, and generating structured outputs. In terms of specific version selection, considering hardware performance and computing time consumption, etc., select the Qwen 2.5:7b model with 7.62 billion parameters as the large language model for local deployment. The specific process of model deployment is as follows: ① Enter the ollama download page (https: / / ollama.com / ), select the Windows system version of ollama for download and installation. ② After installation, enter the Ollama command in Windows PowerShell to verify whether ollama is installed successfully. ③ After ollama is installed successfully, select the Qwen 2.5 model in the ollama model library (https: / / ollama.com / library), and enter ollama run qwen2.5:7b in Windows PowerShell to download, install and run the model, thus realizing the local deployment of the large language model.
[0008] Step 1.4, Verification and Instance Application: The QwenLLM adopted is a model after training and verification. Therefore, after realizing local deployment, there is no need to retrain its parameters. To ensure the accuracy of the model after local deployment, verify the accuracy and reliability of QwenLLM by comparing knowledge questions with other online large language models and accurately answer relevant questions. After realizing the local deployment of the large model, adopt the method of knowledge questions, input the joint operation rules of the reservoir group into the model, and obtain the DT operation rule library based on the large language model by asking what the key decision variables and thresholds are in the operation rules of each reservoir and standardizing them into the form of DT. After obtaining the DT operation rule library, further adopt the method of manual inspection to compare and verify each operation rule based on the large language model with the text one by one to ensure the accuracy of the DT operation rule.
[0009] Step 2, A weighted complexity evaluation model is proposed based on the DT operation rule, and the decision complexity of the operation rule of the reservoir group is evaluated by calculating the decision complexity of the operation rule through DT depth, the number of nodes, and the number of leaf nodes, etc.
[0010] Step 2.1, Analyze the decision complexity of the DT scheduling rule and confirm the decision variables based on the large language model deployed locally. The decision variables are the decision-making nodes of the DT. Their importance varies due to differences in type and position in different DT scheduling rules.
[0011] Step 2.2, Based on the DT scheduling rule, a weighted complexity evaluation model is proposed. The decision complexity of the scheduling rule is calculated through the DT depth, the number of nodes, and the number of leaf nodes, aiming to evaluate the decision complexity of the reservoir group scheduling rule. The complexity evaluation model of the scheduling rule is as follows: ; In the formula: C is the decision complexity of the scheduling rule, D is the DT depth, N is the number of DT nodes, L is the number of DT leaf nodes, ω is the weighting coefficient, taking ω 1 = ω 2 = ω 3 = 1 / 3. Based on the above definitions, the number of key decision variables in the DT scheduling rule is: N - L , and the importance of the corresponding decision variable is measured by the position of the key decision variable in the DT (the lower the position in the DT depth, the higher the importance of the decision variable).
[0012] Step 2.3, Measure the complexity of the reservoir group scheduling decision, the flood control storage capacity of the reservoir and its decision complexity, and the key decision variables of each reservoir scheduling rule in the basin control reservoir group. Differences in the main decision variables lead to different forms of the DT, that is, the number of nodes (Node), the number of leaf nodes (Leaf node), and the DT depth (Depth) are different. The depth of the DT is the number of nodes on the longest path from the root node to the deepest leaf node. Nodes include internal nodes and leaf nodes. Leaf nodes are the final nodes in the DT, representing the final decision made. Therefore, due to the decision variables in the scheduling rule and the form of the DT, there are significant differences in the decision-making process of the reservoir group scheduling. Generally, the more decision variables a reservoir group scheduling rule has and the more complex its form, the higher the complexity of its scheduling decision.
[0013] Step 3, Analyze the association between the scheduling rule and the reservoir characteristics.
[0014] Step 3.1, regarding the issue of whether the larger the flood control storage capacity of the reservoir, the higher its decision-making complexity, further use correlation calculation and principal component analysis (PCA) to determine the correlation between the DT scheduling rule and reservoir characteristics (storage capacity, inflow runoff, location, etc.). PCA first linearly transforms the p dimensional original data X p into a new coordinate system, calculates the eigenvalues and eigenvectors of the sample covariance matrix after mapping, and selects the eigenvectors a k corresponding to the first pi eigenvalues as the principal components, thereby converting the high-dimensional data set into a low-dimensional representation. The i th principal component F i is:
[0015] In the i th principal component F i , the larger the coefficient of the index X p , the greater the contribution of the index to this principal component. By comparing the contribution degrees of each index to the corresponding principal component, each principal component can be divided into different categories, thus realizing the classification of the principal components.
[0016] Step 3.2, by using the PCA dimensionality reduction method, first reduce the characteristic variables of the reservoir group to a few representative principal components, and analyze the contribution degrees of each characteristic variable to the corresponding principal components to achieve the classification analysis of the reservoir group characteristics. The characteristic variables of the reservoir group include storage capacity size, inflow runoff, spatial location, etc., and a specific index is used to completely classify the reservoir group characteristics. After reducing the dimensionality of the reservoir group characteristic variables, further map and analyze the decision-making complexity of the reservoir group scheduling rule and the principal components to determine the correlation between the decision-making complexity and the reservoir group characteristics.
[0017] A system for the standardization processing method of the large language model of the reservoir group scheduling rule adopts the standardization processing method of the large language model of the reservoir group scheduling rule described above.
[0018] A computer device includes: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, the standardization processing method of the large language model of the reservoir group scheduling rule as described above is implemented.
[0019] A computer-readable storage medium stores a computer program thereon. When the computer program is executed, it implements a method for standardizing a large language model according to the reservoir group scheduling rules described above.
[0020] The present invention can achieve the following beneficial effects: Aiming at the problem of standardizing the text description of reservoir scheduling rules, the present invention transforms the joint scheduling rules of reservoir groups into a set of general, transparent, and easy-to-operate decision tree (DT) systems for joint scheduling rules by establishing a locally deployed large language model. On this basis, the present invention deeply analyzes the key decision variables and decision complexity of DT scheduling rules and their correlations with reservoir group attributes, reveals the decisive factors of scheduling rules, and provides a scientific basis for optimizing and improving scheduling rules. The present invention helps to improve the scientificity and effectiveness of joint scheduling of reservoir groups. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present invention will be further described below in conjunction with the drawings and embodiments: Figure 1 It is a schematic diagram of the DT scheduling rule and its feature analysis process based on a large language model provided by the present invention; Figure 2 Distribution of exemplary reservoir groups Figure 3 It is a schematic diagram of the local deployment of the large language model provided by the present invention; Figure 4 It is a schematic diagram of the DT scheduling rule based on the locally deployed QwenLLM provided by the present invention; Figure 5 It is a schematic diagram of the results of reservoir scheduling rules provided by the present invention; Figure 6 It is a schematic diagram of the spatial distribution of the decision complexity of reservoir group scheduling rules provided by the present invention; Figure 7 It is a schematic diagram of the association pattern between scheduling rules and reservoir characteristics provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] As Figure 1 shown, the present invention provides a method for standardizing a large language model of reservoir group scheduling rules. Taking the joint scheduling of reservoir groups in the Yangtze River Basin (basic location information of reservoirs such as Figure 2Taking the Three Gorges Reservoir as the core, supplemented by four cascade reservoirs in the lower reaches of the Jinsha River as the backbone force, and integrating nine upstream and midstream cascade reservoir groups such as the midstream reservoir group of the Jinsha River, the Yalong River reservoir group, the Minjiang River reservoir group, the Jialing River reservoir group, the Wujiang River reservoir group, the Qingjiang River group, the "Four Rivers" group of Dongting Lake, the Hanjiang River group, and the "Five Rivers" group of Poyang Lake, which together constitute an efficient and collaborative reservoir group joint dispatching network in the Yangtze River Basin as an example, the DT dispatching rules of the large language model are established, including the following steps: Step 1: Construct a DT dispatching rule model for the Yangtze River Basin based on the large language model, and obtain the text description of the Yangtze River Basin dispatching rules and identify and process them into a logical form.
[0023] Step 1.1: Localize the deployment of the language model and construct the DT dispatching rules for the Yangtze River Basin. Realize the localization deployment of QwenLLM based on Windows PowerShell and Ollama.
[0024] Step 1.2: Prepare the hardware and software environment. In terms of hardware, deploying the large language model requires high-performance computer equipment. The computer equipment model used is Lenovo Legion Y9000P IRX8, the CPU is Intel Core i9-13900HX, the GPU is NVIDIA RTX4050, the memory is Samsung DDR5 4800MHZ 64GB, and the storage is 1TB PCle 4.0. In terms of software, the computer equipment system version is Windows 11 Home Chinese Edition (version 10.0.22631.4460); Windows PowerShell is a command-line shell and scripting language integrated tool developed by Microsoft, mainly used for automated system management tasks, and uses version 5.0 (version 5.1.22621.4391); Ollama is an open-source framework designed for conveniently deploying and running large language models on local machines, and the version used is 0.4.1.
[0025] Step 1.3, model selection and deployment, select the latest version of QwenLLM, Qwen 2.5, for localized deployment. Qwen2.5 is the latest member of the QwenLLM series, demonstrating high capabilities in programming and mathematics, and provides a series of basic language models ranging in size from 50 million to 72 billion parameters. Qwen2.5 supports contexts up to 128K tags and can generate texts of up to 8K tags. In addition, Qwen2.5 has strong capabilities in instruction following, long text generation, structured data understanding, and generating structured output. In terms of specific version selection, considering hardware performance and computational time, the Qwen 2.5:7b model with 7.62 billion parameters was selected as the large language model for local deployment. The specific process of model deployment is as follows: ① Enter the ollama download page (https: / / ollama.com / ), select the Windows version of ollama to download and install, ② After the installation is complete, enter the Ollama command in Windows PowerShell to verify whether ollama is installed successfully, ③ After ollama is installed successfully, select the Qwen 2.5 model in the ollama model library (https: / / ollama.com / library), enter ollama run qwen2.5:7b in Windows PowerShell to download, install and run the model, thereby realizing the localized deployment of the large language model, such as Figure 3 shown.
[0026] Step 1.4, verification and example application. The QwenLLM used is a trained and verified model. Therefore, after local deployment, there is no need to retrain its parameters. In order to ensure the accuracy of the model after local deployment, the accuracy and reliability of QwenLLM are verified through comparison with knowledge questions and other online large language models to accurately answer relevant questions. After the large model is deployed locally, the knowledge question-and-answer method is used to input the joint flood control and dispatching rules of the reservoir group in the Yangtze River Basin into the model. By asking what are the key decision variables and thresholds in the flood control and dispatching rules of each reservoir in the Yangtze River Basin, and standardizing them into the form of DT, a DT dispatching rule base based on the large language model is obtained. After obtaining the DT dispatching rule base, manual inspection is further used to compare and verify the dispatching rules based on the large language model with the text one by one to ensure the accuracy of the DT dispatching rules in the Yangtze River Basin. Based on the local deployment, QwenLLM standardizes the text description of the flood control and dispatching rules of the controlling reservoir group in the Yangtze River Basin into the form of DT. For example Figure 4 and Figure 5 shown. Figure 4It shows the text description of the flood control operation rules of the Three Gorges Reservoir and the process of converting the text description into DT operation rules through knowledge Q&A. Figure 5 To further standardize the operation rules converted by the large language model. After obtaining the DT operation rules based on the large language model, manual verification is carried out to ensure that the DT operation rules strictly conform to the actual reservoir group operation rules.
[0027] Step 2: A weighted complexity evaluation model is proposed based on the DT operation rules. The decision complexity of the operation rules is calculated through the DT depth, the number of nodes, and the number of leaf nodes, etc., to evaluate the decision complexity of the flood control operation rules of the reservoir group.
[0028] Step 2.1: Analyze the decision complexity of the DT operation rules and confirm the decision variables based on the DT operation rules obtained from the locally deployed large language model. The decision variables are the decision-making judgment nodes of the DT. The types and positions vary in different DT operation rules, resulting in different importance levels. For example, for the flood control operation rules of the controlling reservoir group in the Yangtze River Basin, whether it is the flood season is the primary decision variable for all reservoirs, while rainfall forecast information is an important decision variable in the flood control operation of the Three Gorges Reservoir. However, the flood control operation rules of Liyuan Reservoir in the middle reaches of the Jinsha River do not consider rainfall forecast information.
[0029] Step 2.2: A weighted complexity evaluation model is proposed based on the DT operation rules. The decision complexity of the operation rules is calculated through the DT depth, the number of nodes, and the number of leaf nodes, etc., aiming to evaluate the decision complexity of the reservoir group operation rules. The complexity evaluation model of the operation rules is as follows:
[0030] In the formula: C is the decision complexity of the operation rules, D is the DT depth, N is the number of DT nodes, L is the number of DT leaf nodes, ω is the weighted coefficient, taking ω 1 = ω 2 = ω 3 =1 / 3. Based on the above definitions, the number of key decision variables in the DT operation rules is: N - L , and the importance of the corresponding decision variables is measured by the position of the key decision variables in the DT (the lower the position in the DT depth, the higher the importance of the decision variable).
[0031] Step 2.3, measure the complexity of the reservoir group operation decision-making, the flood control storage capacity of the reservoir and its decision-making complexity, as well as the key decision variables of the operation rules of each reservoir in the basin control reservoir group. The differences in the main decision variables lead to different forms of DT, that is, the number of nodes (Node), the number of leaf nodes (Leaf node), and the depth of DT (Depth) are different. To further analyze the key decision-making characteristics and decision-making complexity of the reservoir group operation, based on the DT operation rules obtained by QwenLLM, the key decision variables are identified and the decision-making complexity of the reservoir group operation rules is analyzed. The results are as Figure 6 shown.
[0032] Figure 6 The spatial distribution of the decision-making complexity of the operation rules of the basin control reservoir group in the Yangtze River Basin. The reservoirs with the lowest decision-making complexity ( C =5) are basically located in the middle and lower reaches of the Yangtze River Basin. The reservoirs with medium decision-making complexity ( C =6) are located in the upper reaches of the main stream and tributaries of the Yangtze River. And the reservoir groups with high complexity ( C =7-10) are located in the lower reaches of the main stream and tributaries of the Yangtze River. Moreover, the cascade reservoir group in the lower reaches of the Jinsha River has the highest degree of decision-making complexity. On the one hand, it is because the flood control storage capacity of the cascade reservoir group in the lower reaches of the Jinsha River needs to be jointly used during flood control operation. On the other hand, the flood control operation in the lower reaches of the Jinsha River needs to cooperate with the flood control operation of the Three Gorges Reservoir to reduce the flood risk in the middle and lower reaches of the Yangtze River. From the spatial distribution of the decision-making complexity of the cascade reservoir group, it shows a trend of gradually increasing from the upper reaches to the lower reaches, and the difference between the south and north of the Yangtze River is not significant.
[0033] Step 3, analyze the correlation between the operation rules and the reservoir characteristics.
[0034] Step 3.1, regarding the problem of whether the greater the flood control storage capacity of the reservoir, the higher its decision-making complexity, further use correlation calculation and principal component analysis (PCA) to determine the correlation between the DT operation rules and the reservoir characteristics (storage capacity, inflow runoff, location, etc.). PCA first maps the p d-dimensional original data X p to a new coordinate system, calculates the eigenvalues and eigenvectors of the sample covariance matrix after mapping, and selects the eigenvectors a k corresponding to the first pi as the principal components, thus converting the high-dimensional data set into a low-dimensional representation. The i th principal component F i is:
[0035] In the i th principal component F i among them, the larger the coefficient of the index X p indicates that the index contributes more to the principal component. By comparing the contribution degrees of each index to the corresponding principal component, each principal component can be divided into different categories, thus realizing the classification of the principal components.
[0036] Step 3.2, by adopting the PCA dimensionality reduction method, first reduce the characteristic variables of the reservoir group to a few representative principal components, and analyze the contribution degrees of each characteristic variable to the corresponding principal component, so as to realize the classification analysis of the reservoir group characteristics. The characteristic variables of the reservoir group include reservoir capacity, inflow runoff, spatial location, etc., and a specific index is used to completely classify the reservoir group characteristics. After reducing the dimensionality of the reservoir group characteristic variables, further map and analyze the decision-making complexity of the reservoir group flood control operation rules and the principal components to determine the correlation between the decision-making complexity and the reservoir group characteristics. The results are shown in Table 1. On this basis, further use PCA to reduce the dimensionality of the reservoir group characteristic variables and analyze the correlation pattern between the decision-making complexity of the operation rules and the reservoir group characteristics. The results are as Figure 7 shown.
[0037] The results in Table 1 show that the basin area and installed capacity are strongly positively correlated with the decision-making complexity. This means that reservoirs with a larger basin area and a higher installed capacity need to handle more complex operation tasks and involve more decision variables. Therefore, larger reservoirs usually require more refined operation strategies, increasing the complexity of decision-making. Similarly, the total reservoir capacity, flood control reservoir capacity, and regulating reservoir capacity of the reservoir are also positively correlated with the decision-making complexity. These characteristics determine that the reservoir needs to manage a larger amount of water, and the operation strategy is more complex. Especially in extreme weather events, decisions such as water level control, water storage, and flood discharge of the reservoir will be more dynamic and complex. However, different from variables such as basin area and reservoir capacity, the water level characteristics of the reservoir (such as normal pool level, flood control level, flood season flood control limit level, etc.) have a relatively small direct impact on the decision-making complexity. In addition, the impact of static characteristic variables such as reservoir capacity coefficient on the decision-making complexity is also weak, mainly affecting the flexibility and emergency response ability of reservoir operation, rather than directly driving the complexity of operation decisions. When analyzing the interaction effect between the reservoir group characteristics and the operation rules, the reservoir group characteristic variables (such as basin area, total reservoir capacity, flood control reservoir capacity, etc.) have a greater impact on the operation rules. Especially when facing the forecasted inflow runoff and forecasted rainfall, a larger basin area and a higher reservoir capacity require more accurate forecasts to support the operation decision, thus increasing the complexity of the decision-making. In addition, upstream floods are also strongly correlated with reservoir characteristics (such as basin area, total reservoir capacity, etc.). During flood operation, the reservoir group characteristics determine the difficulty and flexibility of the operation decision.
[0038] Table 1 Correlation between characteristic variables and operation rules of reservoir groups
[0039] In Table 1, "*" represents the significance level p >0.05, the correlation is not significant
[0040] Figure 7 The results show that in the principal component analysis, the contribution degrees of the first principal component and the second principal component to the characteristic variance of the reservoir group are 42.32% and 35.06% respectively, indicating that the first two principal components can effectively represent the overall characteristics of the reservoir group. Specifically, the characteristic variables with the highest contribution degree of the first principal component include the controlled drainage area, total storage capacity, storage capacity below normal storage level, regulation storage capacity, flood control storage capacity, installed capacity, average annual runoff and storage coefficient, and the remaining characteristic variables mainly contribute to the second principal component. Large-capacity reservoirs such as the Three Gorges and Danjiangkou are mainly concentrated in the direction of the first principal component, while other reservoirs are mainly concentrated in the direction of the second principal component. Further, from the mapping relationship between the principal components and the decision-making complexity of the operation rules, no obvious clustering characteristics are presented in the decision-making complexity. It shows that although the characteristics of the reservoir group are concentrated in the space of the first and second principal components, this does not affect the clustering pattern of the decision-making complexity of the operation rules
[0041] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention
Claims
1. A large language model standardization processing method for reservoir group dispatching rules, characterized by The following steps are involved: Step 1, constructing a DT dispatching rule for a reservoir group based on a large language model. In order to ensure the accuracy of the large language model after local deployment, the accuracy and reliability of the large language model are verified by comparing the knowledge question answering with the online large language model; the DT represents a decision tree; Step 2: Analyze the decision complexity of the DT dispatching rules of the reservoir group, use the weighted complexity evaluation model, calculate the decision complexity of the DT dispatching rules of the reservoir group through the DT depth, the number of nodes and the number of leaf nodes, and estimate the decision complexity of the DT dispatching rules of the basin-controlling reservoir group; Step 3: Analyze the association between the DT dispatching rules and reservoir characteristics, and determine the correlation between the DT dispatching rules and reservoir characteristics by using correlation calculation and principal component analysis; reduce the dimension of the characteristic variables of the reservoir group into representative principal components, and analyze the contribution of each characteristic variable to the corresponding principal component to achieve classification analysis of the characteristics of the reservoir group; after reducing the dimension of the characteristic variables of the reservoir group, map the decision complexity of the DT dispatching rules of the reservoir group and the principal components to determine the correlation between the decision complexity and the characteristics of the reservoir group.
2. According to claim 1, a large language model standardization processing method for DT dispatching rules of a reservoir group is characterized by: In step 1, the following method is used to construct a DT dispatch rule model based on a large language model, and the text description of the DT dispatch rule is recognized and processed into a logical form: Step 1.1: Localize the large language model and build the DT dispatching rules for the reservoir group: Localize the large language model based on the Windows PowerShell platform and the open source Ollama platform; Step 1.2: Deploy the computer equipment required for the large language model, including CPU, GPU, and large-capacity storage space; Step 1.3, model selection and deployment: select Qwen 2.5:7b version model of QwenLLM for local deployment; In step 1.4, the QwenLLM is a trained and verified model, so after localized deployment, there is no need to retrain the parameters of QwenLLM. In order to ensure the accuracy of the large language model after localized deployment, the accuracy and reliability of QwenLLM are verified by comparing the knowledge question and answer with the online large language model.
3. The large language model standardization processing method for reservoir group dispatching rules according to claim 2 is characterized by: In step 1.4, the verification method for the accuracy and reliability of QwenLLM is as follows: after the local deployment of the large language model is realized, the reservoir group dispatching rules are input into the large language model by means of knowledge question and answer, and the key decision variables and thresholds in the dispatching rules of each reservoir are inquired and standardized into the form of DT to obtain the DT dispatching rule library based on the large language model; after obtaining the DT dispatching rule library, further manual inspection is adopted to compare and verify the dispatching rules based on the large language model with the text one by one to ensure the accuracy of the DT dispatching rules.
4. The large language model standardization processing method for reservoir group dispatching rules according to claim 2 is characterized by: The deployment process of the large language model in step 1.3 is as follows: ① Enter the ollama download page, select the Windows version of ollama to download and install; ②After the installation is complete, enter the Ollama command in Windows PowerShell to verify whether ollama is installed successfully; ③After Ollama is successfully installed, select the Qwen 2.5 model in the Ollama model library, enter Ollama run qwen2.5:7b in Windows PowerShell to download, install and run the model, thereby realizing the localized deployment of the large language model.
5. The large language model standardization processing method for reservoir group dispatching rules according to claim 1 is characterized by: In step 2, the sub-steps are: Step 2.1, analyze the decision complexity of the DT scheduling rule, and obtain the decision variables confirmed by the DT scheduling rule based on the locally deployed large language model; The decision variable is the decision judgment node of DT; Step 2.2, build a weighted complexity evaluation model based on the DT dispatch rule, calculate the decision complexity of the dispatch rule by DT depth, number of nodes and number of leaf nodes, and evaluate the decision complexity of the reservoir group dispatch rule; the complexity evaluation model of the dispatch rule is as follows: ; Where: C is the decision complexity of the scheduling rule, D is the DT depth, N is the number of DT nodes, L is the number of DT leaf nodes, ω is the weighting coefficient, ω 1= ω 2= ω 3=1 / 3; The number of key decision variables in the DT scheduling rule is: N - L , the importance of the corresponding decision variables is measured by the position of the key decision variables in the DT, that is, the lower the depth of the DT, the higher the importance of the decision variable; Step 2.3, measure the complexity of reservoir group operation decision, including the flood control storage capacity of the reservoir and the complexity of its decision, as well as the key decision variables of the operation rules of each reservoir in the basin-controlled reservoir group.
6. The large language model standardization processing method for reservoir group dispatching rules according to claim 4 is characterized by: In step 2.3, the difference in key decision variables leads to different forms of DT. The different forms of DT include the number of nodes, the number of leaf nodes and the depth of DT. The depth of DT is the number of nodes on the longest path from the root node to the deepest leaf node. The nodes include internal nodes and leaf nodes. The leaf node is the final node in DT, indicating the final decision made. Therefore, due to the large differences in the decision-making process of reservoir group scheduling due to the decision variables and DT forms in the scheduling rules, the more decision variables the scheduling rules of the reservoir group have and the more complex the form, the higher the complexity of its scheduling decision.
7. The large language model standardization processing method for reservoir group dispatching rules according to claim 1 is characterized by: In step 3, the following method is used to analyze the association between the dispatching rules and reservoir characteristics; Step 3.1, in order to determine the relationship between DT dispatching rules and reservoir characteristics, the correlation calculation and principal component analysis are used to determine the relationship between DT dispatching rules and reservoir characteristics. PCA first transforms the p Dimensional raw data X p Map it to a new coordinate system, calculate the eigenvalues and eigenvectors of the mapped sample covariance matrix, and select the previous k The eigenvalues corresponding to the eigenvector a pi As the principal component, the high-dimensional data set is converted into a low-dimensional representation; i principal components F i for: ; In the i principal components F i Indicators X p The larger the coefficient is, the greater the contribution of the indicator to the principal component is. By comparing the contribution of each indicator to the corresponding principal component, each principal component is divided into different categories, thereby achieving the classification of the principal components. Step 3.2, by adopting the PCA dimensionality reduction method, first reduce the dimension of the characteristic variables of the reservoir group into a few representative principal components, and analyze the contribution of each characteristic variable to the corresponding principal component to achieve classification analysis of the characteristics of the reservoir group; the characteristic variables of the reservoir group include storage capacity, inflow flow, and spatial location, and a specific indicator is used to completely classify the characteristics of the reservoir group; after reducing the dimension of the characteristic variables of the reservoir group, further map the decision complexity of the reservoir group scheduling rules and the principal components to determine the correlation between the decision complexity and the characteristics of the reservoir group.
8. A system for a large language model standardization processing method for reservoir group dispatching rules, characterized by: A large language model standardization processing method for a reservoir group dispatching rule according to any one of claims 1-7 is adopted.
9. A computer device, characterized in that: include: one or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, a large language model standardization processing method for reservoir group scheduling rules according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, a large language model standardization processing method for reservoir group scheduling rules according to any one of claims 1-7 is implemented.