A method and equipment for assessing water footprint
By introducing a random forest regression model and multidimensional feature parameters, combined with one-heat coding and standardization, the accuracy problem of sustainable aviation fuel water footprint assessment was solved, enabling more efficient assessment and optimization of production processes.
Patent Information
- Application Number
- CN202510135366.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Existing technologies for assessing the water footprint of sustainable aviation fuels are not accurate enough. Direct measurement methods suffer from omissions and errors, while the water consumption rate method uses overly simplistic basic parameters.
A random forest regression model was adopted, which combined multiple assessment dimensions such as the attributes of raw material crops, geographical environment and sustainable aviation fuel production environment. Water footprint assessment was carried out through multiple decision trees, and the feature parameters were one-hot encoded and standardized. Feature importance analysis was used to improve the accuracy of assessment.
It improves the accuracy of sustainable aviation fuel water footprint assessment, adapts to different combinations of characteristic parameters, and provides a wealth of references to optimize production processes and promote sustainable development.
Smart Images

Figure CN119962838B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and device for assessing water footprint. Background Technology
[0002] With the rapid development of the air transport industry, the demand for aviation fuel continues to grow. However, the use of traditional fossil fuels has led to substantial carbon emissions and environmental pollution. Therefore, developing low-carbon and environmentally friendly sustainable aviation fuel (SAF) has become one of the important directions for solving this problem. SAF not only aims to reduce carbon emissions but also other environmental impacts; water footprint is one of the important indicators for measuring the environmental impact of fuel production.
[0003] Currently, the water footprint of sustainable aviation fuel is typically assessed using methods such as the "direct measurement method" or the "water consumption rate method." The direct measurement method involves directly measuring water usage at each stage of water consumption involved in sustainable aviation fuel development to calculate the water footprint. The water consumption rate method, on the other hand, predicts the required water footprint based on historical data or industry averages, combined with production scale. However, the direct measurement method frequently suffers from omissions and errors in measurement, while the water consumption rate method uses overly narrow parameter values, resulting in inaccurate water footprint assessments using existing methods. Summary of the Invention
[0004] This application provides a water footprint assessment method and apparatus to improve the accuracy of water footprint assessment for sustainable aviation fuels.
[0005] This application provides a water footprint assessment method, characterized by including:
[0006] In response to a water footprint assessment instruction initiated for the sustainable aviation fuel to be produced, a set of feature parameters to be processed is obtained, the set of feature parameters including attribute-type feature parameters of raw material crops, geographical environment-type feature parameters of raw material crops, and / or production environment-type feature parameters of sustainable aviation fuel;
[0007] The set of feature parameters is input into the pre-trained water footprint assessment model. The water footprint assessment model adopts a random forest regression model and contains multiple preset decision trees. The water footprint assessment model is used to assess the water footprint of the set of feature parameters through multiple decision trees.
[0008] In the water footprint assessment model, the set of feature parameters is input into the multiple decision trees respectively, so that assessment results are generated by the multiple decision trees respectively;
[0009] The evaluation results generated by the multiple decision trees are aggregated to obtain the final evaluation result, which serves as the water footprint evaluation result for the sustainable aviation fuel.
[0010] Further, the set of feature parameters is input into the pre-trained water footprint assessment model, including:
[0011] The feature parameters contained in the feature parameter set are divided into at least one parameter group, and the variable types corresponding to different parameter groups are different.
[0012] If there exists a first parameter group whose variable type is a categorical variable, then the feature parameters in the first parameter group are one-hot encoded to obtain the corresponding preprocessed parameters.
[0013] If there exists a second parameter group with a variable type of numeric, then the feature parameters in the second parameter group are standardized to obtain the corresponding preprocessed parameters.
[0014] The preprocessed set of feature parameters is input into the water footprint assessment model.
[0015] Furthermore, the attribute parameters of the raw material crop include parameter values under one or more characteristic dimensions such as crop category, growth cycle, unit yield, and reference water footprint; the geographical environment parameters of the raw material crop include parameter values under one or more characteristic dimensions such as geographical location, precipitation, and sunshine duration; the production environment parameters of the sustainable aviation fuel include parameter values under one or more characteristic dimensions such as raw material transportation distance and energy proportion; wherein, the first parameter group includes parameter values under the crop category and / or the geographical location; the second parameter group includes parameter values under the growth cycle, the unit yield, the reference water footprint, the precipitation, the sunshine duration, the transportation distance, and / or the energy proportion.
[0016] Furthermore, the method also includes:
[0017] In the process of evaluating the water footprint of the multiple decision trees in the water footprint assessment model, the influence of the evaluation result of each node in the multiple decision trees is determined, and each node corresponds to a feature dimension.
[0018] Based on the impact of the evaluation results determined for each node, the importance of features is statistically analyzed in terms of feature dimensions.
[0019] In the multiple decision trees, there are one or more nodes corresponding to the same feature dimension. If there are multiple nodes, the influence of the evaluation results corresponding to each of the multiple nodes is fused to calculate the feature importance corresponding to the feature dimension.
[0020] Furthermore, the method also includes:
[0021] The feature dimensions are ranked according to their importance as statistically determined for different feature dimensions.
[0022] The importance of the sorted feature dimensions is visualized to show the contribution of different feature dimensions to the water footprint assessment results.
[0023] Furthermore, the training process of the water footprint assessment model includes:
[0024] Under the preset evaluation dimensions, feature dimensions are determined to obtain a set of feature dimensions. The evaluation dimensions include one or more factors in the attributes of the raw material crop, the geographical environment of the raw material crop, and / or the production environment of sustainable aviation fuel.
[0025] For any decision tree to be constructed, select a subset of feature dimensions from the set of feature dimensions.
[0026] Multiple training samples are extracted from the training sample set to form a training set;
[0027] The decision tree is constructed based on the selected feature dimension, and the feature dimension serves as a node in the decision tree.
[0028] Continue to construct other decision trees to form the water footprint assessment model based on the constructed decision trees;
[0029] Based on the training set, the water footprint assessment model is optimized in terms of model parameters and the decision tree structure is adjusted to complete the training of the water footprint assessment model.
[0030] Furthermore, the method also includes:
[0031] Multiple training samples are extracted from the training sample set to serve as the test set;
[0032] The test set is input into the water footprint assessment model to obtain the test results;
[0033] The water footprint assessment model was validated using cross-validation to optimize its hyperparameters.
[0034] The validation metrics used in the cross-validation process include mean squared error and / or mean absolute error.
[0035] Furthermore, the evaluation results generated by the multiple decision trees are aggregated to obtain the final evaluation result, including:
[0036] The average of the evaluations generated by the multiple decision trees is calculated.
[0037] The calculated mean value is taken as the final evaluation result.
[0038] Furthermore, the method also includes:
[0039] The water footprint assessment results corresponding to the sustainable aviation fuel are visualized.
[0040] This application also provides a computing device, including a memory, a processor, and a communication component;
[0041] The memory is used to store one or more computer instructions;
[0042] The processor is coupled to the memory and the communication component to execute one or more computer instructions for performing the aforementioned water footprint assessment method.
[0043] This application proposes a method for assessing the water footprint of sustainable aviation fuels. On one hand, it proposes introducing multiple assessment dimensions, such as the attributes of raw material crops, geographical environment, and the production environment of sustainable aviation fuels, as assessment dimensions for water footprint evaluation. On the other hand, it also proposes using a random forest regression model to assess the water footprint based on the characteristic parameters under each introduced assessment dimension. Thus, by introducing the random forest regression model and multi-dimensional characteristic parameters, the complex influence of various characteristic parameters on water footprint assessment can be fully analyzed, thereby improving the accuracy of water footprint assessment. Furthermore, it can adapt to the diversity of raw material crops, geographical environment, and production environment. By flexibly adjusting the set of input characteristic parameters, the water footprint of sustainable aviation fuels under different combinations of characteristic parameters can be efficiently assessed, providing rich references for the production of sustainable aviation fuels, thereby better optimizing the water footprint of sustainable aviation fuels and accelerating the achievement of sustainable development goals. Attached Figure Description
[0044] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0045] Figure 1 A flowchart illustrating a water footprint assessment method provided in this application;
[0046] Figure 2 This application provides a schematic diagram of a data collection and preprocessing process.
[0047] Figure 3A schematic diagram illustrating a model training and validation process provided in this application;
[0048] Figure 4 A visualization of feature importance provided for this application;
[0049] Figure 5 This is a schematic diagram of the structure of a computing device provided as another exemplary embodiment of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] Before proceeding with a detailed description of the technical solutions provided in the various embodiments of this application, the following is a brief explanation of several technical concepts involved in this application.
[0052] The water footprint can be understood as the amount of water resources required for all products and services consumed by a country, a region, or an individual within a certain period of time.
[0053] Sustainable aviation fuel is a special type of fuel designed specifically for aircraft.
[0054] As described in the background section, the water footprint of sustainable aviation fuel is currently assessed using methods such as the "direct measurement method" or the "water consumption rate method." However, the accuracy of the water footprint assessed by these existing methods is not good.
[0055] To overcome the aforementioned problems in the existing technology, this embodiment proposes a method for assessing the water footprint of sustainable aviation fuels by combining the fields of environmental science, aviation new energy, and artificial intelligence, thereby improving the accuracy of water footprint assessment for sustainable aviation fuels. The technical concept of this application embodiment is described below.
[0056] The technical concept of this application proposes to introduce multiple evaluation dimensions, such as the attributes of raw material crops, geographical environment, and production environment of sustainable aviation fuel, as evaluation dimensions for assessing the water footprint of sustainable aviation fuel.
[0057] Through research and exploration, the inventors discovered that various raw material crops used to produce sustainable aviation fuel, such as reeds and castor beans, have a significant impact on the environment due to water consumption (i.e., water footprint) during their growth and sustainable aviation fuel production processes. They proposed designing a rich set of feature dimensions under the aforementioned two assessment dimensions of raw material crop attributes and geographical environment to explore the complex impact of these two assessment dimensions on the water footprint. The inventors also found that the production environment of sustainable aviation fuel also has a significant impact on the water footprint; therefore, they proposed designing a rich set of feature dimensions under the aforementioned assessment dimension of the sustainable aviation fuel production environment to explore the complex impact of this assessment dimension on the water footprint.
[0058] For example, in the evaluation dimension of raw material crop attributes in this application embodiment, the introduced feature dimensions may include, but are not limited to, crop category, growth cycle, unit yield, and reference water footprint. In the evaluation dimension of the raw material crop's geographical environment, the introduced feature dimensions may include, but are not limited to, geographical location, precipitation, and sunshine duration. In the evaluation dimension of sustainable aviation fuel production environment, the introduced feature dimensions may include, but are not limited to, transportation distance and energy proportion, where energy proportion describes the proportion of different types of energy in the sustainable aviation fuel production process, such as the proportion of energy from coal power, natural gas power, and renewable energy; raw material crops are typical renewable resources.
[0059] The technical concept of this application also proposes to use a random forest regression model to evaluate water footprint based on the feature parameters of each introduced evaluation dimension.
[0060] Random Forest Regression is an ensemble learning algorithm based on decision trees used to solve regression problems. It constructs multiple decision trees and combines their predictions to obtain the final regression prediction. The basic idea of ensemble learning is to combine multiple weak learners (i.e., decision trees) to form a strong learner, thereby improving the model's predictive performance and stability.
[0061] Based on the introduction of the random forest regression model and the aforementioned rich feature dimensions, the water footprint assessment method provided in this application embodiment can fully consider a variety of complex factors affecting the water footprint, such as geographical location, precipitation, sunshine hours, transportation distance and crop characteristics, and fully explore the nonlinear relationship between these complex factors and the water footprint, thereby fully analyzing the complex impact of these complex factors on the water footprint, and thus effectively improving the accuracy of water footprint assessment.
[0062] The technical concept of the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0063] Figure 1 This is a flowchart illustrating a water footprint assessment method provided in this application. The method can be executed by a data processing device, which can be implemented as software, hardware, or a combination of both, and can be integrated into a computing device. (Reference) Figure 1 The method may include:
[0064] Step 100: In response to the water footprint assessment instruction initiated for the sustainable aviation fuel to be produced, obtain a set of feature parameters to be processed, which includes attribute-type feature parameters of raw material crops, geographical environment-type feature parameters of raw material crops, and / or production environment-type feature parameters of sustainable aviation fuel.
[0065] Step 101: Input the set of feature parameters into the pre-trained water footprint assessment model. The water footprint assessment model adopts a random forest regression model and contains multiple preset decision trees. The water footprint assessment model is used to assess the water footprint for the set of feature parameters through multiple decision trees.
[0066] Step 102: In the water footprint assessment model, the set of feature parameters is input into multiple decision trees respectively, so as to generate assessment results through multiple decision trees respectively;
[0067] Step 103: Aggregate the evaluation results generated by multiple decision trees to obtain the final evaluation result, which serves as the water footprint evaluation result for sustainable aviation fuel.
[0068] In this embodiment, the timing of initiating the water footprint assessment command in step 100 is not limited. When it is necessary to conduct water footprint assessment based on different sets of feature parameters, the water footprint assessment command can be initiated as needed.
[0069] The feature parameter set in this embodiment contains several parameter values. This corresponds to the multiple evaluation dimensions introduced above, and the feature dimensions designed under each evaluation dimension. The feature parameter set contains the parameter values for each feature dimension introduced in this embodiment.
[0070] Following the previous example of feature dimensions, in step 100, the feature parameter set includes attribute parameters of raw material crops such as crop category, growth cycle, unit yield, and reference water footprint, with parameter values under one or more feature dimensions; geographical environment parameters of raw material crops such as geographical location, precipitation, and sunshine duration, with parameter values under one or more feature dimensions; and production environment parameters of sustainable aviation fuel such as transportation distance and energy proportion, with parameter values under one or more feature dimensions. It is understood that these feature dimensions are merely exemplary, and this embodiment is not limited to them. More feature dimensions can be designed based on the evaluation dimensions introduced in this embodiment to more comprehensively uncover the feature dimensions that influence the water footprint. Examples of more feature dimensions are not provided here.
[0071] Continue to refer to Figure 1 In step 101, the set of feature parameters can be input into the pre-trained water footprint assessment model. In this embodiment, the water footprint assessment model adopts a random forest regression model. In this embodiment, the input of the water footprint assessment model is the set of feature parameters, and the output is the assessment result.
[0072] refer to Figure 1 After the feature parameter set is input into the water footprint assessment model, in step 102, the feature parameter set will be fed into multiple decision trees contained in the water footprint assessment model. The multiple decision trees can independently perform water footprint assessment based on the feature parameter set and generate assessment results respectively.
[0073] In step 103, the evaluation results generated by multiple decision trees can be aggregated to obtain the final evaluation result, which serves as the water footprint evaluation result for sustainable aviation fuel.
[0074] In step 103, an optional aggregation operation could be: calculating the average of the evaluations generated by multiple decision trees; and using the calculated average as the final evaluation result. Of course, other methods such as taking the median could also be used to achieve the aggregation operation; this is not limited to these methods, nor will further examples be provided.
[0075] It is worth noting that the aggregation operation in step 103 can be completed within the water footprint assessment model. Alternatively, the data processing device in this embodiment can obtain the assessment results generated by multiple decision trees from the water footprint assessment model and then perform the aggregation operation. This embodiment does not limit this approach.
[0076] In addition, the technical concept of this application also proposes that the water footprint assessment results corresponding to sustainable aviation fuel can be visualized.
[0077] In summary, this embodiment proposes a method for assessing the water footprint of sustainable aviation fuels. On one hand, it introduces multiple assessment dimensions, such as the attributes of raw material crops, geographical environment, and the production environment of sustainable aviation fuels, as evaluation dimensions for assessing the water footprint of sustainable aviation fuels. On the other hand, it also proposes using a random forest regression model to assess the water footprint based on the characteristic parameters under each introduced assessment dimension. Thus, by introducing the random forest regression model and multi-dimensional characteristic parameters, the complex influence of various characteristic parameters on water footprint assessment can be fully analyzed, thereby improving the accuracy of water footprint assessment. Furthermore, it can adapt to the diversity of raw material crops, geographical environment, and production environment. By flexibly adjusting the set of input characteristic parameters, the water footprint of sustainable aviation fuels under different combinations of characteristic parameters can be efficiently assessed, providing rich references for the production of sustainable aviation fuels, thereby better optimizing the water footprint of sustainable aviation fuels and accelerating the achievement of sustainable development goals.
[0078] The technical concept of this application embodiment also proposes that the feature parameter set can be preprocessed before being input into the water footprint assessment model. An optional technical concept proposes that the feature parameters contained in the feature parameter set can be divided into at least one parameter group, with different parameter groups corresponding to different variable types; if there is a first parameter group with categorical variables, the feature parameters in the first parameter group are one-hot encoded to obtain the corresponding preprocessed parameters; if there is a second parameter group with numerical variables, the feature parameters in the second parameter group are standardized to obtain the corresponding preprocessed parameters; the preprocessed feature parameter set is then input into the water footprint assessment model.
[0079] In this optional technical concept, the feature parameters in the feature parameter set are grouped according to variable type, and different preprocessing methods are designed for different parameter groups. One-hot encoding technology will be illustrated in detail later. Standardization processing can be understood as unifying the dimensions of the values in different parameter groups to more evenly analyze the impact of relevant feature dimensions on water footprint assessment.
[0080] Following the example of feature dimensions mentioned above, we will now provide examples of the first and second parameter groups. The first parameter group may contain parameter values for crop category and / or geographical location; the second parameter group may contain parameter values for growth cycle, unit yield, reference water footprint, precipitation, sunshine duration, transport distance, and / or energy proportion. There are no restrictions on the number of parameter groups, the types of variables involved in different parameter groups, or the feature dimensions included in different parameter groups.
[0081] By preprocessing the set of feature parameters, the set of feature parameters can be organized into a more standardized and structured manner, which makes it easier for the water footprint assessment model to perform data analysis.
[0082] In the technical concept of this application embodiment, it is also proposed that: during the process of evaluating the water footprint of multiple decision trees in the water footprint assessment model for the feature parameter set, the influence degree of the evaluation result corresponding to each node in the multiple decision trees can be determined, and each node corresponds to a feature dimension; based on the influence degree of the evaluation result determined for each node, the feature importance is statistically analyzed in units of feature dimension; wherein, in the multiple decision trees, there are one or more nodes corresponding to the same feature dimension, and if there are multiple nodes, the influence degree of the evaluation result corresponding to each of the multiple nodes is fused to statistically analyze the feature importance corresponding to the feature dimension.
[0083] Feature importance can be used to reflect the contribution of a feature dimension to the evaluation result. The influence of the evaluation result determined by the node can be expressed using methods such as Mean Decrease in Impurity or Mean Decrease in Impurity with Permutation. Taking Gini impurity as an example, for a feature dimension, at each node of each decision tree, if that feature dimension is selected for splitting, the difference in Gini impurity before and after the split (representing "impurity reduction") is calculated. These differences are then averaged across all nodes and all decision trees to obtain the mean impurity reduction value for that feature dimension. The larger this value, the more important the feature dimension, i.e., the higher the feature importance.
[0084] Furthermore, the feature dimensions can be ranked according to the statistical importance of the features for different feature dimensions; the importance of the ranked feature dimensions can be visualized to show the contribution of different feature dimensions to the water footprint assessment results.
[0085] By analyzing the importance of features, we can fully uncover the key factors influencing water footprint and provide a basis for optimizing water resource management.
[0086] The technical concept of this application also proposes: pre-training the water footprint assessment model. To this end, feature dimensions can be determined under preset assessment dimensions to obtain a feature dimension set; referring to the preceding text, the assessment dimensions may include one or more dimensions such as the attributes of the raw material crop, the geographical environment of the raw material crop, and / or the production environment of sustainable aviation fuel; for any decision tree to be constructed, some feature dimensions are selected from the feature dimension set; multiple training samples are extracted from the training sample set as a training set; based on the selected feature dimensions, a decision tree is constructed, with the feature dimensions serving as nodes in the decision tree; other decision trees are then constructed to form a water footprint assessment model based on the constructed decision trees; based on the training set, the model parameters of the water footprint assessment model are optimized and the decision tree structure is adjusted to complete the training of the water footprint assessment model.
[0087] Additionally, multiple training samples can be extracted from the training sample set as a test set. The test set is then input into the water footprint assessment model to obtain test results. The performance of the water footprint assessment model is then validated using cross-validation to optimize its hyperparameters. The validation metrics used in the cross-validation process include mean squared error and / or mean absolute error. These hyperparameters may include, but are not limited to, the number of trees (n_estimators), the maximum tree depth (max_depth), and the minimum number of samples required for a leaf node (min_samples_leaf), etc. No further examples are provided here.
[0088] In this way, the model parameters, decision tree structure, and hyperparameters in the water footprint assessment model can be optimized based on the training set and the test set, thereby obtaining a water footprint assessment model suitable for the embodiments of this application.
[0089] Based on the foregoing description of the technical concepts in the embodiments of this application, exemplary embodiments of this application are provided below. The steps included in the exemplary embodiments are described below:
[0090] 1) Data collection and preprocessing
[0091] Factors influencing the water footprint are collected, such as crop type, crop growth cycle, crop yield per unit area, geographical location, local average annual precipitation, local average annual sunshine hours, transportation distance from raw materials to fuel processing plants, and the power generation energy structure of a specific year. These influencing factors correspond to the feature dimensions introduced in the embodiments of this application.
[0092] The crop type and geographical location data require special processing. One-hot encoding is used for data preprocessing of these variables to transform them into a numerical format usable by the model. The one-hot encoding process will be described in detail below. The remaining data undergoes standardization to ensure the model can effectively handle different types of data. The model here corresponds to the water footprint assessment model introduced in the embodiments of this application.
[0093] This embodiment outlines the process flow, as follows: Figure 2 As shown, Figure 2 This application provides a schematic diagram of a data collection and preprocessing process.
[0094] 2) Model Building
[0095] This embodiment employs the Random Forest regression algorithm. Random Forest is an ensemble learning algorithm that captures complex nonlinear relationships by training multiple decision trees. Each decision tree independently learns a portion of the dataset, and the prediction accuracy is improved by weighted averaging of the outputs of multiple trees. Model input features are defined, including crop conditions, climate conditions, transportation distance, etc. Specific feature dimensions are described above. To improve the model's predictive performance, grid search and cross-validation techniques are used to optimize the hyperparameters of the random forest model to construct the desired model. By selecting the optimal combination of hyperparameters, the model's high generalization ability under diverse input conditions is ensured.
[0096] The specific components of the random forest regression model are as follows:
[0097] Step 1: Randomly select a subset of samples as the training set for the decision tree.
[0098] Step 2: Randomly select a subset of features (the square root of the total number of features) as the feature set of the decision tree.
[0099] Step 3: Build a decision tree based on the training set and feature set until the predetermined number of leaf nodes is reached or the tree cannot be split.
[0100] Step 4: Repeat the above steps to build multiple decision trees.
[0101] Step 5: For a new sample, input it into each decision tree to obtain multiple prediction results.
[0102] Step 6: Average the multiple prediction results to obtain the final prediction result.
[0103] Its algorithm formula is based on the decision tree regression model, and the prediction function of each decision tree can be expressed as shown in formula (1):
[0104]
[0105] In the formula: K represents the Kth decision tree, x represents the input sample, and J k c represents the number of leaf nodes in the k-th decision tree. kj R represents the predicted value of the j-th leaf node in the k-th decision tree. kj Let represent the sample set of the leaf nodes of the j-th decision tree.
[0106] The prediction function of multiple decision trees can be expressed as:
[0107]
[0108] In the formula: P represents the number of decision trees.
[0109] For model evaluation, the evaluation metrics that can be used include mean squared error (MSE) and coefficient of determination R-squared (R²). 2, It can also be expressed as R^2). Generally speaking, the smaller the MSE value, the better the model fits the data. 2 The closer the value is to 1, the better the model fits the data, and vice versa. The calculation formula is as follows:
[0110]
[0111] In the formula, n represents the sample size, y i This represents the true value of the i-th sample. This represents the predicted value of the i-th sample.
[0112]
[0113] In the formula: This represents the average of the true values of all samples.
[0114] 3) Model training and validation
[0115] The model is trained using collected data, which is then split into training and test sets. Cross-validation is used to evaluate the model's performance. The accuracy of the model is assessed using metrics such as mean squared error (MSE). The specific training and validation process is as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of a model training and validation process provided in this application.
[0116] 4) Feature Importance Analysis
[0117] By utilizing the feature importance analysis function of the random forest model, the contribution of different input variables to water footprint prediction can be evaluated. This feature importance analysis helps industries involved in SAF (Self-Funded Forest) identify key factors influencing water footprint and implement appropriate controls during later production processes to ensure optimal economic benefits.
[0118] 5) Water footprint assessment
[0119] A well-trained model can predict the water footprint of a crop-based fuel accretion flotation (SAF) under specific conditions based on different input variables. These input variables may include: crop growth cycle, yield per unit area, average annual precipitation, average annual sunshine hours, transport distance, and the proportions of coal-fired power, natural gas, and renewable energy in the energy mix. These input variables correspond to the set of feature parameters mentioned earlier.
[0120] 6) Results visualization
[0121] Visualization tools can be used to display actual and predicted water footprint results, and feature importance can be presented in chart form. This not only helps in understanding the model's output, but also visually demonstrates the model's support for decision-making. Figure 4 This is a schematic diagram illustrating the visualization effect of feature importance provided in this application.
[0122] Based on the water footprint assessment method provided in the embodiments of this application, the following technical effects can be achieved in the above exemplary embodiments.
[0123] 1) Processing of multidimensional data: This embodiment introduces the random forest algorithm, which can simultaneously handle the complex impacts of different crops, climate conditions, geographical locations, and energy structures on water footprint, significantly improving the model's prediction accuracy. Random forests have good nonlinear modeling capabilities and can handle complex interactions between high-dimensional variables.
[0124] 2) Dynamic optimization: The model can adapt to changes in different geographical regions, climate conditions, and energy structures through continuous updates and optimization. This dynamic optimization capability makes the method potentially applicable flexibly to different regions and production conditions.
[0125] 3) Feature Importance Analysis: This embodiment can not only predict the water footprint, but also identify the main factors affecting the water footprint through feature importance analysis, such as the water requirements of specific crops, the impact of transportation distance, and the contribution of energy structure to water resource use in production. This function provides valuable reference for SAF production enterprises when designing and optimizing production processes.
[0126] 4) Applicable to various production scenarios: This method is applicable to various sustainable aviation fuel production scenarios, including those using a variety of different biomass feedstocks and energy structures. Enterprises can flexibly adjust input variables according to different crops and production conditions to obtain accurate water footprint predictions.
[0127] Moreover, the water footprint assessment method provided in this application embodiment is applicable to at least the following use cases:
[0128] 1) Biomass feedstock planting planning: Enterprises can select biomass feedstocks with higher water resource utilization efficiency based on the water footprint of different crops, and optimize planting plans. At the same time, they can also select suitable planting areas under the premise of planting a specific crop to ensure maximum benefits.
[0129] 2) Transportation Management: By assessing the impact of different transportation distances on the water footprint, companies can rationally plan raw material transportation routes and reduce unnecessary resource consumption during transportation. Simultaneously, they can adjust the geographical locations of raw material production sites and production areas according to actual conditions to ensure convenience throughout the entire production process.
[0130] 3) Production process design: Against the backdrop of energy structure transformation, the utilization rate of clean energy and renewable energy is gradually increasing. By understanding the impact of energy structure on water footprint, relevant enterprises can optimize energy use in production processes and reduce water consumption.
[0131] Thus, through the water footprint assessment method provided in this application, enterprises can not only optimize resource use in the production process, but also better cope with possible future environmental regulations and standards, and promote the popularization and promotion of sustainable aviation fuel.
[0132] The following describes the specific application of the water footprint assessment method provided in this application embodiment in an exemplary application scenario.
[0133] 1. Data Collection
[0134] 1) First, basic data related to SAF production can be collected. Data sources may include agricultural statistics, meteorological data, energy structure data, and transportation information. The following are examples of data collection instructions (corresponding to the characteristic parameters described above):
[0135] 2) Crop data: Collect data on different biomass feedstocks used in SAF production. Data items include crop type (such as reed, castor bean, rapeseed, etc.), crop growth cycle (unit: days), and unit yield (unit: tons / hectare).
[0136] 3) Climate data: Collect meteorological data from the production site, focusing on average annual precipitation (in millimeters) and average annual sunshine hours (in hours). This data can be obtained from national statistical yearbooks or local meteorological departments.
[0137] 4) Geographical Location and Transportation Data: Transportation distances (in kilometers) between raw material production sites and fuel processing plants were collected, while also considering the impact of different geographical locations on water consumption. Geographical locations were based on major SAF production areas in China, such as Xinjiang Uygur Autonomous Region, Hunan Province, and Hebei Province.
[0138] 5) Energy Structure Data: This data collects information on the types and proportions of energy used in power generation, primarily including the proportion of coal-fired power, natural gas power, and renewable energy. This data can be obtained from annual energy statistics reports or reports provided by power companies.
[0139] 6) Reference water footprint data: Obtain the water footprint data for the planting, transportation, and conversion processes of different biomass raw materials. Water footprint data can be obtained through field experiments or by referring to statistical data in relevant literature.
[0140] 2. Data Preprocessing
[0141] The collected data needs to be cleaned and preprocessed to ensure it can be processed by machine learning models. In practice, the main data preprocessing steps include:
[0142] 1) One-hot encoding of categorical variables: For categorical variables such as crop type and geographical location, one-hot encoding is used. For example, crop type (reed, castor bean, flax, etc.) and geographical location (Xinjiang, Hunan, Hebei, etc.) are converted into multiple binary features (0 or 1) so that the model can process categorical data.
[0143] 2) Standardization of numerical variables: Numerical features such as crop growth cycle, unit yield, average annual precipitation, average annual sunshine hours, transportation distance, proportion of coal-fired power, proportion of natural gas, and proportion of renewable energy are standardized. By converting each feature value into a standard normal distribution with a mean of 0 and a variance of 1, it is ensured that features of different dimensions have similar weights during model training.
[0144] This section will explain one-hot encoding, which is a method of encoding categorical variables (such as color, crop type, geographical location, etc.) into numerical form. The basic idea is that each category is represented by a set of binary numbers (0 and 1), where one category is 1 and the rest are 0.
[0145] For example, suppose there are three crop categories: Arundo donax, castor bean, and Camelina. Using one-hot encoding, these would be converted into three columns representing these three crops, as shown in Table 1 below – One-Hot Encoded Data:
[0146] Crop categories Arundodis castor bean Flaxseed Arundodis 1 0 0 castor bean 0 1 0 Flaxseed 0 0 1
[0147] In this way, each category becomes a unique set of binary numbers, for example, the representation of Reed is 100 and the representation of Castor Bean is 010.
[0148] In this embodiment, the purpose of one-hot encoding is to convert non-numerical classification information such as crop type (e.g., reed, castor bean, flax) and geographical location (e.g., Xinjiang, Hunan, Hebei) into numerical format to facilitate model processing.
[0149] Without one-hot encoding, machine learning models cannot handle non-numerical data. If we directly use the numbers 1, 2, and 3 to represent three crops, the model might misinterpret this as a ranking relationship (e.g., 1 is smaller than 2), which would affect the training results. However, with one-hot encoding, the model avoids this misinterpretation; each category is independent and unordered. This avoids misinterpretations of category priority, making the model more accurate.
[0150] 3. Model Building
[0151] After data preprocessing, the model building stage begins. This embodiment uses the random forest regression algorithm to predict the water footprint. The specific implementation steps are as follows:
[0152] 1) Model initialization: A random forest regression model is built using the `scikit-learn` library in the Python programming language. Random forests obtain more robust prediction results by integrating multiple decision trees and averaging the prediction results of each tree.
[0153] 1.from sklearn.ensemble import RandomForestRegressor
[0154] 2.rf_model=RandomForestRegressor(random_state=42)
[0155] The first line of code above imports the `RandomForestRegressor` class from the `ensemble` module of the `sklearn` library. `sklearn` is a commonly used machine learning library in Python, providing a rich set of machine learning algorithms and tools. The `ensemble` module contains various ensemble learning algorithms, while `RandomForestRegressor` is an implementation class of the random forest regression algorithm, used to solve regression problems and capable of predicting numerical target variables.
[0156] The second line of code above creates an instance of the `RandomForestRegressor` class, `rf_model`, and passes in the parameter `random_state = 42`. The purpose of `random_state` is to set the random number seed, ensuring the repeatability of the model's training and prediction results. With a fixed random number seed, each time the code is run, the random forest model's random sampling and feature selection operations during training will be based on the same random sequence, thus guaranteeing consistent training and prediction results under the same dataset and model parameters. This allows for a more accurate assessment of model performance changes as being caused by parameter adjustments rather than random factors when performing model tuning or comparing the effects of different parameter settings.
[0157] 2) Hyperparameter Optimization: To improve the model's predictive performance, a grid search method is used for hyperparameter optimization. The main parameters optimized include the number of decision trees, maximum depth, and minimum number of samples per node. An example of the hyperparameter optimization steps is as follows:
[0158]
[0159] Here, 'n_estimators' represents the number of decision trees in the forest. A random forest is composed of multiple decision trees, each of which will differ during training. A larger number of decision trees generally results in a more stable and accurate model, but training time will also increase. 100, 200, and 300 all represent the number of tree components. 'max_depth' represents the maximum depth of the decision tree, i.e., the number of nodes in the longest path from the root node to a leaf node. A deeper tree allows the model to fit the data better (especially complex data). However, if the tree is too deep, it may lead to overfitting (i.e., the model performs well on training data but poorly on new data). None: indicates no limit on tree depth; the tree will continue to split until all leaf nodes are pure (or there is no further possibility of splitting). 10, 20, and 30 all represent the maximum depth limit of the tree. 'min_samples_split' represents the minimum number of samples required for a node to split. A node is allowed to split only when the number of samples in it is greater than or equal to this value. 2: Each node requires at least 2 samples before it can continue splitting. This is the default value, which typically results in a more complex (deeper) tree. 5: Each node requires at least 5 samples before splitting, limiting the tree's complexity. 10: Each node requires at least 10 samples before splitting, resulting in a simpler, more regular tree and reducing the risk of overfitting. 'min_samples_leaf' represents the minimum number of samples required in a leaf node (the terminal node of the tree). This parameter controls how many samples a leaf node must contain to prevent excessive splitting and resulting in an overly complex tree.
[0160] 3) Model training: The training data is fitted using a hyperparameter-optimized random forest model. The model continuously learns and adjusts the decision tree structure to capture the complex factors that affect the water footprint.
[0161] 4) Feature Importance Analysis: The Random Forest algorithm has a built-in feature importance assessment function. By evaluating the impact of each feature on the prediction results, it identifies which factors contribute most to the water footprint under different crops, climate conditions, geographical locations, and energy structures. In this way, companies can better understand which factors need to be prioritized for optimization.
[0162]
[0163]
[0164] Here, 'importances': retrieves the model's feature importance score via 'best_rf_model.feature_importances'. Feature importance represents the contribution of a feature to the prediction result; the higher the value, the greater the feature's importance. 'feature_names': manually lists all feature names of the model, corresponding to the features in our previous model. 'np.argsort(importances)[::-1]': sorts the feature importance in descending order and obtains the sorted feature indices. Refer to Table 2, which illustrates the importance of different input variables (i.e., feature dimensions). It is worth noting that this is merely exemplary and should not limit the output values of this embodiment in actual applications.
[0165] feature importance Energy Structure 0.3499 Growth cycle 0.1756 Average annual precipitation 0.1323 Raw material transportation distance 0.0923 Unit output 0.0888 Crop categories 0.0675 Average annual sunshine 0.0654 Geographical location 0.0282
[0166] 4. Model Prediction and Result Analysis
[0167] Model Prediction: By inputting new data, such as a specific crop, geographical location, climate conditions, raw material transportation distance, and energy structure, the model can output predictions of water footprint. For example, assuming a certain biomass raw material has a crop growth cycle of 203 days, a yield of 7.5 tons / hectare, a raw material transportation distance of 500 kilometers, an average annual precipitation of 205 millimeters, an average annual sunshine duration of 2789.7 hours, and an energy structure where coal power accounts for 70%, natural gas for 10%, and renewable energy for 20%, these conditions can be input and predictions can be made.
[0168]
[0169] Results visualization: Visualization tools are used to display the model's prediction results, such as comparing the actual water footprint with the predicted values, helping users intuitively understand the model's accuracy. In addition, feature importance charts can be used to show the impact of different input variables on the water footprint.
[0170] 5. Application Scenarios and Benefit Analysis
[0171] The water footprint assessment method provided in this embodiment can help companies accurately predict the water footprint of sustainable aviation fuels under different conditions. Through feature importance analysis, companies can optimize planting plans, prioritize planting crops with low water footprints, and rationally plan transportation routes to reduce resource waste. In addition, adjustments to the energy structure can also be rationally optimized through model analysis to minimize water resource consumption.
[0172] Based on the description of this application scenario, the water footprint assessment method provided in this embodiment has the following advantages:
[0173] ●High-precision prediction: Compared with traditional linear models, random forests can capture complex nonlinear relationships and improve the accuracy of water footprint prediction.
[0174] ● Highly scalable: The model can handle multidimensional variables and is suitable for water footprint assessment under different crops and geographical conditions. If other input objects need to be added, they can simply be used as input variables in the dataset, facilitating subsequent improvements and modifications.
[0175] ● Flexibility: The model can adjust the input according to different crops, production sites and energy structures, making it suitable for a variety of production environments.
[0176] ●Optimize resource management: Through feature importance analysis, the model can help companies identify the factors that have the greatest impact on water consumption, thereby optimizing production processes and reducing water footprint.
[0177] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different application terminals, messages, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0178] Figure 5 This is a schematic diagram of the structure of a computing device provided as another exemplary embodiment of this application. For example... Figure 5 As shown, the computing device includes: a memory 50, a processor 51, and a communication component 52.
[0179] Processor 51, coupled to memory 50 and communication component 52, is used to execute computer programs in memory 50 for:
[0180] In response to a water footprint assessment instruction initiated for the sustainable aviation fuel to be produced, a set of feature parameters to be processed is obtained, the set of feature parameters including attribute-type feature parameters of raw material crops, geographical environment-type feature parameters of raw material crops, and / or production environment-type feature parameters of sustainable aviation fuel;
[0181] The set of feature parameters is input into the pre-trained water footprint assessment model. The water footprint assessment model adopts a random forest regression model and contains multiple preset decision trees. The water footprint assessment model is used to assess the water footprint of the set of feature parameters through multiple decision trees.
[0182] In the water footprint assessment model, the set of feature parameters is input into the multiple decision trees respectively, so that assessment results are generated by the multiple decision trees respectively;
[0183] The evaluation results generated by the multiple decision trees are aggregated to obtain the final evaluation result, which serves as the water footprint evaluation result for the sustainable aviation fuel.
[0184] In an optional embodiment, when the processor 51 inputs the set of feature parameters into the pre-trained water footprint evaluation model, it may specifically be used to:
[0185] The feature parameters contained in the feature parameter set are divided into at least one parameter group, and the variable types corresponding to different parameter groups are different.
[0186] If there exists a first parameter group whose variable type is a categorical variable, then the feature parameters in the first parameter group are one-hot encoded to obtain the corresponding preprocessed parameters.
[0187] If there exists a second parameter group with a variable type of numeric, then the feature parameters in the second parameter group are standardized to obtain the corresponding preprocessed parameters.
[0188] The preprocessed set of feature parameters is input into the water footprint assessment model.
[0189] In one optional embodiment, the attribute parameters of the raw material crop include parameter values under one or more characteristic dimensions such as crop category, growth cycle, unit yield, and reference water footprint; the geographical environment parameters of the raw material crop include parameter values under one or more characteristic dimensions such as geographical location, precipitation, and sunshine duration; the production environment parameters of the sustainable aviation fuel include parameter values under one or more characteristic dimensions such as transportation distance and energy proportion; wherein, the first parameter group includes parameter values under the crop category and / or the geographical location; the second parameter group includes parameter values under the growth cycle, the unit yield, the reference water footprint, the precipitation, the sunshine duration, the transportation distance, and / or the energy proportion.
[0190] In an alternative embodiment, processor 51 may also be used for:
[0191] In the process of evaluating the water footprint of the multiple decision trees in the water footprint assessment model, the influence of the evaluation result of each node in the multiple decision trees is determined, and each node corresponds to a feature dimension.
[0192] Based on the impact of the evaluation results determined for each node, the importance of features is statistically analyzed in terms of feature dimensions.
[0193] In the multiple decision trees, there are one or more nodes corresponding to the same feature dimension. If there are multiple nodes, the influence of the evaluation results corresponding to each of the multiple nodes is fused to calculate the feature importance corresponding to the feature dimension.
[0194] In an alternative embodiment, processor 51 may also be used for:
[0195] The feature dimensions are ranked according to their importance as statistically determined for different feature dimensions.
[0196] The importance of the sorted feature dimensions is visualized to show the contribution of different feature dimensions to the water footprint assessment results.
[0197] In an optional embodiment, the processor 51 may be specifically used during the training of the water footprint assessment model to:
[0198] Under the preset evaluation dimensions, feature dimensions are determined to obtain a set of feature dimensions. The evaluation dimensions include one or more factors in the attributes of the raw material crop, the geographical environment of the raw material crop, and / or the production environment of sustainable aviation fuel.
[0199] For any decision tree to be constructed, select a subset of feature dimensions from the set of feature dimensions.
[0200] Multiple training samples are extracted from the training sample set to form a training set;
[0201] The decision tree is constructed based on the selected feature dimension, and the feature dimension serves as a node in the decision tree.
[0202] Continue to construct other decision trees to form the water footprint assessment model based on the constructed decision trees;
[0203] Based on the training set, the water footprint assessment model is optimized in terms of model parameters and the decision tree structure is adjusted to complete the training of the water footprint assessment model.
[0204] In an alternative embodiment, processor 51 may also be used for:
[0205] Multiple training samples are extracted from the training sample set to serve as the test set;
[0206] The test set is input into the water footprint assessment model to obtain the test results;
[0207] The water footprint assessment model was validated using cross-validation to optimize its hyperparameters.
[0208] The validation metrics used in the cross-validation process include mean squared error and / or mean absolute error.
[0209] In an optional embodiment, when the processor 51 aggregates the evaluation results generated by the multiple decision trees to obtain the final evaluation result, it may specifically be used to:
[0210] The average of the evaluations generated by the multiple decision trees is calculated.
[0211] The calculated mean value is taken as the final evaluation result.
[0212] In an alternative embodiment, processor 51 may also be used for:
[0213] The water footprint assessment results corresponding to the sustainable aviation fuel are visualized.
[0214] Furthermore, such as Figure 5 As shown, the computing device also includes a display 53, a power supply unit 54, an audio unit 55, and other components. Figure 5 The diagram only shows some components and does not mean that the computing device includes only these components. Figure 5 The components shown.
[0215] It is worth noting that the technical details of the above embodiments of the computing device can be referred to the relevant descriptions in the foregoing method embodiments. To save space, they will not be repeated here, but this should not cause any loss to the scope of protection of this application.
[0216] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps in the above method embodiments.
[0217] Accordingly, this application also provides a computer program product, which, when executed, can implement the steps in the above method embodiments.
[0218] The aforementioned memory is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0219] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0220] The aforementioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.
[0221] The aforementioned power supply components provide power to various components within the device in which they reside. The power supply components may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0222] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0223] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0224] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0225] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0226] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0227] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0228] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for assessing water footprint, characterized in that, include: In response to a water footprint assessment instruction initiated for the sustainable aviation fuel to be produced, a set of feature parameters to be processed is obtained, the set of feature parameters including attribute-type feature parameters of raw material crops, geographical environment-type feature parameters of raw material crops, and / or production environment-type feature parameters of sustainable aviation fuel; The feature parameters contained in the feature parameter set are divided into at least one parameter group, and the variable types corresponding to different parameter groups are different. If there exists a first parameter group whose variable type is a categorical variable, then the feature parameters in the first parameter group are one-hot encoded to obtain the corresponding preprocessed parameters. If there exists a second parameter group with a variable type of numeric, then the feature parameters in the second parameter group are standardized to obtain the corresponding preprocessed parameters. The preprocessed set of feature parameters is input into the pre-trained water footprint assessment model. The water footprint assessment model adopts a random forest regression model and contains multiple preset decision trees. The water footprint assessment model is used to assess the water footprint for the set of feature parameters through multiple decision trees. In the water footprint assessment model, the set of feature parameters is input into the multiple decision trees respectively, so that assessment results are generated by the multiple decision trees respectively; The evaluation results generated by the multiple decision trees are aggregated to obtain the final evaluation result, which serves as the water footprint evaluation result for the sustainable aviation fuel. During the process of evaluating the water footprint of the multiple decision trees for the set of feature parameters, the influence of the evaluation result of each node in the multiple decision trees is determined, and each node corresponds to a feature dimension. Based on the impact of the evaluation results determined for each node, the importance of features is statistically analyzed in terms of feature dimensions. In the multiple decision trees, there are one or more nodes corresponding to the same feature dimension. If there are multiple nodes, the influence of the evaluation results corresponding to each of the multiple nodes is fused to calculate the feature importance corresponding to the feature dimension.
2. The method according to claim 1, characterized in that, The attribute parameters of the raw material crop include parameter values under one or more characteristic dimensions such as crop category, growth cycle, unit yield, and reference water footprint; the geographical environment parameters of the raw material crop include parameter values under one or more characteristic dimensions such as geographical location, precipitation, and sunshine duration; the production environment parameters of the sustainable aviation fuel include parameter values under one or more characteristic dimensions such as raw material transportation distance and energy proportion; wherein, the first parameter group includes parameter values under the crop category and / or the geographical location; the second parameter group includes parameter values under the growth cycle, the unit yield, the reference water footprint, the precipitation, the sunshine duration, the transportation distance, and / or the energy proportion.
3. The method according to claim 1, characterized in that, Also includes: The feature dimensions are ranked according to their importance as statistically determined for different feature dimensions. The importance of the sorted feature dimensions is visualized to show the contribution of different feature dimensions to the water footprint assessment results.
4. The method according to claim 1, characterized in that, The training process of the water footprint assessment model includes: Under the preset evaluation dimensions, feature dimensions are determined to obtain a set of feature dimensions. The evaluation dimensions include one or more dimensions of the raw material crop attributes, the geographical environment of the raw material crop, and / or the production environment of sustainable aviation fuel. For any decision tree to be constructed, select a subset of feature dimensions from the set of feature dimensions. Multiple training samples are extracted from the training sample set to form a training set; The decision tree is constructed based on the selected feature dimension, and the feature dimension serves as a node in the decision tree. Continue to construct other decision trees to form the water footprint assessment model based on the constructed decision trees; Based on the training set, the water footprint assessment model is optimized in terms of model parameters and the decision tree structure is adjusted to complete the training of the water footprint assessment model.
5. The method according to claim 4, characterized in that, Also includes: Multiple training samples are extracted from the training sample set to serve as the test set; The test set is input into the water footprint assessment model to obtain the test results; The water footprint assessment model was validated using cross-validation to optimize its hyperparameters. The validation metrics used in the cross-validation process include mean squared error and / or mean absolute error.
6. The method according to claim 1, characterized in that, The evaluation results generated by the multiple decision trees are aggregated to obtain the final evaluation result, including: The average of the evaluations generated by the multiple decision trees is calculated. The calculated mean value is taken as the final evaluation result.
7. The method according to claim 1, characterized in that, Also includes: The water footprint assessment results corresponding to the sustainable aviation fuel are visualized.
8. A computing device, characterized in that, Includes memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component to execute one or more computer instructions for performing the water footprint assessment method according to any one of claims 1-7.
Citation Information
Patent Citations
Learning condition evaluation method based on course data and terminal equipment
CN115631071A
Intelligent nursing evaluation method, system and equipment based on random forest and entropy weight method
CN118571438A