A Multi-Feature Wind Erosion Modulus Prediction Method and System for Caragana korshinskii Combining Ensemble Learning
By establishing a multi-feature association transmission link for Caragana korshinskii and an integrated learning sub-model set, the problem of underutilization of feature association relationships in the prediction of wind erosion modulus in Caragana korshinskii growing areas was solved, achieving higher accuracy in wind erosion modulus prediction and supporting scientific ecological restoration strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies fail to fully consider the correlation between multiple features of Caragana korshinskii in predicting wind erosion modulus in the growth area, resulting in low prediction accuracy and insufficient generalization ability of a single model when dealing with complex data.
By acquiring a set of multiple features from the growth area of Caragana korshinskii, a multi-feature association and transmission link for Caragana korshinskii is established, an integrated learning sub-model set is constructed, information interaction between different features and model parameter adjustment are realized, and wind erosion modulus prediction results are generated.
It improves the accuracy of wind erosion modulus prediction and the generalization ability of the model, and can more accurately reflect the degree of soil protection provided by Caragana korshinskii vegetation, supporting the formulation of scientific ecological restoration strategies.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a method and system for predicting the multi-feature wind erosion modulus of Caragana korshinskii by combining ensemble learning. Background Technology
[0002] In the fields of ecological protection and desertification control, accurate prediction of the wind erosion modulus in the growing areas of Caragana korshinskii is crucial for assessing the windbreak and sand-fixing effects of Caragana korshinskii vegetation and for formulating scientific and reasonable ecological restoration strategies. The wind erosion modulus reflects the amount of soil eroded by wind per unit area per unit time, and its accurate prediction helps to understand the degree of soil protection provided by Caragana korshinskii vegetation.
[0003] Currently, methods for predicting wind erosion modulus in Caragana korshinskii growing areas have many limitations. Some traditional methods only consider single characteristics of Caragana korshinskii, such as plant height or crown width, while ignoring the combined influence of environmental characteristics (such as soil texture and rainfall) and community distribution characteristics (such as planting density and community structure) on the wind erosion modulus. Other methods, while attempting to consider multiple characteristics, fail to adequately consider the relationships between different characteristics, simply superimposing them, resulting in insufficient utilization of feature information and an inability to accurately reflect the synergistic effect of each characteristic on the wind erosion modulus. Furthermore, existing methods mostly employ single models for prediction. Single models often suffer from insufficient generalization ability and low prediction accuracy when dealing with complex multi-feature data, making it difficult to meet the needs of accurate wind erosion modulus prediction in practical ecological protection work. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for predicting the multi-feature wind erosion modulus of *Caragana korshinskii* by incorporating ensemble learning, the method comprising:
[0005] Acquire a set of multiple features of Caragana korshinskii in the growth area and a record of historical wind erosion modulus corresponding to the growth area. The set of multiple features of Caragana korshinskii includes morphological features, growth environment features, and community distribution features of Caragana korshinskii. The historical wind erosion modulus record corresponds to the acquisition time of the set of multiple features of Caragana korshinskii.
[0006] Based on the correlation between the features in the multi-feature set of Caragana korshinskii, a multi-feature correlation transmission link of Caragana korshinskii is established. The multi-feature correlation transmission link of Caragana korshinskii is used to characterize the information transmission direction and correlation strength between different types of Caragana korshinskii features.
[0007] Based on the multi-feature association and transmission link of Caragana korshinskii, an ensemble of learning sub-models is constructed. Each sub-model in the ensemble of learning sub-models corresponds to an information processing path for a type of Caragana korshinskii feature, and each sub-model achieves information interaction through the multi-feature association and transmission link of Caragana korshinskii.
[0008] The set of multiple features of Caragana korshinskii and the corresponding historical wind erosion modulus records are input into the set of integrated learning sub-models, and the sub-model collaborative training process is executed to obtain the trained Caragana korshinskii multi-feature wind erosion modulus prediction model. The sub-model collaborative training process adjusts the model parameters through information interaction between the sub-models.
[0009] The set of multiple features of the Caragana korshinskii growing area to be predicted is input into the trained Caragana korshinskii multi-feature wind erosion modulus prediction model to generate the wind erosion modulus prediction result of the Caragana korshinskii growing area. The wind erosion modulus prediction result maintains feature correlation and correspondence with the set of multiple features of the Caragana korshinskii in the area to be predicted.
[0010] Furthermore, embodiments of the present invention also provide a multi-feature wind erosion modulus prediction system for Caragana korshinskii combining ensemble learning, characterized in that it includes:
[0011] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described method for predicting the multi-feature wind erosion modulus of *Caragana korshinskii* by incorporating ensemble learning via executing the machine-executable instructions.
[0012] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described method for predicting the multi-feature wind erosion modulus of Caragana korshinskii combined with ensemble learning.
[0013] Based on the above, by acquiring a multi-feature set of Caragana korshinskii trees in the growth area and corresponding historical wind erosion modulus records, this study comprehensively covers the morphological characteristics, growth environment characteristics, and community distribution characteristics of Caragana korshinskii. A multi-feature correlation transmission link for Caragana korshinskii is established based on the correlation relationships between these features, accurately characterizing the information transmission direction and correlation strength between different types of Caragana korshinskii features, and deeply exploring the intrinsic connections between features. Based on this multi-feature correlation transmission link, an ensemble of learning sub-models is constructed, with each sub-model corresponding to an information processing path for a specific type of Caragana korshinskii feature. Information interaction is achieved through the correlation transmission link, fully leveraging the role of different features in prediction. The collaborative training process of the sub-models adjusts model parameters through information interaction between sub-models, effectively improving the overall performance and generalization ability of the model. Finally, the multi-feature set of Caragana korshinskii trees in the area to be predicted is input into the trained model to generate prediction results, achieving accurate prediction of the wind erosion modulus in the Caragana korshinskii growth area. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the execution flow of the method for predicting the multi-feature wind erosion modulus of Caragana korshinskii, which combines ensemble learning, provided in an embodiment of the present invention.
[0015] Figure 2 This is a schematic diagram of exemplary hardware and software components of the Caragana korshinskii multi-feature wind erosion modulus prediction system combined with ensemble learning provided in an embodiment of the present invention. Detailed Implementation
[0016] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for predicting the multi-feature wind erosion modulus of Caragana korshinskii based on ensemble learning, provided in one embodiment of the present invention. The following is a detailed description of this method.
[0017] Step S110: Obtain the multi-feature set of Caragana korshinskii in the growth area and the historical wind erosion modulus record corresponding to the growth area of Caragana korshinskii. The multi-feature set of Caragana korshinskii includes the morphological features of Caragana korshinskii, the growth environment features of Caragana korshinskii, and the distribution features of Caragana korshinskii community. The historical wind erosion modulus record corresponds to the collection time of the multi-feature set of Caragana korshinskii.
[0018] In this embodiment, the aforementioned Caragana korshinskii growing area is a typical Caragana korshinskii planting area in arid and semi-arid regions. Wind erosion is particularly prominent in this area, necessitating the prediction of wind erosion modulus using relevant Caragana korshinskii features. When acquiring the Caragana korshinskii multi-feature set and historical wind erosion modulus records, it is essential to ensure the completeness, timeliness, and correspondence of the data.
[0019] Step S111: Separate the morphological characteristics, growth environment characteristics, and community distribution characteristics of Caragana korshinskii from the set of multiple features of Caragana korshinskii, and determine the specific feature items under each feature type.
[0020] The collected multi-feature sets of *Caragana korshinskii* were classified and processed according to the differences in feature attributes, into three major categories: morphological features, growth environment features, and community distribution features. Specific features under morphological features include plant height, crown width, branch density, and root distribution. Specific features under growth environment features include soil moisture, soil texture, wind speed, and precipitation. Specific features under community distribution features include plant spacing, community coverage, and symbiotic distribution with other plants. After classification, a clear table of correspondence between feature types and features was generated.
[0021] Step S112: For the morphological characteristics of Caragana korshinskii, the plant height is obtained by measuring the vertical distance from the ground to the top of the tree; for the crown width is obtained by measuring the maximum horizontal span of the crown; for the branch density is obtained by counting the number of branches per unit area; and for the root distribution is obtained by measuring the depth and range of the roots in the soil.
[0022] To determine the height characteristics of Caragana korshinskii plants, a portable height measuring instrument was used. At each sampling point, multiple representative Caragana korshinskii plants were selected, and the vertical distance from the base of the plant to the top growth point was measured. The average value of the measurements for each plant was taken as the plant height data. The plant height data of multiple plants constituted the Caragana korshinskii plant height characteristic data for that sampling point. The data was presented in the form of a multi-dimensional array, including the average measurement value and the number of measurements for different plants.
[0023] For the crown width characteristics of Caragana korshinskii, a measuring tape was used to measure the maximum horizontal span of the crown in the east-west and north-south directions. The span data in both directions were used as the crown width characteristic data of the plant. Similarly, after measuring multiple plants, the crown width characteristic data of the sampling points were compiled. The data included the bidirectional span values of each plant and the set of values within the sampling points.
[0024] To obtain the branch density characteristics of Caragana korshinskii, it is necessary to first delineate quadrats of a unit area. The area of the quadrats is determined based on the growth density of Caragana korshinskii. Within each quadrat, the number of Caragana korshinskii branches, including main branches and lateral branches, is counted. After the count is completed, the number of branches per unit area is recorded as the branch density characteristic data of that quadrat. Data from multiple quadrats are integrated into the branch density characteristic data of the sampling point. The data includes the number of branches in each quadrat and statistical quantities such as the mean and standard deviation of the sampling point.
[0025] The data on the root distribution characteristics of Caragana korshinskii were obtained using a layered excavation method. A profile was excavated around the Caragana korshinskii plant, and the distribution of the root system at different soil depths was measured layer by layer. The distribution depth range of the root system was recorded, and the distribution radius of the root system in the horizontal direction was measured at the same time. The root distribution characteristic data were formed by combining the depth and radius data. Multiple plants were selected for each sampling point, and the comprehensive value was obtained. The data is presented in the form of a two-dimensional array of depth and radius.
[0026] Step S113: For the soil moisture characteristics in the growth environment of Caragana korshinskii, characteristic data are obtained by detecting the moisture content in the soil of the growing area; for the soil texture characteristics in the growth environment of Caragana korshinskii, characteristic data are obtained by analyzing the proportion of sand, silt and clay particles in the soil; for the wind speed characteristics in the growth environment of Caragana korshinskii, characteristic data are obtained by recording the air flow speed in the growing area; for the precipitation characteristics in the growth environment of Caragana korshinskii, characteristic data are obtained by statistically analyzing the total precipitation and precipitation frequency in the growing area.
[0027] Soil moisture characteristic data is acquired through soil moisture sensors. Sensors are deployed at different soil depths at each sampling point. The sensors collect soil moisture content data in real time and record the data at fixed time intervals. The data is then summarized to form a soil moisture characteristic data sequence at different depths of the sampling point. The data sequence includes timestamps and corresponding moisture content value arrays for each depth.
[0028] To obtain soil texture characteristics, soil samples must first be collected and brought back to the laboratory for analysis using sieving and hydrometer methods. Sieving separates sand particles of different sizes, while hydrometer methods measure the content of silt and clay particles. The proportions of sand, silt, and clay particles in the soil sample are calculated to obtain soil texture characteristic data for that sampling point. The data are presented as a combination of the proportions of the three particle types.
[0029] Wind speed characteristic data is obtained by anemometers deployed in the growth area. The anemometers record the air flow speed, including instantaneous wind speed and average wind speed, at a set time frequency, and also record the direction of wind speed change. The wind speed data over a period of time is summarized and organized to form a wind speed characteristic data sequence for the area. The data sequence includes a timestamp, instantaneous wind speed value, average wind speed value, and wind direction information array.
[0030] Rainfall characteristic data are collected through rain gauges. The rain gauges record the amount of rainfall for each rainfall event in real time, and calculate the total amount of rainfall within a fixed period. At the same time, the number of rainfall events is recorded to determine the rainfall frequency. The total rainfall and rainfall frequency data are combined to form the rainfall characteristic data of the region. The data includes the amount of rainfall, duration and total amount of rainfall, and frequency statistics for each rainfall event within the period.
[0031] Step S114: For the distribution characteristics of Caragana korshinskii community, the spacing between Caragana korshinskii plants is obtained by measuring the horizontal distance between adjacent Caragana korshinskii plants; for the distribution characteristics of Caragana korshinskii community coverage is obtained by calculating the proportion of the projected area of the Caragana korshinskii community to the total area of the growing area; for the distribution characteristics of Caragana korshinskii community symbiosis with other plants is obtained by statistically analyzing the species and number of other plants symbiotic with Caragana korshinskii in the growing area.
[0032] To obtain the characteristic data of the spacing between Caragana korshinskii plants, multiple transects need to be selected in the sampling area. The horizontal distance between two adjacent Caragana korshinskii plants is measured along the transects. Multiple spacing data are measured for each transect. The spacing data of all transects are summarized, and the average value and standard deviation are calculated to form the characteristic data of the spacing between plants in the area. The data includes the set of spacing values of each transect and the overall statistics.
[0033] The community coverage characteristics of Caragana korshinskii were obtained by combining remote sensing imagery with ground measurements. First, the projected area of the Caragana korshinskii community was extracted using high-resolution remote sensing imagery. Then, the total area of the sampling area was determined by ground measurements. The coverage value was obtained by dividing the projected area by the total area. At the same time, the remote sensing image extraction results were calibrated by combining the measured data from multiple sampling points. Finally, accurate community coverage characteristic data were obtained, which included the measured coverage of each sampling point, the image-extracted coverage, and the calibrated final coverage array.
[0034] Data on the symbiotic distribution characteristics of Caragana korshinskii and other plants were obtained through quadrat sampling. Multiple quadrats were set up in the growth area. In each quadrat, the types of other plants that coexist with Caragana korshinskii were counted, and the name and quantity of each plant were recorded. The proportion of each symbiotic plant in the total number of plants in the quadrat was calculated. The statistical results of all quadrats were summarized to form symbiotic distribution characteristic data, which includes quadrat identifiers, a list of plant species, and a matrix of corresponding quantities and proportions.
[0035] For example, step S115: Assign a unique feature identifier to each feature type and its corresponding feature item, the feature identifier being used to distinguish different feature items.
[0036] Based on different feature types, basic identifier prefixes are assigned to the morphological features of *Caragana korshinskii*, with different basic identifier prefixes assigned to the growth environment features and community distribution features. A unique code for each feature is added after each basic identifier prefix. For example, the identifier for *Caragana korshinskii* plant height is "XTXZ-ZG", the identifier for *Caragana korshinskii* crown width is "XTXZ-GF", the identifier for soil moisture is "XTHJ-TRSD", the identifier for wind speed is "XTHJ-FS", the identifier for *Caragana korshinskii* plant spacing is "XTQL-JJ", and the identifier for *Caragana korshinskii* community coverage is "XTQL-FGD", etc. Each feature identifier is unique and can clearly distinguish different feature items.
[0037] Step S116: Add a collection time identifier to the feature data of each feature item, wherein the collection time identifier is consistent with the collection time identifier of the corresponding historical wind erosion modulus record.
[0038] The data collection time identifier adopts the format of "year-month-day-hour". For example, if the data collection time for a certain feature is 10:00 AM on June 10th of a certain year, the data collection time identifier will be "XXXX-06-10-10". When acquiring feature data for each feature item, the collection time is recorded synchronously and a corresponding time identifier is generated. At the same time, it is ensured that this time identifier is exactly the same as the time identifier of the historical wind erosion modulus record collected at the same time, so as to achieve time alignment between feature data and wind erosion modulus record.
[0039] Step S117: Establish an association mapping table between feature identifiers and acquisition time identifiers. The association mapping table is used to record the feature data corresponding to each feature identifier under different acquisition time identifiers.
[0040] The association mapping table uses feature identifiers as rows and collection time identifiers as columns. Each cell in the table corresponds to the feature data of a certain feature identifier under a specific collection time identifier. For example, in the cell corresponding to the feature identifier "XTXZ-ZG" and the collection time identifier "XXXX-06-10-10", the characteristic data of the height of the Caragana korshinskii measured at that time point is entered. The association mapping table also includes a data quality identifier column to mark the collection status of the feature data, such as "valid", "missing", "abnormal", etc., to facilitate subsequent data filtering.
[0041] Step S118: Sort the feature data under the same feature identifier according to the order of the collection time identifier to form a feature data time series. The feature data time series can reflect the change trend of feature data over time.
[0042] For each feature identifier, all collection time identifiers and corresponding feature data are extracted from the association mapping table, and the data are sorted according to the chronological order of the collection time identifiers. For example, wind speed data for the feature identifier "XTHJ-FS" at different time points are arranged chronologically to form a wind speed feature data time series. In this wind speed feature data time series, each data point contains a collection time identifier and a corresponding feature data value, which can intuitively present the trend of feature data fluctuations, increases, or decreases over time.
[0043] Step S119: Integrate the feature identifier, collection time identifier, feature data, and feature data time series into a unified data structure. The data structure includes a feature type field, a feature item field, a collection time field, a feature data field, and a data time series field.
[0044] The unified data structure is constructed using a structured format. The feature type field is filled with the major category to which the feature belongs, such as "Caragana spp. morphological characteristics," "Caragana spp. growth environment characteristics," or "Caragana spp. community distribution characteristics." The feature item field is filled with the specific name of the feature item and its corresponding feature identifier, such as "Caragana spp. plant height characteristics (XTXZ-ZG)." The collection time field is filled with the collection time identifier corresponding to the feature data. The feature data field is filled with the specific feature data value; if it is multi-dimensional data, it is presented in array form. The data time series field is filled with the feature data time series corresponding to the feature item. For example, a data structure entry might be: Feature type field "Caragana spp. growth environment characteristics," feature item field "wind speed characteristics (XTHJ-FS)," collection time field "XXXX-06-10-10," feature data field "[wind speed data at multiple time points]," and data time series field "[wind speed data series sorted by time]."
[0045] Step S1110: Group and store all data structures according to feature type to form a structured set of multiple features of the Leonurus japonicus. The structured set of multiple features of the Leonurus japonicus can be directly input into the set of ensemble learning sub-models for processing.
[0046] All constructed data structures are divided into three groups based on feature type: the Caragana foliage morphology feature group, the Caragana foliage growth environment feature group, and the Caragana foliage community distribution feature group. Within each group, the data structures are arranged alphabetically by feature name. A group identifier and data volume statistics are added to each group; for example, the Caragana foliage morphology feature group contains four feature items and several data structure entries. After grouping and storage, a structured Caragana foliage multi-feature set is formed. This multi-feature set has a standardized data format and can be directly recognized and read by the ensemble learning sub-model set.
[0047] Step S120: Based on the correlation between features in the set of multiple features of Caragana korshinskii, establish a multi-feature correlation transmission link of Caragana korshinskii. The multi-feature correlation transmission link of Caragana korshinskii is used to characterize the information transmission direction and correlation strength between different types of Caragana korshinskii features.
[0048] In this embodiment, based on the above-mentioned structured set of multiple features of Caragana korshinskii, by analyzing the association types between each feature item, calculating the association parameters, and sorting out the transmission path, an association transmission link that can accurately reflect the information interaction relationship between features is constructed.
[0049] Step S121: Extract specific feature items of morphological features, growth environment features, and community distribution features of Caragana korshinskii from the set of multiple features of Caragana korshinskii, and determine the feature attributes and data representation form of each feature item.
[0050] The structured set of multiple features of *Caragana korshinskii* is traversed, and all specific feature items are extracted from the three feature types. Combining this with information from the data structure, the characteristic attributes of each feature item are clarified. For example, the attributes of *Caragana korshinskii* plant height are "numerical, continuous," the attributes of soil texture are "numerical, proportional," and the attributes of the symbiotic distribution characteristics of *Caragana korshinskii* with other plants are "category + numerical" (including plant species category and quantity). Simultaneously, the data representation format for each feature item is determined. For instance, *Caragana korshinskii* plant height is represented by a numerical value of "length unit," soil texture is represented by a numerical combination of "sand ratio, silt ratio, and clay ratio," and symbiotic distribution characteristics are represented by a combination of "a list of plant species + corresponding quantity."
[0051] Step S122: Analyze the relationship between any two feature items and determine the type of relationship. The type of relationship includes direct relationship and indirect relationship. Direct relationship means that there is a direct influence between the two feature items. Indirect relationship means that the two feature items have an influence through a third feature item.
[0052] For any two features, perform pairwise analysis to determine the type of relationship between them. For example, analyzing the characteristics of Caragana korshinskii plant height and wind speed, wind speed directly affects the growth height of Caragana korshinskii plants, and the two have a direct mutual influence, so they are determined to be directly related. Analyzing the characteristics of Caragana korshinskii branch density and soil moisture, the density of Caragana korshinskii branches affects regional water evaporation, which in turn affects soil moisture. At the same time, soil moisture also affects the growth density of branches. The two affect each other through the intermediate factor of "water evaporation," so they are determined to be indirectly related. As another example, the characteristics of Caragana korshinskii community cover and precipitation, cover directly affects the infiltration and loss of precipitation on the ground surface, so the two are directly related. The characteristics of Caragana korshinskii root distribution and soil texture, root distribution is directly affected by the proportion of sand and clay particles in the soil, so the two are directly related.
[0053] During the analysis, it is necessary to combine the basic theories of botany, ecology, and meteorology, and to refer to the synchronicity of changes in the two characteristic data in the time series of characteristic data to ensure the accuracy of the correlation type determination. For combinations that are difficult to determine directly, supplementary confirmation can be obtained by consulting research literature on the growth of Caragana korshinskii in the region.
[0054] Step S123: For directly related feature items, calculate the correlation influence parameter between feature items. The correlation influence parameter is determined by the synchronicity trend of feature items changing over time.
[0055] We selected directly related feature combinations, such as the plant height and wind speed of Caragana korshinskii, and the community coverage and precipitation of Caragana korshinskii, and extracted the corresponding feature data time series. Taking the plant height and wind speed of Caragana korshinskii as an example, we aligned the two time series according to the collection time markers and analyzed whether the changing trends of wind speed data and plant height data were synchronized within the same time interval.
[0056] When calculating the correlation parameters, first determine the time interval of the time series, divide the data into multiple time periods according to the intervals, and calculate the rate of change of wind speed and plant height characteristics within each time period. The rate of change is calculated as: (data value at the end of the time period - data value at the beginning of the time period) / data value at the beginning of the time period. Then, count the number of times the rate of change of the two characteristics are synchronized. Synchronization includes changes in the same direction (both increasing or both decreasing) and the correlation of the magnitude of change.
[0057] Calculate the synchronization ratio parameter, which is the proportion of synchronization times to the total number of time periods. Then calculate the variation amplitude correlation parameter by comparing the absolute value of the difference between the change rates of the two characteristics in each time period, taking the average of all absolute values of the difference, and subtracting the ratio of this average to the maximum possible difference from 1 to obtain the variation amplitude correlation parameter. The synchronization ratio parameter and the variation amplitude correlation parameter are then weighted and summed according to preset weights to obtain the correlation influence parameter. The value of this correlation influence parameter ranges from 0 to 1; the closer the value is to 1, the stronger the correlation influence between the two characteristics. For example, the synchronization ratio parameter for the Caragana korshinskii community coverage characteristic and precipitation characteristic is close to 1, and the variation amplitude correlation parameter is also close to 1; the weighted sum of the two results in a correlation influence parameter between 0.85 and 1.0. Similarly, the synchronization ratio parameter for the Caragana korshinskii plant height characteristic and wind speed characteristic is between 0.6 and 0.8, and the variation amplitude correlation parameter is between 0.6 and 0.8; the weighted sum of the two results in a correlation influence parameter between 0.6 and 0.8.
[0058] Step S124: For the indirectly related feature items, identify the intermediate related feature items, and determine the transmission order and transmission amplitude parameters of the influence of the intermediate related feature items on the two indirectly related feature items.
[0059] For combinations of indirectly related features, such as the density of Caragana korshinskii branches and soil moisture, intermediate related features are identified through theoretical analysis and observation of data trends. In this combination, branch density affects the surface evaporation rate of the region, and surface evaporation rate directly affects soil moisture; therefore, "surface evaporation rate" is the intermediate related feature. If the Caragana korshinskii multi-feature set does not directly contain this intermediate feature, it is derived from existing feature data, such as using branch density, wind speed, and temperature to estimate surface evaporation rate, forming a derived time series of data for the intermediate feature.
[0060] When determining the order of influence transmission, the time series changes of the three characteristic items are analyzed. Generally, changes in branch density first affect surface evaporation, and changes in surface evaporation then affect soil moisture. Therefore, the transmission order is "Caragana branch density characteristics → surface evaporation characteristics → soil moisture characteristics".
[0061] Determining the transmission amplitude parameter requires calculating the influence amplitude parameter of the preceding feature on the intermediate feature, and the influence amplitude parameter of the intermediate feature on the subsequent feature. The calculation logic is consistent with the calculation method of the directly related influence parameter. Taking the transmission link of Caragana korshinskii branch density characteristic → surface evaporation characteristic → soil moisture characteristic as an example, first calculate the influence parameter of Caragana korshinskii branch density characteristic and surface evaporation characteristic, and use it as the first transmission amplitude parameter; then calculate the influence parameter of surface evaporation characteristic and soil moisture characteristic, and use it as the second transmission amplitude parameter.
[0062] Multiplying the first and second transmission amplitude parameters yields the overall transmission amplitude parameter for this indirect association combination. This overall transmission amplitude parameter also ranges from 0 to 1; a larger value indicates a stronger transmission effect of the indirect association. For example, the first transmission amplitude parameter between the *Caragana korshinskii* branch density characteristic and the surface evaporation characteristic is between 0.7 and 0.9, and the second transmission amplitude parameter between the surface evaporation characteristic and the soil moisture characteristic is between 0.8 and 1.0. The overall transmission amplitude parameter obtained by multiplying these two parameters is between 0.56 and 0.9.
[0063] For example, in the indirect correlation between the symbiotic distribution characteristics of Caragana korshinskii and other plants and wind speed characteristics, the intermediate correlation feature is the overall community roughness characteristic, and the transmission order is "symbiotic distribution characteristics of Caragana korshinskii and other plants → overall community roughness characteristic → wind speed characteristic". The correlation influence parameter between the symbiotic distribution characteristics and the roughness characteristic is calculated to be between 0.8 and 1.0, the correlation influence parameter between the roughness characteristic and the wind speed characteristic is between 0.8 and 1.0, and the overall transmission amplitude parameter is between 0.64 and 1.0.
[0064] Step S125: Construct an initial feature association network based on the type of association, association influence parameters, intermediate association feature items, and influence transmission order. In the initial feature association network, each node corresponds to a feature item, and the connections between nodes correspond to association relationships.
[0065] Each feature item is treated as an independent node, and the feature item name, feature identifier, and feature attribute are labeled on the node. For directly related feature item nodes, solid lines are used to connect them, and the corresponding association influence parameters are labeled on the connecting lines. For indirectly related feature item nodes, dashed lines are used to connect them, and the intermediate related feature items and their corresponding overall transmission amplitude parameters are labeled on the dashed lines in the transmission order. Arrows are used to indicate the direction of information transmission.
[0066] For example, the Caragana chinensis community coverage feature node and the precipitation feature node are connected by a solid line, and the associated influence parameter is labeled with a value between 0.85 and 1.0; the Caragana chinensis branch density feature node and the soil moisture feature node are connected by a dashed line, and the line is labeled "Caragana chinensis branch density feature → surface evaporation feature → soil moisture feature, 0.56 to 0.9", with the arrow pointing from the branch density feature node to the surface evaporation feature node, and then from the surface evaporation feature node to the soil moisture feature node.
[0067] All feature node pairs and their corresponding connections are combined according to the rules described above to form an initial feature association network. In this initial feature association network, the size of the nodes is set according to the number of dimensions of the feature items; the more dimensions, the larger the nodes. The thickness of the connections is set according to the association influence parameter or the overall transmission amplitude parameter; the larger the parameter value, the thicker the connection, which visually presents the differences in the association strength between feature items.
[0068] Step S126: The initial feature association network is sorted out for information transmission paths. Paths whose information transmission efficiency parameters meet the association requirements are retained, and paths whose information transmission efficiency parameters do not meet the association requirements are deleted, forming a multi-feature association transmission link.
[0069] First, we define how to calculate the information transmission efficiency parameter. For directly related paths, the information transmission efficiency parameter is equal to the association impact parameter of that path; for indirectly related paths, the information transmission efficiency parameter is equal to the overall transmission amplitude parameter of that path. We then set an efficiency threshold for the association requirement. This threshold is determined based on the distribution of information transmission efficiency parameters across all paths, typically selecting the median of the distribution as the threshold.
[0070] Traverse all associated paths in the initial feature association network, extract the information transmission efficiency parameter for each path, and compare it with a set efficiency threshold. If the parameter is greater than or equal to the threshold, the path is deemed to meet the association requirements and is retained; if the parameter is less than the threshold, the path is deemed not to meet the association requirements and is deleted.
[0071] For example, if the efficiency threshold is set to 0.5, and the correlation impact parameter of a directly related path is 0.4, which is less than the threshold, then the path is deleted; if the overall transmission amplitude parameter of an indirectly related path is 0.56, which is greater than the threshold, then the path is retained. During the analysis process, if a feature node becomes isolated (without any connections) due to path deletion, then the isolated node is removed from the network to avoid invalid nodes occupying resources.
[0072] After the analysis, the remaining nodes and connections form a multi-feature association and transmission link for *Caragana korshinskii*. In this link, direct and indirect association paths are clearly distinguished, and the direction and intensity of information transmission are well-defined, accurately reflecting the effective information interaction relationships between different types of *Caragana korshinskii* characteristics. For example, the final link includes a direct association path between plant height and crown width within the morphological characteristics of *Caragana korshinskii* (association influence parameter 0.7 to 0.9), a direct association path between precipitation and soil moisture within the growth environment characteristics of *Caragana korshinskii* (association influence parameter 0.8 to 1.0), and an indirect association path between branch density and soil moisture between the morphological characteristics and growth environment characteristics of *Caragana korshinskii* (transmission amplitude parameter 0.56 to 0.9), etc.
[0073] Step S130: Based on the multi-feature association transmission link of Caragana korshinskii, construct an integrated learning sub-model set. Each sub-model in the integrated learning sub-model set corresponds to an information processing path for a Caragana korshinskii feature type, and each sub-model achieves information interaction through the multi-feature association transmission link of Caragana korshinskii.
[0074] In this embodiment, based on the above-mentioned multi-feature association and transmission link of Caragana korshinskii, the processing modules are divided according to feature type, a basic prediction sub-model is configured, an information interaction channel and a result fusion module are established, and finally integrated to form an ensemble of learning sub-models.
[0075] Step S131: Based on the type of feature items in the Caragana korshinskii multi-feature association transmission link, divide the feature processing modules. Each feature processing module is responsible for processing one type of Caragana korshinskii feature. The feature processing modules include a Caragana korshinskii morphological feature processing module, a Caragana korshinskii growth environment feature processing module, and a Caragana korshinskii community distribution feature processing module.
[0076] All feature nodes in the multi-feature association transmission link of *Caragana korshinskii* were analyzed and categorized into three feature processing modules according to feature type. The *Caragana korshinskii* morphological feature processing module handles plant height, crown width, branch density, and root distribution characteristics; the *Caragana korshinskii* growth environment feature processing module handles soil moisture, soil texture, wind speed, and precipitation characteristics; and the *Caragana korshinskii* community distribution feature processing module handles plant spacing, community coverage, and symbiotic distribution characteristics with other plants.
[0077] Each feature processing module includes a data receiving unit, a data cleaning unit, a feature standardization unit, and a feature filtering unit. The data receiving unit receives feature data of the corresponding type of *Caragana korshinskii*. The data cleaning unit removes outliers and missing values from the data. Outliers are determined by their deviation from the mean, and missing values are filled with the mean of the same type of feature. The feature standardization unit converts feature data of different dimensions to the same numerical range. The conversion method is (feature data value - minimum feature value) / (maximum feature value - minimum feature value). The feature filtering unit filters key features based on the correlation between the feature and the wind erosion modulus.
[0078] For example, after the Caragana morphology feature processing module receives data containing features such as plant height and crown width, the data cleaning unit removes outliers in plant height, the feature standardization unit standardizes data such as plant height (length unit) and branch density (number unit) to the range of 0 to 1, and the feature screening unit retains feature items that are highly correlated with wind erosion modulus.
[0079] Step S132: Configure a corresponding basic prediction sub-model for each feature processing module. The input of the basic prediction sub-model is the Caragana korshinskii feature data processed by the corresponding feature processing module, and the output is the preliminary wind erosion modulus prediction value.
[0080] A gradient boosting tree-based basic prediction sub-model is configured for the morphological feature processing module of *Caragana korshinskii*. This basic prediction sub-model includes an input layer, multiple decision tree hidden layers, and an output layer. The input layer receives standardized morphological feature data output from the feature processing module, with each feature corresponding to an input neuron. The decision tree hidden layers perform hierarchical processing of the feature data through node splitting rules, which are determined based on the information gain of the features. The output layer outputs preliminary wind erosion modulus predictions, which are multi-dimensional arrays containing prediction results for different confidence intervals.
[0081] A basic prediction sub-model based on a recurrent neural network is configured for the environmental feature processing module of Caragana korshinskii. This basic prediction sub-model includes an input layer, a recurrent hidden layer, a fully connected layer, and an output layer. The input layer receives standardized environmental feature time series data; the recurrent hidden layer retains historical feature information through memory units to capture the changing patterns of features over time; the fully connected layer integrates the output of the recurrent hidden layer; and the output layer outputs preliminary wind erosion modulus prediction values, which also include multi-dimensional confidence interval data.
[0082] A basic prediction sub-model based on a convolutional neural network is configured for the Caragana korshinskii community distribution feature processing module. This basic prediction sub-model includes an input layer, a convolutional layer, a pooling layer, and an output layer. The input layer receives the standardized community distribution feature matrix data; the convolutional layer extracts the local correlation patterns of features through convolutional kernels, with the kernel size set according to the feature dimension; the pooling layer downsamples the convolution results, retaining key feature information; and the output layer outputs preliminary wind erosion modulus prediction values, which are multi-dimensional numerical sets.
[0083] Each basic prediction sub-model contains a model parameter initialization unit. The initialization unit sets initial parameter values based on the number of feature dimensions to ensure that the model can receive feature data of the corresponding dimension in its initial state.
[0084] Step S133: Based on the information transmission direction in the multi-feature association transmission link of the Caragana korshinskii, establish an information interaction channel between each basic prediction sub-model. The information interaction channel is used to realize the transmission of feature information between different basic prediction sub-models.
[0085] Based on the arrow directions in the multi-feature association transmission chain of *Caragana korshinskii*, the information transmission direction between each basic prediction sub-model is determined. For example, if there is information transmission from *Caragana korshinskii* morphological features to *Caragana korshinskii* growth environment features, a one-way information interaction channel is established between the gradient boosting tree sub-model (morphological features) and the recurrent neural network sub-model (environmental features); if there is bidirectional information transmission from *Caragana korshinskii* growth environment features to *Caragana korshinskii* community distribution features, a bidirectional information interaction channel is established between the recurrent neural network sub-model and the convolutional neural network sub-model.
[0086] Each information exchange channel includes a data encoding unit, a data transmission unit, and a data decoding unit. The data encoding unit converts the feature data and preliminary prediction values of the sending sub-model into interactive data in a unified format. The encoding method is determined based on the mapping relationship between feature identifiers and data values. The data transmission unit adopts a queue-based transmission mechanism to transmit interactive data in chronological order, ensuring the orderly transmission of data. The data decoding unit restores the received interactive data to the original feature data and prediction value format for use by the receiving sub-model.
[0087] For example, when the gradient boosting tree sub-model transmits information to the recurrent neural network sub-model, the encoding unit encodes the plant height feature data, branch density feature data, and corresponding preliminary prediction values into a unified format containing feature identifiers, data values, and prediction values. The transmission unit transmits the encoded data in the order of acquisition time, and the decoding unit restores the received encoded data into feature data and prediction values that can be recognized by the recurrent neural network sub-model.
[0088] Step S134: Configure information transmission weights for each information interaction channel. The information transmission weights are determined based on the correlation strength of the corresponding correlation relationships in the multi-feature correlation transmission link of the *Leymus chinensis*.
[0089] For each information exchange channel, information transmission weights are configured through steps such as extracting correlation strength information, analyzing influencing factors, and normalization processing. The specific process is as follows.
[0090] Step S1341: Extract the association strength description information between the two feature items associated with the corresponding information interaction channel from the multi-feature association transmission link of the Caragana korshinskii.
[0091] Each information interaction channel corresponds to an association path in the multi-feature association transmission chain of *Caragana korshinskii*, and the association strength description information of this path is extracted. For information interaction channels corresponding to direct association paths, the association strength description information is the association influence parameter of this path; for information interaction channels corresponding to indirect association paths, the association strength description information is the overall transmission amplitude parameter of this path. For example, the information interaction channel from the morphological feature sub-model to the environmental feature sub-model corresponds to the indirect association path of branch density-soil moisture, and the overall transmission amplitude parameter of this path is extracted as a value between 0.56 and 0.9 as the association strength description information.
[0092] Step S1342: Analyze the duration and range of influence between feature items in the correlation strength description information. The duration of influence refers to the time span during which a change in one feature item affects another feature item, and the range of influence refers to the number of feature dimensions that affect another feature item.
[0093] Extract the time series of feature data for the two feature items in the corresponding association path and analyze the duration of the impact. Taking the association between branch density and soil moisture as an example, record the time point when branch density changes significantly, and calculate the time span from that time point to the corresponding change in soil moisture. This span is taken as the duration of the impact. To determine the scope of impact, count the number of dimensions of the soil moisture feature affected by changes in branch density. For example, if soil moisture includes three dimensions—surface, middle, and deep layers—and all three dimensions are affected, then the scope of impact is three dimensions.
[0094] For example, in the direct correlation between precipitation characteristics and community cover characteristics, the time span of cover change after precipitation change is the duration of influence, the cover characteristics include numerical dimensions of multiple sampling points, and the number of affected dimensions is the range of influence.
[0095] Step S1343: Normalize the duration and range of the impact respectively. Based on the normalized duration and range parameters, determine the quantitative basis for the correlation strength. The quantitative basis is used to convert the correlation strength description information into calculable correlation parameters.
[0096] When normalizing the duration of the impact, first determine the maximum and minimum values of the impact duration corresponding to all information interaction channels. Then, use the calculation method of (impact duration of a channel - minimum value) / (maximum value - minimum value) to convert the impact duration into an impact duration parameter between 0 and 1. When normalizing the scope of the impact, similarly determine the maximum and minimum values of the scope of the impact corresponding to all information interaction channels, and convert the scope of the impact into an scope parameter between 0 and 1 using the same calculation method.
[0097] The normalized duration of influence parameter and the range of influence parameter are weighted and summed according to a preset ratio to obtain the quantitative basis of the correlation strength. In the preset ratio, the weights of the duration of influence parameter and the range of influence parameter are set according to the importance of the feature type. For example, in the interaction channel between morphological features and environmental features, the weight of duration of influence is higher than that of range of influence.
[0098] Step S1344: Based on the quantification criteria, the correlation strength description information is converted into initial correlation parameters, which can reflect the relative magnitude of the correlation strength between the two feature items.
[0099] The initial correlation parameter is obtained by multiplying the parameter value in the correlation strength description information with the quantization basis. For example, if the overall transmission amplitude parameter in the correlation strength description information is a value between 0.56 and 0.9, and the quantization basis is a value between 0.6 and 0.8, the initial correlation parameter after multiplying the two is a value between 0.336 and 0.72. The initial correlation parameter remains within the range of 0 to 1, with a larger value indicating a relatively higher correlation strength.
[0100] Step S1345: Refer to the prediction error feedback information corresponding to this information interaction channel in the historical training data to calibrate the initial correlation parameters and obtain the adjusted correlation parameters.
[0101] Extract the prediction error data of this information interaction channel from historical training data across different training rounds, and calculate the correlation between the prediction error and the initial correlation parameter. If the prediction error decreases when the initial correlation parameter increases, then calibrate in the direction of positive correlation; if the prediction error increases when the initial correlation parameter increases, then calibrate in the direction of negative correlation.
[0102] The calibration method is as follows: calculate the rate of change of the prediction error, and adjust the initial correlation parameter according to the sign and magnitude of the rate of change. For example, if the initial correlation parameter is 0.5 and the corresponding rate of change of the prediction error is negative (the error decreases as the parameter increases), then the initial correlation parameter is increased by 0.1, resulting in an adjusted correlation parameter of 0.6; if the rate of change of the prediction error is positive, then the initial correlation parameter is decreased by 0.1, resulting in an adjusted correlation parameter of 0.4.
[0103] Step S1346: Map the adjusted association parameters to information transmission weights. The value of the information transmission weights is positively correlated with the adjusted association parameters, thus completing the information transmission weight configuration for each information interaction channel.
[0104] A mapping relationship is established between the adjusted correlation parameters and the information transmission weights. The mapping rule is that the information transmission weight equals the adjusted correlation parameter divided by the sum of the adjusted correlation parameters of all information interaction channels. For example, if the adjusted correlation parameter of a certain information interaction channel is 0.6, and the sum of the adjusted correlation parameters of all channels is 3.0, then the information transmission weight of that channel is 0.2. This mapping method ensures that the sum of the information transmission weights of all information interaction channels is 1, satisfying the basic requirements of weight allocation.
[0105] Once configured, each information exchange channel is associated with a unique information transmission weight, which is dynamically adjusted based on error feedback during subsequent model training.
[0106] Step S135: Set a result fusion module at the output end of each basic prediction sub-model. The result fusion module is used to receive the preliminary wind erosion modulus prediction values output by all basic prediction sub-models and allocate fusion weights according to the prediction accuracy of each basic prediction sub-model.
[0107] The results fusion module includes a prediction accuracy evaluation unit, a fusion weight allocation unit, and a weighted fusion unit. The prediction accuracy evaluation unit is used to calculate the prediction accuracy of each basic prediction sub-model on historical data. The calculation method is to subtract the average absolute error between the predicted value and the true value from 1 and divide it by the range of the true value. The resulting accuracy value ranges from 0 to 1.
[0108] The fusion weight allocation unit assigns fusion weights based on the prediction accuracy values. The allocation rule is that the fusion weight of a sub-model is equal to the accuracy value of that sub-model divided by the sum of the accuracy values of all sub-models. For example, if the accuracy value of the gradient boosting tree sub-model is 0.8, the accuracy value of the recurrent neural network sub-model is 0.9, and the accuracy value of the convolutional neural network sub-model is 0.8, and the sum of the three is 2.5, then the corresponding fusion weights are 0.32, 0.36, and 0.32, respectively.
[0109] The weighted fusion unit receives the preliminary wind erosion modulus predictions from each sub-model, multiplies each prediction value by its corresponding fusion weight, and then concatenates them to form a multi-dimensional fusion prediction result. The concatenation method is arranged in order of sub-model type, and the confidence interval information of each sub-model prediction value is retained to ensure the integrity of the fusion result.
[0110] Step S136: Integrate the feature processing module, the basic prediction sub-model, the information interaction channel, and the result fusion module to form an integrated learning sub-model set. Each sub-model in the integrated learning sub-model set keeps information synchronized through the information interaction channel.
[0111] The three feature processing modules, three basic prediction sub-models, all information interaction channels, and the result fusion module are connected and integrated according to the data flow. The output of the feature processing module is connected to the input of the corresponding basic prediction sub-model. The basic prediction sub-models are connected bidirectionally or unidirectionally through information interaction channels, and the output of each basic prediction sub-model is connected to the input of the result fusion module.
[0112] During the integration process, a data synchronization control unit is set up. This unit synchronizes the data streams of each module based on the acquisition time stamp, ensuring that feature data at the same point in time is processed synchronously in each sub-model and interaction channel. Simultaneously, a module status monitoring unit is configured to monitor the operational status of each module in real time, such as data processing speed and parameter stability. When a module malfunctions, a backup module switching mechanism is triggered to ensure the stable operation of the integrated learning sub-model set.
[0113] The resulting ensemble of integrated learning sub-models can receive a structured set of multiple features of *Caragana korshinskii*, and output preliminary wind erosion modulus fusion prediction results through collaborative processing of internal modules. Furthermore, the sub-models communicate with each other in real time via information exchange channels to ensure information synchronization and collaborative processing. For example, the *Caragana korshinskii* morphological feature processing module inputs selected plant height, crown width, and other feature data into the corresponding basic prediction sub-model. After generating preliminary prediction values, this basic prediction sub-model transmits the prediction values and corresponding feature data to the basic prediction sub-model corresponding to the *Caragana korshinskii* growth environment features through the information exchange channel. Simultaneously, it receives environmental feature processing results and prediction data from this sub-model, and both maintain a consistent data processing rhythm based on the time stamp of the synchronization control unit.
[0114] Step S140: Input the set of multiple features of Caragana korshinskii and the corresponding historical wind erosion modulus records into the set of integrated learning sub-models, execute the sub-model collaborative training process, and obtain the trained Caragana korshinskii multi-feature wind erosion modulus prediction model. The sub-model collaborative training process adjusts the model parameters through information interaction between the sub-models.
[0115] In this embodiment, the structured multi-feature set of Caragana korshinskii growing areas in arid and semi-arid regions and the corresponding historical wind erosion modulus records are used as training data. The data is input into the above-mentioned ensemble learning sub-model set. Through multiple rounds of feature selection, prediction, interaction and parameter adjustment, the sub-models are trained collaboratively, and finally a convergent and stable prediction model is obtained.
[0116] Step S141: Assign the Caragana multi-feature set to the corresponding feature processing modules in the integrated learning sub-model set according to the feature type. Each feature processing module performs feature filtering processing on the input Caragana feature data and retains the feature items that meet the requirements of correlation with wind erosion modulus.
[0117] The morphological feature data of Caragana korshinskii in the structured multi-feature set are assigned to the Caragana korshinskii morphological feature processing module, the Caragana korshinskii growth environment feature data are assigned to the Caragana korshinskii growth environment feature processing module, and the Caragana korshinskii community distribution feature data are assigned to the Caragana korshinskii community distribution feature processing module.
[0118] Each feature processing module contains a feature correlation calculation unit. For the input feature data, combined with the corresponding historical wind erosion modulus records, it calculates the correlation between each feature and the wind erosion modulus. The calculation logic is as follows: extract the feature data time series of the feature item and the historical wind erosion modulus time series of the same period, calculate the synchronization ratio and magnitude of their changing trends, and then sum them according to preset weights to obtain the correlation parameter. This correlation parameter ranges from 0 to 1.
[0119] A correlation threshold is set, which is determined based on the distribution of correlation parameters for all feature items, typically selecting the lower limit of the upper interval of the distribution. Each feature processing module retains feature items whose correlation parameters reach the threshold and removes those that do not. For example, in the Caragana korshinskii morphological feature processing module, the correlation parameters between Caragana korshinskii crown width, branch density, and wind erosion modulus reach the threshold and are retained; in the Caragana korshinskii growth environment feature processing module, the correlation parameters between soil moisture and wind speed reach the threshold and are retained; in the Caragana korshinskii community distribution feature processing module, the correlation parameters between community coverage and plant spacing reach the threshold and are retained.
[0120] Step S142: Each feature processing module inputs the filtered Caragana korshinskii feature data into the corresponding basic prediction sub-model to generate the preliminary wind erosion modulus prediction value of each basic prediction sub-model.
[0121] The Caragana morphology feature processing module inputs the retained crown width, branch density and other feature data into the corresponding basic prediction sub-model in time series order. The basic prediction sub-model processes the feature data through internal neural network layers. The first layer is the input layer that receives feature data, the second layer is the hidden layer that performs dimensional transformation and feature fusion on the feature data, and the third layer is the output layer that generates preliminary wind erosion modulus prediction values. The prediction values are presented in the form of a multi-dimensional array, which includes the predicted value at each time point and the corresponding confidence interval.
[0122] Similarly, the Caragana chinensis growth environment feature processing module inputs the selected soil moisture, wind speed and other feature data into the corresponding basic prediction sub-model, and the Caragana chinensis community distribution feature processing module inputs the selected community coverage, plant spacing and other feature data into the corresponding basic prediction sub-model. Both sub-models generate their respective preliminary wind erosion modulus prediction values according to the above neural network layer processing logic, ensuring that the output format is consistent with the basic prediction sub-model corresponding to the Caragana chinensis morphological features.
[0123] Step S143: Through a preset information interaction channel, the preliminary wind erosion modulus prediction value and corresponding feature data of each basic prediction sub-model are transmitted to other associated basic prediction sub-models to realize information interaction between sub-models.
[0124] After each basic prediction sub-model generates an initial prediction value, it triggers the information exchange channel transmission mechanism. For example, the basic prediction sub-model corresponding to the morphological characteristics of Caragana korshinskii transmits its initial wind erosion modulus prediction value, filtered feature data, and feature correlation parameters to the basic prediction sub-model corresponding to the growth environment characteristics of Caragana korshinskii through a two-way information exchange channel; at the same time, this basic prediction sub-model transmits the above information to the basic prediction sub-model corresponding to the community distribution characteristics of Caragana korshinskii through a one-way information exchange channel.
[0125] During information transmission, the information interaction channel marks the transmitted data according to preset information transmission weights. Information with higher weight values has a higher processing priority at the receiving end. For example, when the basic prediction sub-model corresponding to the growth environment characteristics of *Caragana korshinskii* receives information from the morphological feature sub-model, because the correlation between the two is high and the information transmission weight is large, this information is preferentially included in the subsequent prediction correction process. The receiving sub-model parses the received information, extracts the predicted values, feature data, and weight identifiers, and stores them in a temporary data buffer for further processing.
[0126] Step S144: After receiving the information transmitted by the associated sub-model, each basic prediction sub-model calculates the prediction deviation value by combining its own preliminary wind erosion modulus prediction value and the corresponding historical wind erosion modulus record.
[0127] Each basic prediction sub-model extracts information passed from associated sub-models from a temporary data cache and integrates it with its own generated preliminary wind erosion modulus predictions. Taking the basic prediction sub-model corresponding to the growth environment characteristics of Caragana korshinskii as an example, this basic prediction sub-model extracts preliminary predictions from the morphological feature sub-model, as well as crown width and branch density feature data, and combines them with its own generated preliminary predictions based on soil moisture and wind speed to form a comprehensive prediction reference dataset.
[0128] The system retrieves the actual wind erosion modulus values corresponding to the current processing time from historical wind erosion modulus records. It then compares each preliminary prediction value in the comprehensive prediction reference dataset with the actual values to calculate the prediction deviation. The calculation logic is as follows: For each time point, the absolute value of the difference between each preliminary prediction value and the actual wind erosion modulus value is calculated. This difference is then weighted and summed using the information transfer weight or model confidence level corresponding to each prediction value to obtain the prediction deviation value for that time point. Finally, the prediction deviation values for all time points are averaged to obtain the overall prediction deviation value of the basic prediction sub-model. This overall prediction deviation value is a non-negative, multi-dimensional numerical set containing deviation statistics for different time intervals.
[0129] Step S145: Based on the prediction deviation value, adjust the internal parameters of the basic prediction sub-model and the information transmission weights of the corresponding information interaction channels to reduce the prediction deviation value.
[0130] For the overall prediction deviation of each basic prediction sub-model, initiate the parameter and weight adjustment process to ensure that the deviation is gradually reduced to a reasonable range.
[0131] Step S1451: Analyze the source of the prediction deviation value to determine whether the deviation originates from the internal parameter settings of the basic prediction sub-model or the information transmission weight settings of the information interaction channel.
[0132] Construct a deviation source analysis matrix, where rows correspond to the internal parameter types of the basic prediction sub-model, and columns correspond to the identifiers of associated information interaction channels. Calculate the sensitivity coefficient of each internal parameter change on the prediction deviation value; this sensitivity coefficient is obtained by observing the magnitude of the deviation value change after adjusting the parameter values. Simultaneously, calculate the influence coefficient of each information interaction channel's information transmission weight change on the prediction deviation value; this influence coefficient is obtained by observing the magnitude of the deviation value change after adjusting the weight values.
[0133] By comparing the magnitudes of the sensitivity coefficient and the influence coefficient, if the sensitivity coefficient of a certain internal parameter is greater than all the influence coefficients, the judgment bias mainly comes from the setting of that internal parameter; if the influence coefficient of a certain information interaction channel is greater than all the sensitivity coefficients, the judgment bias mainly comes from the information transmission weight setting of that channel; if there are multiple cases where both the sensitivity coefficient and the influence coefficient are large, the judgment bias is caused by both.
[0134] Step S1452: If the deviation originates from the internal parameter settings of the basic prediction sub-model, extract the internal parameters in the basic prediction sub-model that are directly related to the prediction of wind erosion modulus, and calculate the contribution of each internal parameter to the prediction deviation value. The contribution refers to the magnitude of the change in the prediction deviation value caused by the change of a single internal parameter.
[0135] Internal parameters directly related to wind erosion modulus prediction, such as weight parameters of the hidden layers of the neural network and slope parameters of the activation function, are extracted from the basic prediction sub-model, and several categories of core parameters are selected. For each category of parameter, its value is gradually changed within a preset adjustment range, and the change in prediction deviation value after each adjustment is recorded. The contribution of the parameter is obtained by dividing the absolute value of the change by the absolute value of the parameter adjustment.
[0136] The contribution calculation needs to be repeated multiple times to reduce random errors, and the final contribution of the parameter is the average of the multiple calculations. For example, in the basic prediction sub-model corresponding to the morphological features of *Caragana korshinskii*, adjusting a certain weight parameter in the hidden layer causes a large change in the prediction deviation value, and its contribution value is high; while adjusting a certain parameter of the activation function has a small impact on the deviation value, and its contribution value is low.
[0137] Step S1453: Based on the contribution, adjust the internal parameters in a direction that reduces the prediction deviation value to obtain the adjusted internal parameters.
[0138] The internal parameters are sorted from highest to lowest contribution, with priority given to adjusting parameters with higher contributions. For example, for the hidden layer weight parameter with the highest contribution, if increasing the parameter value reduces the prediction bias, the parameter value is gradually increased until the bias no longer decreases significantly; if increasing the parameter value causes the bias to increase, the parameter value is gradually decreased until the bias reaches a local minimum.
[0139] Perform the above adjustment operation on each type of parameter, monitoring the trend of prediction deviation values in real time during the adjustment process to avoid overfitting the model due to excessive parameter adjustment. After the adjustment is completed, record the new values of all internal parameters to form the adjusted internal parameter set.
[0140] Step S1454: If the deviation originates from the information transmission weight setting of the information interaction channel, extract the correlation strength description information and historical adjustment records corresponding to the information interaction channel, and analyze the gap between the current information transmission weight and the optimal weight.
[0141] We extract the correlation strength description information such as the feature item association influence parameter and transmission amplitude parameter corresponding to the information interaction channel, and combine it with the adjustment record of the information transmission weight of the channel during historical training to construct a weight change trend curve. The curve is plotted with the training round as the horizontal axis and the information transmission weight value as the vertical axis, and the change of prediction deviation value after each round of adjustment is also marked.
[0142] By analyzing the curve, the difference between the current weight value and the historical best weight value (i.e. the weight value with the smallest deviation) is calculated, and the absolute value of the difference and the difference ratio are calculated. The difference ratio is the ratio of the absolute value of the difference to the historical best weight value, which is used to quantify the size of the difference.
[0143] Step S1455: Based on the gap, adjust the value of the information transmission weight to make the adjusted information transmission weight closer to the optimal weight, thereby reducing the prediction deviation value; at the same time, after adjusting the internal parameters of the basic prediction sub-model and the information transmission weight of the information interaction channel, recalculate the prediction deviation value and determine whether the adjusted prediction deviation value has decreased.
[0144] If the current weight value is lower than the historical best weight value and the difference is significant, the weight value will be increased by a preset step size; if the current weight value is higher than the historical best weight value and the difference is significant, the weight value will be decreased by a preset step size. The step size is determined based on the difference ratio; the larger the difference ratio, the larger the step size, in order to speed up the weight adjustment.
[0145] After adjusting the information transmission weights, and combining them with the adjusted internal parameters of the basic prediction sub-model, the feature selection, preliminary prediction, information interaction, and deviation calculation steps are re-executed to obtain a new prediction deviation value. The new deviation value is then compared with the deviation value before adjustment to determine whether the deviation value has decreased.
[0146] Step S1456: If the adjusted prediction deviation value decreases, retain the adjusted internal parameters and information transmission weights; if the adjusted prediction deviation value increases, restore the internal parameters of the basic prediction sub-model and the information transmission weights of the information interaction channel before adjustment, and reselect the adjustment direction.
[0147] If the new prediction bias value is significantly reduced compared to before the adjustment, and the reduction exceeds the preset threshold, the adjustment is confirmed to be effective. The adjusted internal parameters and information transmission weights are retained and used as the initial parameters and weights for the next round of training.
[0148] If the new prediction deviation value is larger than before the adjustment, or the reduction does not reach the preset threshold, the adjustment is deemed invalid, triggering the parameter and weight restoration mechanism. This restores the internal parameters of the basic prediction sub-model and the information transmission weights of the information interaction channels to their pre-adjustment values. Simultaneously, the source of the deviation is re-analyzed, and the adjustment step size or direction of the parameters and weights is adjusted. For example, the step size is reduced, and the adjustment is attempted again to ensure more accurate subsequent adjustments.
[0149] Step S146: Repeat the process of feature screening, generating preliminary wind erosion modulus predictions, information interaction, calculating prediction deviations, and adjusting parameter weights until the prediction deviations stabilize within the preset range.
[0150] Initiate a cyclic training mechanism. After each round of feature selection, prediction, interaction, bias calculation, and parameter weight adjustment, record the current prediction bias value and the corresponding parameter and weight values. Calculate the fluctuation range of the prediction bias value over several consecutive training rounds. The fluctuation range is the ratio of the maximum difference to the average value of the bias values in consecutive rounds.
[0151] If the fluctuation amplitude is less than the preset stability threshold and the prediction deviation is lower than the preset upper limit of deviation, the prediction deviation is determined to be stable within the preset range. If the fluctuation amplitude is greater than the stability threshold or the deviation is higher than the upper limit of deviation, the process returns to step S141 to re-execute the relevant operations and continue the next round of training and adjustment. During the iterative training process, after several rounds of training, the model is tested on a validation set to avoid overfitting. The validation set data consists of the multi-feature set of *Caragana korshinskii* that did not participate in the training and the corresponding historical wind erosion modulus records.
[0152] Step S147: When the prediction deviation value stabilizes within the preset range, stop the sub-model collaborative training process and determine the ensemble of learning sub-models at this time as the trained Caragana korshinskii multi-feature wind erosion modulus prediction model.
[0153] When the fluctuation range of the prediction deviation value is less than the stable threshold for several consecutive rounds during the cyclic training process, and the deviation value of the validation set test also meets the requirements, the training stop instruction is triggered to terminate the sub-model co-training process.
[0154] All module parameters, information interaction channel weights, and result fusion module weights of the ensemble learning sub-model set are saved to generate a model parameter configuration file. This configuration file includes parameter values for each module, data processing rules, information transmission weights, and synchronization control strategies. The saved ensemble learning sub-model set is then designated as the trained Caragana korshinskii multi-feature wind erosion modulus prediction model. This Caragana korshinskii multi-feature wind erosion modulus prediction model can be directly used for wind erosion modulus prediction in the area to be predicted.
[0155] Step S150: Input the set of multiple features of the Caragana korshinskii growing area to be predicted into the trained Caragana korshinskii multi-feature wind erosion modulus prediction method model to generate the wind erosion modulus prediction result of the Caragana korshinskii growing area. The wind erosion modulus prediction result maintains feature correlation and correspondence with the set of multiple features of the Caragana korshinskii growing area to be predicted.
[0156] In this embodiment, the region to be predicted as the growing area of Caragana korshinskii is another arid or semi-arid Caragana korshinskii planting area with similar climate and topography to the training area. The set of multiple features of Caragana korshinskii in this region is input into the trained model, and the prediction result is generated through internal collaborative processing of the model, ensuring that the result corresponds to the input features.
[0157] Step S151: Collect a set of multiple features of the Caragana korshinskii growing area to be predicted, ensuring that the set of multiple features of the Caragana korshinskii includes morphological features, growth environment features, and community distribution features, and that the number of features is consistent with the set of multiple features of the Caragana korshinskii during training.
[0158] Using the same method as the training data collection, multiple sampling points were set up in the growth area of Caragana korshinskii to be predicted. Data on the morphological characteristics of Caragana korshinskii (plant height, crown width, branch density, root distribution), the growth environment characteristics of Caragana korshinskii (soil moisture, soil texture, wind speed, precipitation), and the community distribution characteristics of Caragana korshinskii (plant spacing, community coverage, symbiotic distribution) were collected at each sampling point.
[0159] During the data collection process, data was acquired strictly according to the feature definitions and data representation formats used during training, ensuring that the number of feature terms and data dimensions were completely consistent with the Caragana korshinskii multi-feature set used during training. For example, soil texture features were collected along with the proportions of sand, silt, and clay particles, and wind speed features were recorded along with instantaneous wind speed, average wind speed, and wind direction information, to avoid model processing anomalies due to missing feature terms or inconsistent formats.
[0160] Step S152: Input the set of multiple features of Caragana korshinskii to be predicted into the feature processing module of the trained Caragana korshinskii multi-feature wind erosion modulus prediction model. Each feature processing module filters the input feature data to be predicted according to the filtering rules during training. The filtered feature data to be predicted is input into the corresponding basic prediction sub-model. Each basic prediction sub-model generates the preliminary wind erosion modulus prediction value of the area to be predicted based on the trained internal parameters.
[0161] The set of multiple features of the Caragana korshinskii to be predicted is assigned to the corresponding feature processing modules according to feature type. Each module calls the feature filtering rules saved during training, namely the correlation threshold and filtering logic, to filter the input feature data to be predicted. For example, the Caragana korshinskii morphological feature processing module retains the crown width and branch density feature data in the data to be predicted that reach the correlation parameter threshold, and removes other morphological feature data that do not reach the threshold.
[0162] The filtered feature data to be predicted is input into the corresponding basic prediction sub-model according to the time series. The sub-model loads the trained internal parameters (such as the weights of the hidden layer of the neural network, activation function parameters, etc.) and processes the feature data layer by layer. After receiving the feature data, the input layer passes it to the hidden layer. The hidden layer performs feature fusion according to the dimensionality transformation rules at the time of training. The output layer generates the preliminary wind erosion modulus prediction value for each time point of the region to be predicted based on the trained parameters. The prediction value includes numerical value and confidence interval information.
[0163] Step S153: Through the information exchange channel in the trained Caragana korshinskii multi-feature wind erosion modulus prediction model, the preliminary wind erosion modulus prediction values of each basic prediction sub-model and the filtered feature data to be predicted are transmitted to other related basic prediction sub-models to realize the information exchange of sub-models in the prediction stage.
[0164] After each basic prediction sub-model generates preliminary prediction values, the information exchange channel transmits the preliminary wind erosion modulus prediction values, the filtered feature data to be predicted, and the prediction confidence intervals to other associated sub-models according to the information transmission weights and directions after training. For example, the basic prediction sub-model corresponding to the distribution characteristics of Caragana korshinskii community transmits its own preliminary prediction values, community coverage, and plant spacing feature data to the basic prediction sub-model corresponding to the growth environment characteristics of Caragana korshinskii through a one-way information exchange channel, retaining the time stamp and weight markers of the data during the transmission process.
[0165] The receiving terminal model parses the received information, extracts key data and stores it in a temporary buffer to ensure that the subsequent prediction and correction process can directly call this information. The synchronization of information interaction is guaranteed by the model's built-in synchronization control unit, which is consistent with the synchronization mechanism in the training phase.
[0166] Step S154: Each basic prediction sub-model combines its own generated preliminary wind erosion modulus prediction value with the information transmitted by the associated sub-models to correct the preliminary wind erosion modulus prediction value, thus obtaining the corrected wind erosion modulus prediction value.
[0167] Each basic prediction sub-model extracts information passed from the associated sub-models from the temporary cache, compares and analyzes it with its own preliminary prediction results, and obtains a more accurate prediction value through multiple rounds of correction.
[0168] Step S1541: Each basic prediction sub-model receives the preliminary wind erosion modulus prediction value and the corresponding filtered feature data to be predicted from the associated sub-model.
[0169] The basic prediction sub-model receives data packets from associated sub-models through the information exchange channel's receiving interface. These data packets contain an array of preliminary wind erosion modulus predictions, a matrix of filtered feature data to be predicted, data collection time identifiers, and information transmission weights. For example, the basic prediction sub-model corresponding to the growth environment characteristics of *Caragana korshinskii* receives crown width and branch density feature data from the morphological feature sub-model, along with an array of preliminary wind erosion modulus predictions generated based on these features. It also receives community coverage and plant spacing feature data from the community distribution feature sub-model, along with corresponding arrays of preliminary predictions. The data packets also include the collection time identifier "XXXX-07-15-09" and information transmission weights for each data point. The basic prediction sub-model decapsulates the data packets, storing them according to data type in the feature data cache and prediction value cache, while also recording the sub-model identifier and information transmission weight for each data source.
[0170] Step S1542: Analyze the differences between the filtered feature data to be predicted transmitted by the correlation sub-model and the filtered feature data to be predicted by itself, and determine the difference feature items and the degree of difference. The degree of difference refers to the magnitude of the numerical difference of the same feature item in the two feature data.
[0171] The basic prediction sub-model extracts the filtered feature data to be predicted from its own feature processing results and compares it with the filtered feature data to be predicted transmitted by the associated sub-model. Taking the basic prediction sub-model corresponding to the growth environment characteristics of Caragana korshinskii as an example, the feature data it filters includes soil moisture and wind speed feature data matrices, while the feature data transmitted by the associated sub-model includes crown width, branch density, community coverage, and plant spacing feature data matrices.
[0172] For common features (if any) or related features shared by both datasets, calculate the degree of difference. For numerical features, such as wind speed-related derived feature data from different sub-models, calculate the absolute value of the difference between the corresponding values at each time point, and take the average of the absolute values of the differences at all time points as the degree of difference parameter for that feature. For proportional features, such as the correlation between soil moisture and community cover, calculate the proportional difference between the two within the same time interval, and combine it with time weights to obtain the degree of difference parameter. The degree of difference parameter ranges from 0 to 1, with a larger value indicating a more significant difference in the feature data.
[0173] Simultaneously, features whose difference parameters exceed a preset difference threshold are marked as difference features. For example, if the soil moisture feature data of the Caragana korshinskii growth environment feature submodel and the moisture correlation data derived from branch density by the correlation submodel have a difference parameter exceeding the threshold, this feature is marked as a difference feature.
[0174] Step S1543: Based on the difference feature items and the degree of difference, determine the reference value of the preliminary wind erosion modulus prediction value transmitted by the associated sub-model. The smaller the degree of difference, the higher the reference value.
[0175] A reference value assessment function is constructed, using the degree of difference parameter as the core input, and combining it with parameters such as the prediction confidence interval and information transmission weight of the associated sub-model for comprehensive calculation. The assessment logic is as follows: subtract the degree of difference parameter from 1 to obtain the feature consistency parameter; then, the feature consistency parameter, information transmission weight, and prediction confidence interval parameter are weighted and summed according to a preset ratio to obtain the reference value parameter, which ranges from 0 to 1.
[0176] A reference value threshold is set. If the reference value parameter is greater than or equal to the threshold, the preliminary wind erosion modulus prediction value transmitted by the associated sub-model is determined to have high reference value; if the reference value parameter is lower than the threshold, the reference value is determined to be low. For example, the preliminary prediction value transmitted by the Caragana korshinskii morphological feature sub-model has a small difference parameter and a high information transmission weight, and the calculated reference value parameter exceeds the threshold, so its reference value is determined to be high; while the prediction value transmitted by a certain associated sub-model has an excessively large difference parameter, and the reference value parameter does not reach the threshold, so its reference value is determined to be low.
[0177] Step S1544: If the reference value meets the preset value condition, extract the reasonable fluctuation range in the preliminary wind erosion modulus prediction value transmitted by the associated sub-model, and compare the preliminary wind erosion modulus prediction value generated by itself with the reasonable fluctuation range.
[0178] The preset value condition is that the reference value parameter is greater than or equal to the reference value threshold. When the preliminary predicted value transmitted by the correlation sub-model meets this condition, the confidence interval corresponding to the predicted value array is extracted, and the upper and lower limits of the confidence interval are used as the boundaries of the reasonable fluctuation range. For example, in the preliminary predicted value array transmitted by the correlation sub-model, the predicted value at each time point corresponds to a confidence interval of [lower limit value, upper limit value]. By integrating the confidence intervals of all time points, a reasonable fluctuation range matrix for the predicted value is formed.
[0179] Each time point value in the preliminary wind erosion modulus prediction array generated by itself is compared with the upper and lower limits of the corresponding time point in the reasonable fluctuation range matrix, and the number of time points where its own prediction value is within the fluctuation range and the number of time points outside the range are recorded.
[0180] Step S1545: If the preliminary wind erosion modulus prediction value generated by itself is within a reasonable fluctuation range, then keep the preliminary wind erosion modulus prediction value unchanged; if the preliminary wind erosion modulus prediction value generated by itself exceeds the reasonable fluctuation range, then adjust it to the reasonable fluctuation range to obtain the first corrected prediction value.
[0181] For time points where the predicted value is within a reasonable fluctuation range, the original preliminary predicted value is retained without adjustment; for time points exceeding the reasonable fluctuation range, if the predicted value is higher than the upper limit of the fluctuation range, the predicted value for that time point is adjusted to the upper limit of the fluctuation range; if the predicted value is lower than the lower limit of the fluctuation range, it is adjusted to the lower limit of the fluctuation range.
[0182] During the adjustment process, the original predicted value, the adjusted value, and the basis for adjustment (such as exceeding the upper or lower limit) at each adjustment time point must be recorded to form the first correction record. All adjusted predicted values at all time points are then integrated to obtain an array of predicted values after the first correction. This array still retains the confidence interval information for each time point. For example, if the initial predicted value generated by the model itself at a certain time point is higher than the upper limit of the reasonable fluctuation range of the predicted value of the associated sub-model, it is adjusted to the upper limit value, completing the first correction.
[0183] Step S1546: If the reference value does not meet the preset value condition, extract the key influencing feature items from the filtered predictable feature data passed by the associated sub-model, and analyze the influence trend of the key influencing feature items on the wind erosion modulus.
[0184] When the initial predicted values transmitted by the associated sub-model do not meet the preset value conditions, key features are extracted from the filtered feature data to be predicted. By calculating the historical correlation parameters between each feature item and the wind erosion modulus, several feature items with the highest correlation parameters are selected as key influencing feature items. For example, among the feature data transmitted by the associated sub-model, the branch density feature has the highest historical correlation with the wind erosion modulus and is identified as a key influencing feature item.
[0185] Analyzing the time series of key influencing features and combining the relationship between these features and wind erosion modulus recorded during the training phase, we can determine their influence trends. These trends include positive correlation (wind erosion modulus increases as feature value increases), negative correlation (wind erosion modulus decreases as feature value increases), and nonlinear correlation (feature value exhibits different correlation directions in different intervals). For example, the key influencing feature is community coverage, whose feature data shows an upward trend over time. Historical data shows a negative correlation between this feature and wind erosion modulus, meaning that wind erosion modulus tends to decrease as community coverage increases.
[0186] Step S1547: Based on the influence trend, adjust the initial wind erosion modulus prediction value generated by itself so that the adjusted prediction value is consistent with the influence trend, and obtain the first corrected prediction value.
[0187] Based on the influence trends of key influencing features and the changing direction of current feature data, determine the adjustment direction of the initial forecast value. If the key influencing features show a positive correlation trend and the current feature data shows an upward trend, while the initial forecast value shows a downward trend, then the forecast value will be adjusted upward; if the key influencing features show a negative correlation trend and the current feature data shows an upward trend, while the initial forecast value shows an upward trend, then the forecast value will be adjusted downward.
[0188] The adjustment range is determined based on the magnitude of change in the feature data and the strength of the influence trend. The greater the magnitude of change and the stronger the trend, the larger the adjustment range. After adjustment, an array of predicted values after the first correction is formed, ensuring that the direction of change of the predicted value at each time point is consistent with the influence trend of the key influencing feature. For example, if the key influencing feature, wind speed, shows a positive correlation trend and the current wind speed data is rising, and the initial predicted value is insufficient, the predicted value is increased according to the calculated adjustment range to match the influence trend.
[0189] Step S1548: Repeatedly compare the information transmitted by the associated sub-model with the preliminary wind erosion modulus prediction information generated by itself, and perform a second correction on the prediction value after the first correction until the correction prediction value tends to stabilize, and obtain the final corrected wind erosion modulus prediction value.
[0190] The first revised prediction value is used as the new base prediction value. Steps S1541 to S1547 are repeated, that is, the updated information transmitted by the associated sub-model is received again (if it exists), the differences in feature data are analyzed, the reference value is evaluated, and adjustments are made again based on the evaluation results.
[0191] After each correction, the overall change in the predicted value array before and after the correction is calculated. The change is the average of the absolute values of the differences between the corresponding time points in the two predicted value arrays. If the change is less than a preset stability threshold, the corrected predicted value is determined to be stable, and the correction process stops; if the change is greater than or equal to the stability threshold, the next correction is performed.
[0192] Typically, the predicted value will reach a stable state after two to three corrections. For example, if the variation range parameter between the predicted value after the first correction and the predicted value after the second correction is lower than the stability threshold, the predicted value at this point is determined as the final corrected wind erosion modulus predicted value.
[0193] Step S155: Input the corrected wind erosion modulus prediction values output by all basic prediction sub-models into the result fusion module. The result fusion module performs weighted fusion processing on the corrected wind erosion modulus prediction values according to the fusion weights determined during training. After weighted fusion processing, the final wind erosion modulus prediction result of the Caragana korshinskii growth area to be predicted is generated. The final wind erosion modulus prediction result is associated with the key feature items in the Caragana korshinskii multi-feature set of the area to be predicted.
[0194] All basic prediction sub-models (sub-models corresponding to the morphological characteristics, growth environment characteristics, and community distribution characteristics of Caragana korshinskii) input the final corrected wind erosion modulus prediction value array, along with the corresponding confidence intervals and feature association information, into the result fusion module. The result fusion module calls the fusion weight configuration saved during the training phase. Each basic prediction sub-model's corrected prediction value corresponds to a fixed fusion weight. The value of the fusion weight is determined based on the prediction accuracy of each sub-model during training; the higher the accuracy, the greater the weight.
[0195] During weighted fusion processing, for each time point, the corrected predicted values of the three sub-models are multiplied by their respective fusion weights to obtain a weighted predicted value. These three weighted predicted values are then concatenated to form the multi-dimensional fusion prediction result for that time point, with the concatenation order consistent with the sub-model type order (morphological features → growth environment features → community distribution features). Simultaneously, the confidence interval and fusion weight information corresponding to each weighted predicted value are retained as supplementary explanations of the fusion result.
[0196] The multi-dimensional fusion prediction results from all time points are integrated to form a final wind erosion modulus prediction result array. Within this array, the prediction results for each time point are associated with and labeled with corresponding key feature terms and their values. For example, the core features affecting the wind erosion modulus at that time point are labeled as "wind speed feature (numerical)" and "community coverage feature (numerical)," thus establishing a correlation between the prediction results and the input features. The final output prediction results also include result reliability assessment parameters, which are calculated based on the confidence intervals of each sub-model and the degree of consistency during the fusion process.
[0197] Based on the same inventive concept, please refer to Figure 2 The diagram shows a schematic block diagram of a multi-feature wind erosion modulus prediction system 100 for performing the above-described multi-feature wind erosion modulus prediction method for Caragana korshinskii using integrated learning, provided in an embodiment of this application. The multi-feature wind erosion modulus prediction system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0198] In this embodiment, both the machine-readable storage medium 120 and the processor 130 are located within the Caragana korshinskii multi-feature wind erosion modulus prediction system 100 incorporating ensemble learning, and are separately configured. However, it should be understood that the machine-readable storage medium 120 may also be independent of the Caragana korshinskii multi-feature wind erosion modulus prediction system 100 incorporating ensemble learning, and may be accessed by the processor 130 via a bus interface. Alternatively, the machine-readable storage medium 120 may also be integrated into the processor 130 and may communicate with external systems via the communication unit 110.
[0199] The processor 130 is the control center of the integrated learning-based multi-feature wind erosion modulus prediction system 100 for Caragana korshinskii. It connects various parts of the system via various interfaces and lines, and performs overall monitoring of the system by running or executing software programs and / or modules stored in the machine-readable storage medium 120, and by accessing data stored in the machine-readable storage medium 120. Optionally, the processor 130 may include one or more processing cores; for example, it may integrate an application processor and a modem processor, where the application processor primarily handles the operating system, user interface, and applications, and the modem processor primarily handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to implement the multi-feature wind erosion modulus prediction method of Caragana korshinskii combined with ensemble learning provided in the aforementioned method embodiments.
[0200] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for predicting the wind erosion modulus of a multi-feature of a Caragana, characterized in that, The method comprises: acquiring a multi-feature set of a caragana growth area and a historical wind erosion modulus record corresponding to the caragana growth area, the multi-feature set of the caragana comprising caragana morphological features, caragana growth environment features and caragana community distribution features, and the historical wind erosion modulus record corresponding to the collection time of the multi-feature set of the caragana; based on the correlation between the features in the multi-feature set of the caragana, a multi-feature correlation transmission link of the caragana is established, the multi-feature correlation transmission link of the caragana being used to represent the information transmission direction and correlation strength between different types of caragana features; according to the multi-feature correlation transmission link of the caragana, an ensemble learning sub-model set is constructed, each sub-model in the ensemble learning sub-model set corresponding to an information processing path of one type of caragana feature, and the sub-models realizing information interaction through the multi-feature correlation transmission link of the caragana; the multi-feature set of the caragana and the corresponding historical wind erosion modulus record are input into the ensemble learning sub-model set, a sub-model cooperative training process is performed, a trained multi-feature wind erosion modulus prediction model of the caragana is obtained, and the model parameters are adjusted through information interaction between the sub-models in the sub-model cooperative training process; the multi-feature set of a caragana growth area to be predicted is input into the trained multi-feature wind erosion modulus prediction model of the caragana, and a wind erosion modulus prediction result of the caragana growth area is generated, the wind erosion modulus prediction result corresponding to the multi-feature set of the caragana growth area to be predicted in feature correlation; the multi-feature correlation transmission link of the caragana, an ensemble learning sub-model set is constructed, each sub-model in the ensemble learning sub-model set corresponding to an information processing path of one type of caragana feature, and the sub-models realizing information interaction through the multi-feature correlation transmission link of the caragana; according to the type of the feature item in the multi-feature correlation transmission link of the caragana, a feature processing module is divided, each feature processing module corresponding to processing one type of caragana feature, the feature processing module comprising a caragana morphological feature processing module, a caragana growth environment feature processing module and a caragana community distribution feature processing module; a corresponding basic prediction sub-model is configured for each feature processing module, the input of the basic prediction sub-model being the caragana feature data processed by the corresponding feature processing module, and the output being a preliminary wind erosion modulus prediction value; based on the information transmission direction in the multi-feature correlation transmission link of the caragana, an information interaction channel is established between the basic prediction sub-models, the information interaction channel being used to realize feature information transmission between different basic prediction sub-models; an information transmission weight is configured for each information interaction channel, the information transmission weight being determined according to the correlation strength of the corresponding correlation in the multi-feature correlation transmission link of the caragana; a result fusion module is arranged at the output end of each basic prediction sub-model, the result fusion module being used to receive the preliminary wind erosion modulus prediction values output by all the basic prediction sub-models and distribute fusion weights according to the prediction accuracy of each basic prediction sub-model; the feature processing module, the basic prediction sub-model, the information interaction channel and the result fusion module are integrated to form an ensemble learning sub-model set, and each sub-model in the ensemble learning sub-model set keeps information synchronization through the information interaction channel; The multi-feature set of the caragana and the corresponding historical wind erosion modulus record are input into the integrated learning sub-model set, a sub-model cooperative training process is performed, and a trained multi-feature wind erosion modulus prediction model of the caragana is obtained. The multi-feature set of the caragana is distributed to corresponding feature processing modules in the integrated learning sub-model set according to feature types, each feature processing module performs feature screening processing on the input caragana feature data, and retains feature items associated with wind erosion modulus that meet the requirements; Each feature processing module inputs the screened caragana feature data into a corresponding basic prediction sub-model to generate a preliminary wind erosion modulus prediction value of each basic prediction sub-model; Through a preset information interaction channel, the preliminary wind erosion modulus prediction value of each basic prediction sub-model and the corresponding feature data are transmitted to other associated basic prediction sub-models to realize information interaction between the sub-models; After each basic prediction sub-model receives the information transmitted by the associated sub-model, the preliminary wind erosion modulus prediction value of the basic prediction sub-model and the corresponding historical wind erosion modulus record are combined to calculate a prediction deviation value; According to the prediction deviation value, the internal parameters of the basic prediction sub-model and the information transmission weight of the corresponding information interaction channel are adjusted to reduce the prediction deviation value; The processes of feature screening processing, preliminary wind erosion modulus prediction value generation, information interaction, prediction deviation value calculation, and parameter weight adjustment are repeatedly performed until the prediction deviation value is stable within a preset range; When the prediction deviation value is stable within the preset range, the sub-model cooperative training process is stopped, and the integrated learning sub-model set at this time is determined as the trained multi-feature wind erosion modulus prediction model of the caragana.
2. The method according to claim 1, wherein, Based on the association relationship between each feature in the multi-feature set of the caragana, a caragana multi-feature association transmission link is established, including: Specific feature items of caragana shape features, caragana growth environment features, and caragana community distribution features are extracted from the multi-feature set of the caragana, and the feature attributes and data representation forms of each feature item are determined; The association relationship between any two feature items is analyzed to determine the type of the association relationship, which includes direct association and indirect association. The direct association refers to the direct influence between two feature items, and the indirect association refers to the influence between two feature items through a third feature item. For directly associated feature items, the association influence degree between the feature items is calculated, which is determined by the synchronization trend of the feature items over time. For indirectly associated feature items, the intermediate association feature items are identified, and the influence transmission order and transmission amplitude of the intermediate association feature items on the two indirectly associated feature items are determined. According to the type of the association relationship, the association influence degree, the intermediate association feature items, and the influence transmission order, an initial feature association network is constructed, in which each node corresponds to a feature item, and the connection between the nodes corresponds to the association relationship. The initial feature association network is combed for information transmission paths, paths with information transmission efficiency meeting the association requirements are retained, and paths with information transmission efficiency not meeting the association requirements are deleted to form the caragana multi-feature association transmission link.
3. The method according to claim 1, wherein, The information transmission weight is configured for each information interaction channel, including: Extracting the association strength description information between the two feature items associated with the corresponding information interaction channel from the Salix mongolica multi-feature association transmission link; Analyzing the influence duration and influence range between the feature items in the association strength description information, the influence duration refers to the time span of the influence of the change of one feature item on another feature item, and the influence range refers to the number of feature dimensions of the influence of the change of one feature item on another feature item; The influence duration and influence range are normalized respectively, and the quantization basis of the association strength is determined according to the normalized influence duration parameter and influence range parameter, which is used to convert the association strength description information into a calculable association parameter; Based on the quantization basis, the association strength description information is converted into an initial association parameter, which can reflect the relative size of the association strength between the two feature items; Referring to the prediction error feedback information corresponding to the information interaction channel in the historical training data, the initial association parameter is calibrated to obtain an adjusted association parameter; The adjusted association parameter is mapped to the information transmission weight, and the value of the information transmission weight is positively correlated with the adjusted association parameter, and the information transmission weight configuration of each information interaction channel is completed.
4. The method according to claim 1, wherein, The internal parameters of the base prediction sub-model and the information transmission weight of the corresponding information interaction channel are adjusted according to the prediction deviation value, including: Analyzing the source of the prediction deviation value to determine whether the deviation is from the internal parameter setting of the base prediction sub-model or the information transmission weight setting of the information interaction channel; If the deviation is from the internal parameter setting of the base prediction sub-model, extract the internal parameters in the base prediction sub-model that are directly related to the wind erosion modulus prediction, calculate the contribution degree of each internal parameter to the prediction deviation value, and the contribution degree refers to the magnitude of the change of the prediction deviation value caused by the change of a single internal parameter; According to the contribution degree, the internal parameters are adjusted to the direction that reduces the prediction deviation value to obtain the adjusted internal parameters; If the deviation is from the information transmission weight setting of the information interaction channel, extract the association strength description information and historical adjustment record corresponding to the information interaction channel, and analyze the gap between the current information transmission weight and the optimal weight; According to the gap, the value of the information transmission weight is adjusted so that the adjusted information transmission weight is closer to the optimal weight, thereby reducing the prediction deviation value; after adjusting the internal parameters of the base prediction sub-model and the information transmission weight of the information interaction channel, the prediction deviation value is recalculated to determine whether the adjusted prediction deviation value is reduced; If the adjusted prediction deviation value is reduced, the adjusted internal parameters and information transmission weight are retained; if the adjusted prediction deviation value is increased, the internal parameters of the base prediction sub-model and the information transmission weight of the information interaction channel before adjustment are restored, and the adjustment direction is reselected.
5. The method according to claim 1, wherein, The Salix mongolica multi-feature set of the to-be-predicted Salix mongolica growth area is input into the trained Salix mongolica multi-feature wind erosion modulus prediction model to generate the wind erosion modulus prediction result of the Salix mongolica growth area, including: Collect a plurality of features of the shrub in the area to be predicted, and ensure that the plurality of features of the shrub include morphological features of the shrub, growth environment features of the shrub and community distribution features of the shrub, and the number of feature items is consistent with the plurality of features of the shrub during training; input the plurality of features of the shrub to be predicted into the feature processing module of the trained prediction model of the plurality of features of the shrub, each feature processing module performs screening processing on the input feature data to be predicted according to the screening rule during training, and the screened feature data to be predicted is input into the corresponding basic prediction sub-model, and each basic prediction sub-model generates a preliminary wind erosion modulus prediction value of the area to be predicted based on the internal parameters of the trained prediction model; through the information interaction channel in the trained prediction model of the plurality of features of the shrub, the preliminary wind erosion modulus prediction value of each basic prediction sub-model and the screened feature data to be predicted are transmitted to other associated basic prediction sub-models to realize the information interaction of the sub-models in the prediction stage; each basic prediction sub-model combines the preliminary wind erosion modulus prediction value generated by itself and the information transmitted by the associated sub-model to modify the preliminary wind erosion modulus prediction value and obtain a modified wind erosion modulus prediction value; input the modified wind erosion modulus prediction values output by all basic prediction sub-models into the result fusion module, and the result fusion module performs weighted fusion processing on the modified wind erosion modulus prediction values according to the fusion weight determined during training, and generates a final wind erosion modulus prediction result of the area to be predicted after the weighted fusion processing, and the final wind erosion modulus prediction result is associated with the key feature items in the plurality of features of the shrub in the area to be predicted.
6. The method according to claim 5, wherein, The modification of the preliminary wind erosion modulus prediction value by each basic prediction sub-model in combination with the preliminary wind erosion modulus prediction value generated by itself and the information transmitted by the associated sub-model includes: each basic prediction sub-model receives the preliminary wind erosion modulus prediction value and the corresponding screened feature data to be predicted transmitted by the associated sub-model; analyze the difference between the screened feature data to be predicted transmitted by the associated sub-model and the screened feature data to be predicted of itself, determine the difference feature items and the difference degree, and the difference degree refers to the numerical difference between the same feature items in the two kinds of feature data; according to the difference feature items and the difference degree, judge the reference value of the preliminary wind erosion modulus prediction value transmitted by the associated sub-model, the smaller the difference degree, the higher the reference value; if the reference value meets the preset value condition, extract the reasonable fluctuation range in the preliminary wind erosion modulus prediction value transmitted by the associated sub-model, and compare the preliminary wind erosion modulus prediction value generated by itself with the reasonable fluctuation range; if the preliminary wind erosion modulus prediction value generated by itself is within the reasonable fluctuation range, the preliminary wind erosion modulus prediction value remains unchanged; if the preliminary wind erosion modulus prediction value generated by itself is beyond the reasonable fluctuation range, it is adjusted to be within the reasonable fluctuation range to obtain a first modified prediction value; if the reference value does not meet the preset value condition, extract the key influence feature items in the screened feature data to be predicted transmitted by the associated sub-model, and analyze the influence trend of the key influence feature items on the wind erosion modulus; According to the influence trend, the preliminary wind erosion modulus prediction value generated by itself is adjusted, the adjusted prediction value is consistent with the influence trend, and a first modified prediction value is obtained; The information transmitted by the comparative correlation sub-model and the preliminary wind erosion modulus prediction value information generated by itself are repeatedly compared, the first modified prediction value is secondarily modified until the modified prediction value tends to be stable, and a final modified wind erosion modulus prediction value is obtained.
7. The method according to claim 1, wherein, The morphological characteristics of the caragana include caragana plant height characteristics, caragana crown width characteristics, caragana branch density characteristics and caragana root system distribution characteristics; the growth environment characteristics of the caragana include soil humidity characteristics, soil texture characteristics, wind speed characteristics and precipitation characteristics of a growth area; and the caragana community distribution characteristics include caragana plant spacing characteristics, caragana community coverage characteristics and symbiotic distribution characteristics of the caragana and other plants, and include: The morphological characteristics of the caragana, the growth environment characteristics of the caragana and the caragana community distribution characteristics are separated from the caragana multi-feature set, and specific feature items under each type of feature are determined; For the caragana plant height characteristics in the morphological characteristics of the caragana, feature data is obtained by measuring the vertical distance from the ground to the top of the caragana; for the caragana crown width characteristics in the morphological characteristics of the caragana, feature data is obtained by measuring the maximum span of the caragana crown in the horizontal direction; for the caragana branch density characteristics in the morphological characteristics of the caragana, feature data is obtained by counting the number of caragana branches in a unit area; and for the caragana root system distribution characteristics in the morphological characteristics of the caragana, feature data is obtained by measuring the distribution depth and distribution range of the caragana root system in the soil; For the soil humidity characteristics in the growth environment characteristics of the caragana, feature data is obtained by detecting the water content in the soil of the growth area; for the soil texture characteristics in the growth environment characteristics of the caragana, feature data is obtained by analyzing the proportion of sand particles, silt particles and clay particles in the soil; for the wind speed characteristics in the growth environment characteristics of the caragana, feature data is obtained by recording the air flow speed of the growth area; and for the precipitation characteristics in the growth environment characteristics of the caragana, feature data is obtained by counting the total amount of precipitation and the precipitation frequency of the growth area; For the caragana plant spacing characteristics in the caragana community distribution characteristics, feature data is obtained by measuring the horizontal distance between adjacent caragana plants; for the caragana community coverage characteristics in the caragana community distribution characteristics, feature data is obtained by calculating the proportion of the projected area of the caragana community to the total area of the growth area; and for the symbiotic distribution characteristics of the caragana and other plants in the caragana community distribution characteristics, feature data is obtained by counting the types and quantities of other plants that symbiotically grow with the caragana in the growth area; All the feature data obtained is sorted according to the classification of the morphological characteristics of the caragana, the growth environment characteristics of the caragana and the caragana community distribution characteristics, to form a structured caragana multi-feature set, and each feature item in the structured caragana multi-feature set corresponds to unique feature data and collection time.
8. A system for predicting the wind erosion modulus of a multi-featured Caragana shrub in combination with ensemble learning, characterized in that, It includes: a processor; a machine readable storage medium for storing machine executable instructions of the processor; wherein the processor is configured to execute the machine executable instructions to perform the caragana multi-feature wind erosion modulus prediction method combined with integrated learning according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-model integrated load prediction method based on wavelet transform
CN110443417A
Insect population density prediction method and system based on neural network multi-model combination
CN112862168A