Forest grass science and technology data analysis and prediction method and system based on cloud computing

By improving the GWO algorithm and optimizing the LSTM hyperparameters, and combining various data processing techniques, the problem of inaccurate forestry and grassland data processing was solved, achieving accuracy in forestry and grassland trend prediction and scientific ecological protection, and providing a rational planning and management strategy for forestry and grassland resources.

CN121920607APending Publication Date: 2026-04-24河南省林业生态建设发展中心 +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
河南省林业生态建设发展中心
Filing Date
2026-01-14
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies in forestry and grassland data processing suffer from inaccurate analysis, an inability to quickly meet the processing needs of massive amounts of forestry and grassland data, an inability to provide accurate predictions in complex forestry and grassland data scenarios, an inability to adapt to diverse scenarios, and an inability to provide accurate predictions for different categories of forestry and grassland data.

Method used

An improved GWO algorithm is used to optimize the hyperparameters of an LSTM model. Data is collected from IoT sensors, remote sensing equipment, and scientific research platforms. Unstructured text information is extracted using NLP (Natural Language Processing) technology. Long-distance semantic dependencies are captured through a Transformer model to generate unified metadata tags. An LSTM prediction model is then established. The improved GWO algorithm is used to optimize the hyperparameters to generate a target GWO-LSTM prediction model for trend prediction. Finally, dynamic data visualization charts are generated using GAN (Generative Artificial Intelligence).

Benefits of technology

It improves the accuracy of forestry and grassland trend forecasting, provides a basis for the planning and utilization of forestry and grassland resources, enables the rational arrangement of logging and protected areas, protects biodiversity, provides scientific forestry and grassland resource management strategies, and promotes the sustainable development of forestry and grassland ecology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920607A_ABST
    Figure CN121920607A_ABST
Patent Text Reader

Abstract

The invention discloses a forest and grass science and technology data analysis and prediction method and system based on cloud computing, initial forest and grass data are transmitted to a cloud computing platform, a Spark framework is utilized to perform real-time denoising, format conversion and semantic analysis on the data to generate a unified metadata tag, and forest and grass processing data are obtained; capturing a long-distance semantic dependency relationship of the forest grass processing data through a multi-head attention mechanism, and classifying project data to obtain classified forest grass data; establishing an LSTM prediction model based on an LSTM long short-term memory network, and optimizing hyper-parameters of the LSTM prediction model by using an improved GWO grey wolf optimization algorithm to obtain a target prediction model; and performing trend prediction on the classified forest and grass data by using a target prediction model, and outputting a species distribution change trend. The improved algorithm is adopted to adapt to various scenes, the accuracy of data analysis and processing can be improved, and accurate analysis and prediction are provided for forest and grass data of different classifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forestry and grassland data management technology, and in particular to cloud computing-based methods and systems for analyzing and predicting forestry and grassland science and technology data. Background Technology

[0002] As a key area for ecological protection and restoration, the forestry and grassland industry has accumulated massive data resources, covering multiple aspects such as forest resource monitoring, ecological assessment, pest and disease control, and scientific research results. Traditional data management methods are insufficient to cope with the complexity and diversity of forestry and grassland data. This data includes not only structured tabular data but also unstructured text reports, remote sensing images, and time-series sensor data, among other multimodal information.

[0003] The existing technology CN117557400A employs a cloud computing-based intelligent tree growth monitoring system, including a data acquisition module, a data transmission module, a data storage module, a machine learning algorithm module, a data analysis and processing module, a visualization interface module, and a database. It collects data through sensors and uses a CNN model from deep learning algorithms for analysis and growth prediction. However, it suffers from inaccurate analysis when processing this heterogeneous data, resulting in low efficiency in the lifecycle management of forestry and grassland science and technology data. It cannot quickly meet the processing needs of massive amounts of forestry and grassland data, lacks the processing capabilities required for complex forestry and grassland data scenarios, cannot adapt to diverse scenarios, and cannot provide accurate predictions for different categories of forestry and grassland data. Summary of the Invention

[0004] The purpose of this invention is to solve the above-mentioned problems. It designs a cloud computing-based method for analyzing and predicting forestry and grassland science and technology data. It adopts an improved GWO algorithm to optimize the hyperparameters of LSTM, avoids getting trapped in local optima, and the optimized model can capture long-term dependencies, improve the accuracy of forestry and grassland trend prediction, provide a basis for the planning and utilization of forestry and grassland resources, rationally arrange logging, plan protected areas, and protect biodiversity.

[0005] To achieve the above objectives, the technical solution of the present invention further includes the following steps in the above-mentioned cloud computing-based forestry and grassland science and technology data analysis and prediction method:

[0006] Forestry and grassland data are collected through IoT sensors, remote sensing equipment and scientific research platforms, and unstructured text information is extracted by combining NLP natural language processing technology to obtain initial forestry and grassland data.

[0007] The initial forest and grassland data is transmitted to a cloud computing platform, and the Spark framework is used to perform real-time noise reduction, format conversion, and semantic parsing on the data to generate unified metadata tags, thus obtaining the processed forest and grassland data.

[0008] By capturing the long-distance semantic dependencies of the forestry and grassland processing data through the multi-head attention mechanism in the Transformer model, the project data is classified to obtain classified forestry and grassland data.

[0009] An LSTM prediction model is established based on the LSTM long short-term memory network. The hyperparameters of the LSTM prediction model are optimized using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model.

[0010] The target GWO-LSTM prediction model is used to predict the trend of the classified forest and grassland data, and the species distribution change trend is output.

[0011] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the initial forestry and grassland data is obtained by collecting forestry and grassland data through IoT sensors, remote sensing equipment, and scientific research platforms, and then extracting unstructured text information using NLP (Natural Language Processing) technology, including:

[0012] Collect forestry and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, including at least forestry and grassland growth environment data, forestry and grassland coverage area and growth status data, forestry and grassland pest and disease situation and seed germination rate data.

[0013] By combining NLP (Natural Language Processing) technology to segment the collected data into word units, the text is divided into word segmentation data.

[0014] The segmented data is labeled with part-of-speech tags to determine the part of speech of each word. Entity recognition is performed on the labeled words to identify forest and grassland species, place names, and time information in the text, thus obtaining initial forest and grassland data.

[0015] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the step of transmitting the initial forestry and grassland data to a cloud computing platform, and using the Spark framework to perform real-time noise reduction, format conversion, and semantic parsing of the data to generate unified metadata tags, thereby obtaining processed forestry and grassland data, includes:

[0016] The initial forestry and grassland data is transmitted to the cloud computing platform via network using encrypted transmission; a range threshold is set for the data, and data exceeding the range threshold is treated as noise and removed using the Spark framework, and then converted into JSON format data;

[0017] By analyzing the context and meaning of JSON format data and combining it with domain knowledge graphs, semantic parsing is performed on the data to generate preliminary semantic identifiers;

[0018] Semantic identifiers are organized and standardized according to a unified metadata standard to generate unified metadata tags. Metadata tags include at least the data source, collection time, data type, and semantic information.

[0019] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the step of capturing the long-distance semantic dependencies of the forestry and grassland processing data through the multi-head attention mechanism in the Transformer model, classifying the project data, and obtaining classified forestry and grassland data includes:

[0020] By concatenating and linearly transforming the results of multiple attention heads in the Transformer model, long-distance semantic dependencies in the data can be captured.

[0021] Based on the captured long-distance semantic dependencies, the forestry and grassland processing data are classified, and the classification categories include forestry and grassland species, growth stage, and geographical location.

[0022] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the step of establishing an LSTM prediction model based on an LSTM long short-term memory network and optimizing the hyperparameters of the LSTM prediction model using an improved GWO gray wolf optimization algorithm to obtain a target GWO-LSTM prediction model includes:

[0023] An initial population is generated using chaotic initialization, and the exploration and development capabilities of an adaptive weight factor balancing algorithm are introduced to obtain an improved GWO (Grey Wolf) optimization algorithm.

[0024] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the step of establishing an LSTM prediction model based on an LSTM long short-term memory network and optimizing the hyperparameters of the LSTM prediction model using an improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model further includes:

[0025] The improved GWO algorithm is used to optimize the hyperparameters of the LSTM prediction model, including the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient.

[0026] Determine the search range and fitness function of hyperparameters, initialize the gray wolf population, with each individual representing a set of hyperparameter combinations, calculate the fitness value of each individual, and determine the α, β, and δ wolves in the wolf pack based on the fitness value;

[0027] The positions of other individuals in the population are updated according to the predation behavior of gray wolves, and the hyperparameter combination is iteratively output to obtain the target GWO-LSTM prediction model.

[0028] Furthermore, in the aforementioned cloud-based forestry and grassland science and technology data analysis and prediction method, the step of using the target GWO-LSTM prediction model to perform trend prediction on the classified forestry and grassland data and outputting the species distribution change trend includes:

[0029] The classification of forest and grassland data is input into the target GWO-LSTM prediction model. Based on the temporal characteristics and historical change patterns of the data, the distribution change trend of forest and grassland species is predicted, and the species distribution change trend is output.

[0030] Based on the changing trends of species distribution, dynamic data visualization charts are generated through generative adversarial networks (GANs), and combined with augmented reality (AR) technology to provide decision-making support for managers.

[0031] Furthermore, user permissions are dynamically adjusted based on DRL deep reinforcement learning in the cloud computing platform, and privacy is protected when sharing data across departments using federated learning.

[0032] Furthermore, in the cloud-based forestry and grassland science and technology data analysis and prediction system, the forestry and grassland science and technology data full life cycle management system includes the following modules:

[0033] The forestry and grassland data acquisition module is used to collect forestry and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, and extract unstructured text information by combining NLP natural language processing technology to obtain initial forestry and grassland data.

[0034] The forestry and grassland data processing module is used to transmit the initial forestry and grassland data to the cloud computing platform, and use the Spark framework to perform real-time noise reduction, format conversion and semantic parsing of the data to generate unified metadata tags, thereby obtaining the processed forestry and grassland data.

[0035] The forestry and grassland data classification module is used to capture the long-distance semantic dependencies of the forestry and grassland processed data through the multi-head attention mechanism in the Transformer model, classify the project data, and obtain classified forestry and grassland data.

[0036] The prediction model building module is used to build an LSTM prediction model based on the LSTM long short-term memory network, and optimize the hyperparameters of the LSTM prediction model using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model.

[0037] The species distribution analysis module is used to perform trend prediction on the classified forest and grassland data using the target GWO-LSTM prediction model and output the species distribution change trend.

[0038] Furthermore, in the system for implementing the above-mentioned cloud computing-based forestry and grassland science and technology data analysis and prediction method, the species distribution analysis module includes the following sub-modules:

[0039] An improved submodule is used to generate the initial population using chaotic initialization. The ability to explore and develop an adaptive weight factor balancing algorithm is introduced, resulting in an improved GWO (Grey Wolf) optimization algorithm.

[0040] Furthermore, in the system for implementing the above-mentioned cloud computing-based forestry and grassland science and technology data analysis and prediction method, the prediction model building module includes the following sub-modules:

[0041] The optimization submodule is used to optimize the hyperparameters of the LSTM prediction model using the improved GWO algorithm, including the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient.

[0042] The parameter submodule is used to determine the search range and fitness function of hyperparameters, initialize the gray wolf population, with each individual representing a set of hyperparameter combinations, calculate the fitness value of each individual, and determine the α, β, and δ wolves in the wolf pack based on the fitness value;

[0043] The output submodule is used to update the positions of other individuals in the population according to the predation behavior of the gray wolves, iteratively output the hyperparameter combination, and obtain the target GWO-LSTM prediction model.

[0044] Its beneficial effects are,

[0045] 1. This invention employs multi-head attention to standardize the output through a LayerNormalization layer, thereby stabilizing parameter updates during training, accelerating model convergence, and enabling the discovery of potential correlations across long-distance intervals in the data. This provides a more comprehensive and accurate basis for classification, thereby improving the performance of the classification model in complex forestry and grassland data scenarios and providing strong technical support for the scientific management and accurate classification of forestry and grassland resources.

[0046] 2. The present invention utilizes the improved GWO algorithm to optimize the hyperparameters of the LSTM prediction model, which has significant technical advantages in the special scenario of trend prediction of classified forest and grassland data. The improved GWO algorithm uses chaotic initialization to generate the initial population, which greatly increases the population diversity, avoids getting trapped in local optima, and lays the foundation for subsequent search for the global optimal hyperparameter combination.

[0047] 3. The optimized GWO-LSTM prediction model of this invention can better capture long-term dependencies and potential patterns in the data. When predicting the trend of species distribution changes, it can fully consider the temporal changes of forest and grassland growth affected by environmental factors, and improve the accuracy of prediction by combining historical data with current characteristics.

[0048] 4. This invention generates dynamic data visualization charts through GAN (Generative Adversarial Network). The generated charts are highly interactive, allowing users to freely adjust parameters and switch perspectives, and to explore in depth the impact of different factors on species distribution. This enables the formulation of more targeted and scientific forest and grassland resource protection and management strategies, and promotes the sustainable development of forest and grassland ecosystems.

[0049] 5. The accurate prediction of species distribution change trends in this invention can provide a basis for the rational planning and utilization of forest and grassland resources. Based on the prediction results, logging plans can be rationally arranged to avoid ecological damage caused by over-logging. In view of the possible migration directions of species, the scope and layout of protected areas can be planned in advance to protect biodiversity. Attached Figure Description

[0050] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0051] Figure 1 This is a schematic diagram of the first embodiment of the cloud computing-based forestry and grassland science and technology data analysis and prediction method in this invention.

[0052] Figure 2 This is a schematic diagram of the second embodiment of the cloud computing-based forestry and grassland science and technology data analysis and prediction method in this invention.

[0053] Figure 3 This is a schematic diagram of the first embodiment of the cloud computing-based forestry and grassland science and technology data analysis and prediction system of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0055] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms "one," "an," "say," and "this" used herein may also include plural forms. It should be further understood that the terminology used in this specification includes the presence of the stated feature, integer, step, operation, element, and / or component, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0056] The present invention will now be described in detail with reference to the accompanying drawings, such as... Figure 1 As shown, the cloud-based forestry and grassland science and technology data analysis and prediction method includes the following steps in its full lifecycle management of forestry and grassland science and technology data:

[0057] Step 101: Collect forest and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, and extract unstructured text information by combining NLP natural language processing technology to obtain initial forest and grassland data;

[0058] Specifically, in this embodiment, forest and grassland data are collected through IoT sensors, remote sensing equipment and scientific research platforms, including at least forest and grassland growth environment data, forest and grassland coverage area and growth status data, forest and grassland pest and disease situation and seed germination rate data.

[0059] By combining NLP (Natural Language Processing) technology to segment the collected data into word units, the text is divided into word segmentation data.

[0060] Part-of-speech tagging is performed on the segmented data to determine the part of speech of each word. Entity recognition is then performed on the tagged words to identify forest and grassland species, place names, and time information in the text, thus obtaining the initial forest and grassland data.

[0061] Specifically,

[0062] IoT sensor data acquisition

[0063] Data Acquisition Equipment and Data: IoT sensors can collect environmental data on forest and grassland growth, such as temperature, humidity, soil nutrients, light intensity, and carbon dioxide concentration. Different types of sensors have different performance characteristics. For example, a certain model of air temperature and humidity sensor has a temperature range of -40℃ to 80℃ with an accuracy of ±0.5℃; and a humidity range of 0 to 100%RH with an accuracy of ±2%RH. Soil nutrient sensors can collect the content of elements such as nitrogen, phosphorus, and potassium in the soil with an accuracy of up to 0.1 mg / kg.

[0064] Data collection frequency: Different collection frequencies are set according to the characteristics of forest and grassland growth and research needs. For parameters that change rapidly in the growth environment, such as air temperature and humidity, data can be collected every 10 minutes; for parameters that change slowly, such as soil nutrients, data can be collected once a day.

[0065] Remote sensing equipment data collection

[0066] Equipment and Data: Remote sensing equipment, such as high-resolution satellites, can provide macroscopic data on forest and grassland coverage and growth status at meter-level resolution. Remote sensing imagery can acquire information such as the National Density Variable Index (NDVI) and leaf area index of forests and grasslands, reflecting their growth vitality and coverage. Furthermore, unmanned aerial vehicle (UAV) remote sensing can be used for small-scale, high-precision data collection on forests and grasslands, such as their height and density.

[0067] Data processing: Remote sensing data needs to be preprocessed, including radiometric correction and geometric correction, to eliminate the influence of atmospheric and topographical factors on the data and improve the accuracy of the data.

[0068] Scientific research platform collection

[0069] The research platform can provide experimental data on forest and grassland pests and diseases, growth rates, and seed germination rates. This data is typically obtained through manual observation and experimental analysis, and is highly accurate and targeted.

[0070] NLP technology extracts unstructured text information.

[0071] Operation process: NLP technology processes unstructured texts such as scientific research literature, observation reports, and expert experience. First, word segmentation is performed to divide the text into individual words or lexical units; then, part-of-speech tagging is performed to determine the part of speech of each word; next, entity recognition is performed to identify entity information such as forest and grassland species, place names, and time in the text.

[0072] Problem Solving: During the extraction process, issues such as polysemy and ambiguity may be encountered. These can be addressed by considering the context and using pre-trained language models to improve the accuracy of information extraction. For unstructured text from different sources, there may be differences in format and expression, requiring standardization to unify the expression methods.

[0073] Step 102: Transmit the initial forestry and grassland data to the cloud computing platform, and use the Spark framework to perform real-time noise reduction, format conversion and semantic parsing of the data to generate unified metadata tags, thus obtaining the forestry and grassland processed data.

[0074] Specifically, in this embodiment, encrypted transmission is used to transmit the initial forestry and grassland data to the cloud computing platform via the network; a range threshold is set for the data, and the Spark framework is used to treat data exceeding the range threshold as noise and remove it, and then convert it into JSON format data;

[0075] By analyzing the context and meaning of JSON format data and combining it with domain knowledge graphs, semantic parsing is performed on the data to generate preliminary semantic identifiers;

[0076] Semantic identifiers are organized and standardized according to a unified metadata standard to generate unified metadata tags. Metadata tags include at least the data source, collection time, data type, and semantic information.

[0077] Specifically,

[0078] Deep operations of data transmission to cloud computing platforms

[0079] Transmission protocol selection: The transmission protocol is flexibly selected based on the data volume and real-time requirements. For sensor data with high real-time requirements, the MQTT protocol is used. Its lightweight nature reduces network bandwidth consumption and is suitable for low-bandwidth and unstable network environments, ensuring real-time data delivery. For batches of remote sensing imagery and scientific research platform data, the HTTP / HTTPS protocol is used. With the help of the breakpoint resume function, data transmission failure caused by network interruption is avoided, ensuring that large amounts of data can be transmitted completely.

[0080] Data preprocessing before transmission: Initial forestry and grassland data is compressed before data transmission. For time-series data from sensors, differential coding compression is used, transmitting only the difference between adjacent data, significantly reducing the data volume. For spatial data such as remote sensing imagery, the JPEG2000 compression algorithm is used to reduce storage and transmission costs while maintaining image quality. The compressed data generates a checksum. After transmission to the cloud computing platform, the checksum is compared to confirm data integrity; if incomplete, a retransmission mechanism is triggered.

[0081] Cloud computing platform receiving mechanism: The cloud computing platform is equipped with distributed data receiving nodes, which can simultaneously receive data from multiple acquisition terminals. The receiving nodes perform preliminary screening of the data, removing data with obvious format errors, and assign a unique identifier to each data block, associating it with its source information and transmission time, facilitating subsequent tracking and tracing.

[0082] Deep analysis of data processing in the Spark framework

[0083] Multi-layered verification for real-time denoising: In addition to setting basic threshold ranges, a dynamic threshold adjustment mechanism is introduced. A dynamic threshold model is established by combining historical data from the same period and regional environmental characteristics. During the high-temperature period in summer, the upper limit of the temperature threshold is appropriately increased; during the rainy season, the humidity threshold range is adjusted. For suspected noisy data, cross-validation with data from adjacent sensors is performed. If the deviation of a sensor's data from multiple surrounding sensors exceeds 15%, it is determined to be noise. Simultaneously, an isolated forest algorithm is used to identify outliers from data distribution patterns, compensating for the limitations of the threshold method and ensuring comprehensive denoising.

[0084] Field mapping rules for format conversion: A detailed field mapping table is established to clearly define the correspondence between each field in different original data formats and the target fields in the JSON format. The soil nitrogen content (mg / kg) field in the Excel spreadsheet is mapped to `soil_nitrogen_content` in the JSON, with the unit standardized to g / kg. For multi-valued fields, such as multi-band data from remote sensing imagery, they are stored in the JSON as arrays, with band descriptions added. After conversion, an automated script performs field integrity checks to ensure that each required field has a corresponding value. Missing fields are marked as needing to be added, and their sources are recorded for future data completion.

[0085] A collaborative mechanism for semantic parsing and metadata tag generation: During semantic parsing, a semantic dictionary specific to the forestry and grassland domain is constructed, containing standardized vocabulary such as forestry and grassland species names, environmental parameter terms, and scientific research indicator definitions to ensure consistency in parsing. Combined with a domain knowledge graph, entities in the data are associated with nodes in the knowledge graph, and an NDVI value of 0.8 is associated with the semantic concept of vigorous vegetation growth. Metadata tag generation adopts a hierarchical structure: first-level tags include data type and collection area; second-level tags refine to specific parameters, such as soil nutrients—phosphorus content; third-level tags record data quality levels (excellent, good, medium, poor) and processing status (raw, denoised, transformed). After tag generation, the tags are stored in a metadata database and indexed with the original data, supporting rapid data retrieval and filtering by tag.

[0086] Step 103: Capture the long-distance semantic dependencies of forestry and grassland processing data through the multi-head attention mechanism in the Transformer model, classify the project data, and obtain classified forestry and grassland data;

[0087] Specifically, in this embodiment, the results of multiple attention heads in the Transformer model are concatenated and linearly transformed to capture long-distance semantic dependencies in the data;

[0088] Based on the captured long-distance semantic dependencies, the forestry and grassland processing data are classified, and the classification categories include forestry and grassland species, growth stage, and geographical location.

[0089] Specifically,

[0090] Training and optimization details of Transformer models

[0091] Training data preparation: High-quality samples are selected from historical data, covering different forest and grassland types, growth stages, and environmental conditions to ensure sample diversity and representativeness. Data augmentation is performed on the samples by expanding the training set through random cropping and feature perturbation (slightly adjusting parameters such as temperature and humidity) to avoid model overfitting. The training set is divided into a training subset and a validation subset in a 7:3 ratio. The training subset is used for model parameter learning, and the validation subset is used for real-time monitoring of model training performance. Training is stopped when the accuracy of the validation set does not improve for 5 consecutive epochs to prevent overtraining.

[0092] Feature allocation in the multi-head attention mechanism: Each attention head focuses on a specific dimension of feature combination. Head 1 focuses on the relationship between temperature, humidity, and growth stage; Head 2 focuses on the relationship between soil nutrients, species type, and geographical location; and Head 3 focuses on the interaction between light intensity, vegetation index, and growth rate. Attention weights are visualized to analyze the focus of each head. If the weight of a particular attention head is dispersed and its contribution is low, its feature input dimension is adjusted to enhance the model's ability to capture key features. The output of the multi-head attention mechanism is standardized through a LayerNormalization layer to stabilize parameter updates during training and accelerate model convergence.

[0093] Forestry and grassland data are rich and diverse, encompassing images, text, sensor data, and more, with different data types containing complex and varied feature information and correlations. Multi-head attention mechanisms, by distributing features across multiple attention heads, can process different aspects of feature information in parallel. Each attention head can focus on capturing specific types of feature associations. For example, when processing forestry and grassland image data, some attention heads can focus on the color features of vegetation, while others focus on its shape and texture features, thus comprehensively and meticulously mining key information in the data and avoiding important features that a single attention mechanism might miss. Furthermore, forestry and grassland data are affected by factors such as season, climate, and geographical location, resulting in dynamic feature distribution. Multi-head attention mechanisms, with their powerful adaptive feature allocation capabilities, can dynamically adjust the focus of each attention head according to different input data. For instance, the growth status of forestry and grassland varies significantly in different seasons; this mechanism can automatically allocate more attention to features related to the current season, improving classification accuracy and adaptability. Multi-head attention mechanisms can also effectively handle long-distance dependencies. In forestry and grassland data, some seemingly unrelated features may be closely related in deep semantics. Multi-head attention mechanisms can cross long distances in the data to discover these potential relationships, providing a more comprehensive and accurate basis for classification, thereby improving the performance of classification models in complex forestry and grassland data scenarios and providing strong technical support for the scientific management and accurate classification of forestry and grassland resources.

[0094] End-to-end management of data classification: Dynamic adjustment of classification rules: Based on the initial classification results, a classification correction rule base is built by combining expert experience. When the model misclassifies pine trees with a height of 2 meters and a diameter at breast height of 5 centimeters as growing trees, a rule is added to classify pine trees with a height of <3 meters and a diameter at breast height of <8 centimeters as seedlings, and the classification results are corrected a second time through the rule engine. New classification samples and feedback information are collected regularly, and the model is updated using incremental learning to adapt the classification rules to the dynamic changes in forestry and grassland data.

[0095] Construction of a multi-dimensional classification system: In addition to basic classification categories, composite classification dimensions are added, such as combined categories of arid regions, mature stages, poplar wetlands, seedling stages, and reeds, to meet the needs of refined management. Classification results are stored hierarchically, with the upper level being major categories (by species) and the lower level being subcategories (by growth stage), supporting data aggregation and analysis at different levels.

[0096] Comprehensive evaluation of classification performance: In addition to cross-validation, confusion matrix analysis was introduced to statistically analyze the precision, recall, and F1 score for each class, with a focus on the precision of rare species classification to ensure accurate identification. The Kappa coefficient was used to assess the consistency between the classification results and the manually labeled results; a Kappa value ≥ 0.8 was considered excellent. Simultaneously, feedback from forestry and grassland researchers was collected through user feedback channels to form subjective evaluation indicators. These indicators, combined with objective indicators, comprehensively measured the classification quality and provided direction for model optimization.

[0097] Step 104: Establish an LSTM prediction model based on the LSTM long short-term memory network, and optimize the hyperparameters of the LSTM prediction model using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model.

[0098] Specifically, in this embodiment, a chaotic initialization method is used to generate the initial population, and the exploration and development capabilities of the adaptive weight factor balancing algorithm are introduced to obtain the improved GWO gray wolf optimization algorithm.

[0099] The improved GWO algorithm is used to optimize the hyperparameters of the LSTM prediction model, including the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient.

[0100] Determine the search range and fitness function of hyperparameters, initialize the gray wolf population, with each individual representing a set of hyperparameter combinations, calculate the fitness value of each individual, and determine the α, β, and δ wolves in the wolf pack based on the fitness value;

[0101] The positions of other individuals in the population are updated according to the predation behavior of gray wolves, and the hyperparameter combination is iteratively output to obtain the target GWO-LSTM prediction model.

[0102] Specifically,

[0103] LSTM network characteristics: LSTM (Long Short-Term Memory) is a special type of recurrent neural network that can effectively process time-series data and solve the gradient vanishing or gradient exploding problems existing in traditional RNNs. It controls the flow of information through gating mechanisms (input gate, forget gate, output gate), enabling it to remember long-term dependent information. It is suitable for processing time-series data such as forestry and grassland data, including growth trends of forests and grasslands and changing trends of environmental factors.

[0104] Improved GWO Algorithm: GWO (Grey Wolf Optimization) is an optimization algorithm based on the predatory behavior of grey wolves in nature. The improved GWO algorithm builds upon the traditional algorithm, potentially improving aspects such as population initialization, search strategy, and convergence speed. It employs chaotic initialization to generate the initial population, increasing population diversity; and introduces an adaptive weighting factor to balance the algorithm's exploration and development capabilities, thereby improving optimization accuracy and convergence speed.

[0105] Hyperparameter optimization process: The improved GWO algorithm is used to optimize the hyperparameters of the LSTM prediction model, such as the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient. The specific process is as follows:

[0106] Determine the search range of hyperparameters and the fitness function (aiming to minimize prediction error). Initialize the gray wolf population, with each individual representing a set of hyperparameter combinations. Calculate the fitness value of each individual, and determine the α, β, and δ wolves in the pack based on the fitness values. Update the positions of other individuals in the population according to the gray wolves' predation behavior, i.e., update the hyperparameter combinations. Repeat the iterative process until the maximum number of iterations is reached or the fitness value converges, yielding the optimal hyperparameter combination.

[0107] Target GWO-LSTM Prediction Model: Optimized hyperparameters are applied to an LSTM network to construct a target GWO-LSTM prediction model. This model combines the advantages of LSTM in processing time-series data with the optimization capabilities of the improved GWO algorithm, thereby enhancing the accuracy of trend prediction for forestry and grassland data.

[0108] Applying the improved GWO algorithm to optimize the LSTM prediction model demonstrates significant technical advantages in the specific scenario of trend prediction for categorical forestry and grassland data. The improved GWO algorithm employs chaotic initialization to generate the initial population, greatly increasing population diversity and avoiding getting trapped in local optima, laying the foundation for subsequent searches for the globally optimal hyperparameter combination. Introducing an adaptive weighting factor dynamically balances exploration and development capabilities according to the algorithm's search progress. In the early stages of the search, it enhances exploration capabilities, broadly searching the hyperparameter space and discovering potential high-quality regions; in the later stages, it strengthens development capabilities, finely searching high-quality regions, improving optimization accuracy and convergence speed, and quickly locating the optimal hyperparameter combination. When optimizing the hyperparameters of the LSTM prediction model, key parameters such as the learning rate and the number of hidden layer nodes can be accurately determined to achieve optimal model performance. For categorical forestry and grassland data, which contains complex temporal features and multi-dimensional information, the optimized GWO-LSTM prediction model can better capture long-term dependencies and potential patterns in the data. When predicting species distribution trends, it can fully consider the temporal changes in forestry and grassland growth influenced by environmental factors, combining historical data with current characteristics to improve prediction accuracy. Moreover, given the differences in forestry and grassland data across different regions and species, the improved algorithm can flexibly adjust search strategies to adapt to diverse scenarios and provide accurate predictions for different categories of forestry and grassland data.

[0109] Step 105: Use the target GWO-LSTM prediction model to predict the trend of categorized forest and grassland data and output the species distribution change trend.

[0110] Specifically, in this embodiment, the classified forest and grassland data are input into the target GWO-LSTM prediction model. Based on the temporal characteristics and historical change patterns of the data, the distribution change trend of forest and grassland species is predicted, and the species distribution change trend is output.

[0111] Based on the changing trends of species distribution, dynamic data visualization charts are generated through generative adversarial networks (GANs), and combined with augmented reality (AR) technology to provide decision-making support for managers.

[0112] Furthermore, user permissions are dynamically adjusted based on DRL deep reinforcement learning in the cloud computing platform, and privacy is protected when sharing data across departments using federated learning.

[0113] Specifically,

[0114] Pre-processing of Classified Forestry and Grassland Data

[0115] Data temporal alignment: The classified forestry and grassland data may exhibit inconsistent collection time intervals (some sensors collect data daily, while remote sensing data is updated monthly), necessitating a unified temporal granularity. Interpolation is employed to fill in missing intermediate values: short-term gaps (1-3 days) are handled using linear interpolation to maintain a smooth data transition; long-term gaps (more than one week) are interpolated based on historical trends from the same period to ensure temporal continuity. The processed data is then arranged at fixed time intervals (weekly) to form a regular temporal sequence.

[0116] Feature selection and weight allocation: Not all categorical data have an equal impact on prediction. Through feature importance analysis, features strongly correlated with species distribution are selected. For moisture-loving species, the weights of humidity and precipitation features are increased; for light-loving species, the weights of light intensity and vegetation index features are increased. The selected features are combined according to their weight ratios and used as model input to reduce the interference of redundant information on prediction.

[0117] Generative Adversarial Networks (GANs) possess powerful data generation and simulation capabilities. They can deeply learn complex patterns and potential laws in historical forestry and grassland data, accurately capturing the growth and migration characteristics of different species under various environmental factors. Based on this, they generate highly realistic future forestry and grassland data scenarios, providing a solid data foundation for trend prediction and making the prediction results more reliable and forward-looking. In generating dynamic data visualization charts, they can present species distribution trends in an intuitive and vivid way. Traditional charts are often static and simple, while dynamic charts generated by GANs can clearly show the spatial diffusion, contraction, or migration processes of species over time. For example, through color gradations and dynamic graphic movement, they can intuitively present the dynamic trajectory of a rare species migrating from its original habitat to a more suitable area due to climate change, allowing managers to quickly grasp key information. GANs can generate visualization results under various hypothetical scenarios. Faced with uncertain future environmental changes, such as extreme climates and human disturbance, they can simulate the distribution changes of forestry and grassland species under different scenarios, providing rich references for decision-making. Moreover, the generated charts are highly interactive, allowing users to freely adjust parameters and switch perspectives to delve into the impact of different factors on species distribution, thereby formulating more targeted and scientific strategies for the protection and management of forest and grassland resources and promoting the sustainable development of forest and grassland ecosystems.

[0118] Application process of target GWO-LSTM prediction model

[0119] Multi-scenario forecast settings: Select the forecast duration according to management needs. Short-term forecasts (1-6 months) focus on seasonal changes, such as species dispersal trends before the rainy season; medium-term forecasts (1-3 years) focus on the gradual impact of environmental factors, such as the potential changes in species distribution due to rising annual average temperature; long-term forecasts (5-10 years) combine climate change models to infer the migration direction of species habitats. Adjust the time window length of the model input for different scenarios. Long-term forecasts require inputting longer historical data (data from the last 10 years).

[0120] Dynamic correction of prediction results: After the model outputs preliminary prediction results, it is corrected by combining field observation data. If a species is predicted to migrate to higher altitude areas, but field surveys reveal geographical barriers (cliffs) in those areas, the predicted path is adjusted based on the barrier information to make the results more consistent with actual terrain constraints. Simultaneously, the model is periodically retrained with newly collected forestry and grassland data to update prediction parameters and ensure the timeliness of long-term predictions.

[0121] Output and interpretation of species distribution change trends

[0122] Diverse output formats: Prediction results are presented in a visual format, including dynamic heat maps (showing changes in species distribution density at different times), migration path lines (marking the direction of movement of species' core distribution areas), and trend data tables (quantifying the proportion of increase or decrease in species numbers in each region). For rare species, a separate key monitoring report is generated to highlight the risk of shrinkage or expansion of their distribution range.

[0123] Supporting information for trend interpretation: The output results include an analysis of influencing factors, illustrating the key driving factors leading to distribution changes. For example, the northward shift of coniferous forests in a certain region is mainly affected by the increase in average annual temperature, with related environmental factors contributing up to 65%. An uncertainty assessment is also provided, indicating the confidence interval of the prediction results. For example, it indicates that the distribution range of this species will expand by 10%-15% in the next 3 years (90% confidence level), helping policymakers understand the reliability of the prediction.

[0124] Its beneficial effects are as follows: 1. Significantly improves data processing efficiency. Compared with traditional manual processing methods, it shortens data processing time, enabling rapid response to the processing needs of massive amounts of forestry and grassland data, improving the accuracy of data analysis and the efficiency of the entire lifecycle management of forestry and grassland science and technology data. 2. Through the full lifecycle management and trend prediction of forestry and grassland data, it is possible to understand the dynamic information such as the growth status of forestry and grassland and changes in species distribution in real time. It allows for the timely detection of potential spread trends of pests and diseases, abnormal changes in the forestry and grassland growth environment, and provides time for taking corresponding prevention and protection measures. 3. Accurate prediction of species distribution trends can provide a basis for the rational planning and utilization of forestry and grassland resources. Based on the prediction results, logging plans can be rationally arranged to avoid ecological damage caused by over-logging; and the scope and layout of protected areas can be planned in advance based on the possible migration directions of species to protect biodiversity.

[0125] Please see Figure 2 In the cloud-based forestry and grassland science and technology data analysis and prediction method, the initial forestry and grassland data is transmitted to a cloud computing platform. The Spark framework is used to perform real-time noise reduction, format conversion, and semantic parsing of the data to generate unified metadata tags. The resulting forestry and grassland processed data includes the following steps:

[0126] Step 201: Transmit the initial forestry and grassland data to the cloud computing platform via the network using encrypted transmission; set a data range threshold, use the Spark framework to treat data exceeding the range threshold as noise and remove it, and convert it into JSON format data;

[0127] Step 202: By analyzing the context and meaning of the JSON format data and combining it with the domain knowledge graph, perform semantic parsing on the data to generate preliminary semantic identifiers;

[0128] Step 203: Organize and standardize semantic identifiers according to unified metadata standards to generate unified metadata tags. Metadata tags should include at least the data source, collection time, data type, and semantic information.

[0129] Please see Figure 3 In the cloud-based forestry and grassland science and technology data analysis and prediction system, the forestry and grassland science and technology data full life cycle management system includes the following modules:

[0130] The forestry and grassland data acquisition module is used to collect forestry and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, and extract unstructured text information by combining NLP natural language processing technology to obtain initial forestry and grassland data.

[0131] The forestry and grassland data processing module is used to transmit the initial forestry and grassland data to the cloud computing platform. The Spark framework is used to perform real-time noise reduction, format conversion and semantic parsing of the data to generate unified metadata tags, thus obtaining the processed forestry and grassland data.

[0132] The forestry and grassland data classification module is used to capture the long-distance semantic dependencies of forestry and grassland processing data through the multi-head attention mechanism in the Transformer model, classify the project data, and obtain classified forestry and grassland data.

[0133] The prediction model building module is used to build an LSTM prediction model based on the LSTM long short-term memory network, and to optimize the hyperparameters of the LSTM prediction model using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model.

[0134] The species distribution analysis module is used to predict the trend of classified forest and grassland data using the target GWO-LSTM prediction model and output the species distribution change trend.

[0135] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A cloud computing-based method for analyzing and predicting forestry and grassland science and technology data, characterized in that, The method for managing the entire lifecycle of forestry and grassland science and technology data includes the following steps: Forestry and grassland data are collected through IoT sensors, remote sensing equipment and scientific research platforms, and unstructured text information is extracted by combining NLP natural language processing technology to obtain initial forestry and grassland data. The initial forest and grassland data is transmitted to a cloud computing platform, and the Spark framework is used to perform real-time noise reduction, format conversion, and semantic parsing on the data to generate unified metadata tags, thus obtaining the processed forest and grassland data. By capturing the long-distance semantic dependencies of the forestry and grassland processing data through the multi-head attention mechanism in the Transformer model, the project data is classified to obtain classified forestry and grassland data. An LSTM prediction model is established based on the LSTM long short-term memory network. The hyperparameters of the LSTM prediction model are optimized using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model. The target GWO-LSTM prediction model is used to predict the trend of the classified forest and grassland data, and the species distribution change trend is output.

2. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The initial forestry and grassland data is obtained by collecting forestry and grassland data through IoT sensors, remote sensing equipment, and scientific research platforms, and then extracting unstructured text information using NLP (Natural Language Processing) technology, including: Collect forestry and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, including at least forestry and grassland growth environment data, forestry and grassland coverage area and growth status data, forestry and grassland pest and disease situation and seed germination rate data. By combining NLP (Natural Language Processing) technology to segment the collected data into word units, the text is divided into word segmentation data. The segmented data is labeled with part-of-speech tags to determine the part of speech of each word. Entity recognition is performed on the labeled words to identify forest and grassland species, place names, and time information in the text, thus obtaining initial forest and grassland data.

3. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The initial forestry and grassland data is transmitted to a cloud computing platform, where the Spark framework is used to perform real-time noise reduction, format conversion, and semantic parsing to generate unified metadata tags, resulting in processed forestry and grassland data, including: The initial forestry and grassland data is transmitted to the cloud computing platform via network using encrypted transmission; a range threshold is set for the data, and data exceeding the range threshold is treated as noise and removed using the Spark framework, and then converted into JSON format data; By analyzing the context and meaning of JSON format data and combining it with domain knowledge graphs, semantic parsing is performed on the data to generate preliminary semantic identifiers; Semantic identifiers are organized and standardized according to a unified metadata standard to generate unified metadata tags. Metadata tags include at least the data source, collection time, data type, and semantic information.

4. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The process involves capturing long-distance semantic dependencies in the forestry and grassland processing data using the multi-head attention mechanism in the Transformer model, classifying the project data, and obtaining categorized forestry and grassland data, including: By concatenating and linearly transforming the results of multiple attention heads in the Transformer model, long-distance semantic dependencies in the data can be captured. Based on the captured long-distance semantic dependencies, the forestry and grassland processing data are classified, and the classification categories include forestry and grassland species, growth stage, and geographical location.

5. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The LSTM prediction model based on the LSTM long short-term memory network is established, and the hyperparameters of the LSTM prediction model are optimized using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model, including: An initial population is generated using chaotic initialization, and the exploration and development capabilities of an adaptive weight factor balancing algorithm are introduced to obtain an improved GWO (Grey Wolf) optimization algorithm.

6. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The process of establishing an LSTM prediction model based on an LSTM long short-term memory network, optimizing the hyperparameters of the LSTM prediction model using an improved GWO (Grey Wolf) optimization algorithm to obtain the target GWO-LSTM prediction model, further includes: The improved GWO algorithm is used to optimize the hyperparameters of the LSTM prediction model, including the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient. Determine the search range and fitness function of hyperparameters, initialize the gray wolf population, with each individual representing a set of hyperparameter combinations, calculate the fitness value of each individual, and determine the α, β, and δ wolves in the wolf pack based on the fitness value; The positions of other individuals in the population are updated according to the predation behavior of gray wolves, and the hyperparameter combination is iteratively output to obtain the target GWO-LSTM prediction model.

7. The cloud computing-based forestry and grassland science and technology data analysis and prediction method as described in claim 1, characterized in that, The step of using the target GWO-LSTM prediction model to predict trends in the classified forest and grassland data and outputting species distribution change trends includes: The classification of forest and grassland data is input into the target GWO-LSTM prediction model. Based on the temporal characteristics and historical change patterns of the data, the distribution change trend of forest and grassland species is predicted, and the species distribution change trend is output. Based on the changing trends of species distribution, dynamic data visualization charts are generated through generative adversarial networks (GANs), and combined with augmented reality (AR) technology to provide decision-making support for managers. Furthermore, user permissions are dynamically adjusted based on DRL deep reinforcement learning in the cloud computing platform, and privacy is protected when sharing data across departments by combining federated learning.

8. A cloud-based forestry and grassland science and technology data analysis and prediction system, characterized in that: The forestry and grassland science and technology data lifecycle management system includes the following modules: The forestry and grassland data acquisition module is used to collect forestry and grassland data through IoT sensors, remote sensing equipment and scientific research platforms, and extract unstructured text information by combining NLP natural language processing technology to obtain initial forestry and grassland data. The forestry and grassland data processing module is used to transmit the initial forestry and grassland data to the cloud computing platform, and use the Spark framework to perform real-time noise reduction, format conversion and semantic parsing of the data to generate unified metadata tags, thereby obtaining the processed forestry and grassland data. The forestry and grassland data classification module is used to capture the long-distance semantic dependencies of the forestry and grassland processed data through the multi-head attention mechanism in the Transformer model, classify the project data, and obtain classified forestry and grassland data. The prediction model building module is used to build an LSTM prediction model based on the LSTM long short-term memory network, and optimize the hyperparameters of the LSTM prediction model using the improved GWO gray wolf optimization algorithm to obtain the target GWO-LSTM prediction model. The species distribution analysis module is used to perform trend prediction on the classified forest and grassland data using the target GWO-LSTM prediction model and output the species distribution change trend.

9. The cloud computing-based forestry and grassland science and technology data analysis and prediction system as described in claim 8, characterized in that, The species distribution analysis module includes the following sub-modules: An improved submodule is used to generate the initial population using chaotic initialization. The ability to explore and develop an adaptive weight factor balancing algorithm is introduced, resulting in an improved GWO (Grey Wolf) optimization algorithm.

10. The cloud computing-based forestry and grassland science and technology data analysis and prediction system as described in claim 8, characterized in that, The prediction model building module includes the following sub-modules: The optimization submodule is used to optimize the hyperparameters of the LSTM prediction model using the improved GWO algorithm, including the learning rate, number of hidden layer nodes, number of iterations, and regularization coefficient. The parameter submodule is used to determine the search range and fitness function of hyperparameters, initialize the gray wolf population, with each individual representing a set of hyperparameter combinations, calculate the fitness value of each individual, and determine the α, β, and δ wolves in the wolf pack based on the fitness value; The output submodule is used to update the positions of other individuals in the population according to the predation behavior of the gray wolves, iteratively output the hyperparameter combination, and obtain the target GWO-LSTM prediction model.

Citation Information

Patent Citations

  • Tree growth intelligent monitoring system based on cloud computing platform

    CN117557400A