Transform agricultural pest intelligent prediction model based on text big data and state space
By adopting a Transformer model based on text big data and state space in agricultural pest monitoring, the problems of low efficiency, high cost, limited accuracy and lack of real-time dynamic prediction capabilities in the existing technology are solved, and efficient, low-cost, accurate and real-time pest prediction are achieved to support farmers' decision-making.
Patent Information
- Application Number
- CN202510069977.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing agricultural pest and disease monitoring technologies have problems such as low efficiency, high cost, limited accuracy and generalization capabilities, and lack of real-time dynamic prediction capabilities.
The Transformer model based on text big data and state space is adopted to collect and label crop disease image data, preprocess and information extraction, build a knowledge graph, and set state space transfer equation, state space based Transformer model and state space loss function to achieve efficient, low-cost, accurate and real-time dynamic prediction of diseases and pests.
It reduces the cost and operational complexity of pest monitoring, improves the accuracy and real-time prediction, realizes the transformation from passive response to active prevention, provides farmers with better decision-making support, and reduces the losses caused by pests and diseases.
Smart Images

Figure CN120012987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural science and technology and data analysis technology, and in particular to a Transformer agricultural pest intelligent prediction model based on text big data and state space. Background Art
[0002] In modern agricultural production, the monitoring and prevention of crop diseases and pests is one of the key issues, which is directly related to the efficiency of agricultural production and the yield and quality of crops. With the development of information technology and biotechnology, the prediction and control of agricultural diseases and pests has gradually developed from traditional manual identification and experience processing to a more scientific, precise and automated direction. However, existing technologies still face many limitations and challenges in practical applications.
[0003] Traditional pest and disease monitoring methods mainly rely on the experience of agricultural workers and regular field inspections. Although these methods can control the spread of pests and diseases to a certain extent, they have problems such as low efficiency, time-consuming and labor-intensive, and low accuracy. Agricultural workers need to conduct one-by-one inspections in large areas of farmland, which is not only inefficient, but also difficult to achieve comprehensive coverage and easy to miss or misjudge. In addition, this manual-dependent method is limited by professional knowledge and experience, and is particularly unfriendly to novices or non-professionals. Although high-tech products such as drones, satellite images, and automated monitoring equipment have begun to be used in modern agriculture to identify and predict pests and diseases, these technologies have improved the accuracy and scope of monitoring, but their high equipment costs and operational complexity limit the use of ordinary farmers. For example, although drones and remote sensing technologies can quickly cover large areas of farmland and provide real-time data, they require professional operation and maintenance, as well as high-performance computing resources required to process high-resolution images. Even when large amounts of data are available, how to effectively process and utilize these data remains a challenge. Existing agricultural decision support systems often rely on simple data processing models and cannot fully explore the complex patterns and trends in the data, which to a certain extent reduces the accuracy and reliability of predictions. In addition, existing models often cannot be updated in real time and cannot adapt to rapidly changing environmental conditions and the development of pests and diseases. These models are often very sensitive to the quality of data and the accuracy of labeling. Small changes or errors in the data may lead to large deviations in the prediction results.
[0004] As agricultural production develops towards efficiency and precision, the demand for pest and disease monitoring systems that can perform real-time dynamic predictions is increasing. However, most existing technologies can only provide static predictions, making it difficult to track and predict the development of pests and diseases in real time, which limits the response speed and timeliness of agricultural management.
[0005] In summary, although the existing technologies have made some progress in agricultural pest monitoring and management, they still have problems such as low efficiency, high cost, limited accuracy and generalization ability, and lack of real-time dynamic prediction capabilities. Therefore, it is necessary to propose a new state-space-based Transformer model, which aims to achieve efficient, low-cost, high-accuracy and real-time dynamic prediction of crop pests and diseases through the combination of deep learning and state-space theory to meet the high standards of modern agricultural production. Summary of the invention
[0006] The purpose of the present invention is to provide a Transformer agricultural pest and disease intelligent prediction model based on text big data and state space to solve the problems existing in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides a Transformer agricultural pest and disease intelligent prediction model based on text big data and state space, comprising the following steps:
[0008] S1. Collect a large amount of crop disease image data and annotate the collected image data;
[0009] S2, preprocessing the labeled data;
[0010] S3, extract information from the preprocessed data and then construct a knowledge graph;
[0011] S4, setting the state space transfer equation of the model, the state space based Transformer model and the state space loss function;
[0012] S5: The data processed by S3 is used as the input of the model. After being processed by the model, the token embedding is output to predict the future state.
[0013] Preferably, S1 uses web crawler technology to collect image data to ensure the accuracy and consistency of the data. The data is annotated using the open source annotation tool LabelImg, which is a graphical image annotation tool with a good user interface and simple operation. It can easily draw bounding boxes on images and provide corresponding labels, thereby achieving accurate annotation of target objects in images. The latest version of LabelImg has been optimized in terms of function, performance and stability, and can better meet image annotation needs. Specifically, it includes:
[0014] To ensure the quality of annotation, first, establish an annotation guideline that contains annotation standards and procedures for various data types;
[0015] Secondly, in order to ensure the consistency and accuracy of annotation, a dual annotation strategy is adopted, which includes: preliminary annotation, which is independently annotated by several annotators, and several annotators regularly conduct cross-checks and discussions to solve complex problems that arise in the annotation process and standardize annotation standards; annotation review, in which the reviewers verify and arbitrate the annotators' annotations. To further ensure the quality of data annotation, a comprehensive quality control mechanism is established, including regular training of annotators to familiarize them with annotation guidelines and standards, improve the accuracy and consistency of annotations, regularly check annotated data to evaluate the quality of annotations, and adjust annotation guidelines and procedures based on the inspection results. A feedback mechanism is established to encourage annotators to raise questions and suggestions during the annotation process and continuously optimize the annotation process.
[0016] Preferably, the preprocessing task is to clean and format the data to improve the accuracy and efficiency of subsequent tasks. The data preprocessing in S2 specifically includes:
[0017] For images without clear boundaries, segmentation is more complicated and relies on a specific segmentation algorithm to identify word boundaries. The segmentation formula is:
[0018] W = Segment(T);
[0019] Where T represents the original data image after S1 processing; W represents the segmentation result; Segment(·) represents the segmentation function;
[0020] Cleaning involves removing noise and unnecessary information from the image. The cleaned image is more standardized, which is conducive to model processing. The formula is:
[0021] T′=Clean(T);
[0022] Where, T' is the cleaned image; Clean(·) is the cleaning operation;
[0023] Normalized, the formula is:
[0024] T″=Normalize(T′);
[0025] Where T” is the normalized image; Normalize(·) is the normalization function.
[0026] Preferably, information extraction in S3 is to identify and extract entities, attributes and relationships from preprocessed data. Entity extraction is used to identify named entities in text, and relationship extraction determines the semantic connection between entities, which is crucial for building knowledge graphs. Entity extraction can identify named entities in text, such as names of people, places and organizations; relationship extraction determines the semantic connection between entities; attribute extraction focuses on descriptive information about entities; the extraction of entities and relationships is expressed as:
[0027] Unified Knowledge=
[0028] Knowledge Fusion (Extracted Information);
[0029] Using TF-IDF and WordVec to construct knowledge graphs can significantly improve the information retrieval capability and semantic parsing efficiency of knowledge graphs. TF-IDF evaluates the importance of a word by multiplying the term frequency TF of the word by the inverse document frequency IDF. The term frequency TF is the number of times a word appears in a document divided by the total number of words in the document, while the inverse document frequency IDF is the total number of documents divided by the logarithm of the number of documents containing the word. WordVec converts words into vector form through training, captures the contextual relationship between words, and predicts the current word from the context through the Skip-gram model.
[0030] Preferably, in S4, the state space transfer equation is introduced to enhance the prediction ability of the complex system, which is expressed as:
[0031]
[0032] y t =Ch t +Dx t ;
[0033] Among them, h t represents the hidden state at time t; x t represents the input at time t; y t Represents the output of the model; and ΔB represent the updates of the state transition matrix and the input control matrix, respectively; C and D represent the output matrix and the direct transfer matrix, respectively;
[0034] Use the exponential map to update the state transition matrix as follows:
[0035]
[0036] The rationale for designing with state-space transfer equations is that it can simulate and predict dynamic changes in actual complex systems. Through the state-space model, the changes in system states over time can be captured more accurately, which is crucial for complex system predictions based on time series data.
[0037] Preferably, the state-space-based Transformer model in S4 specifically includes:
[0038] The block consists of several blocks, each of which is a variant of the layer in the Transformer model, which integrates the characteristics of the state-space model. Each block contains a self-attention mechanism and a feed-forward neural network, and also embeds a state space for the representation layer to simulate the time evolution characteristics of the input data;
[0039] Inter-block connections,Blocks are connected via residual connections and layer normalization.,Residual connections allow information to flow directly from one block to another,,while layer normalization helps maintain stability during training.
[0040] Preferably, the state space loss function in S4 is designed for a new prediction model that integrates the state space model with the Transformer architecture. It is very different from the traditional Transformer model loss function, mainly in that it considers both the dynamic characteristics of the time series and the accuracy of the prediction. The state space loss function not only measures the difference between the predicted output and the true value, but also considers the smoothness and coherence of the model state transition. Its mathematical expression is:
[0041] L(θ)=λ1L predict (θ)+λ2L smooth (θ);
[0042] Among them, L predict (θ) is the traditional prediction loss component; the mean square error (MSE) is usually used to quantify the difference between the model output and the true value, expressed as:
[0043]
[0044] Where N is the total number of data points; y t is the true value at time t; is the prediction output of the model;
[0045] L smooth (θ) is the smoothness loss component of the state space model, which is used to ensure the continuity and rationality of the state transfer. A typical choice is to use the Frobenius norm of the state transition matrix change as a smoothing term, as follows:
[0046]
[0047] Among them, ΔA represents the change of the state transition matrix; ||·|| Frepresents the Frobenius norm; λ1 and λ2 represent weighting parameters for balancing the contribution of the prediction loss component and the smoothness loss component to the total loss. The selection of these parameters depends on the specific task and data characteristics and needs to be determined by cross-validation or other model selection techniques. The design of the state-space loss function takes into account not only the importance of prediction accuracy, but also the importance of the smooth evolution of the system state over time in the complex system prediction task. The state-space loss function encourages the model to capture the dynamic characteristics of the system through terms. In addition, relying solely on prediction loss during model training usually leads to overfitting. By introducing smoothness loss, the state-space loss function enhances the generalization ability of the model and makes it more robust on unseen data. In practical applications, there may be a conflict between prediction accuracy and the smoothness of state transitions. By adjusting λ1 and λ2, the state-space loss function allows the best balance between the two to be found according to actual needs.
[0048] Preferably, S5 specifically includes:
[0049] The processed text data is taken as input and converted into vector representation through an embedding layer (called input token embedding). These embeddings are designed to capture the semantic information of each word or character in the text. Learned context IDs are introduced in each module of the model to identify and learn different context states, enabling the model to distinguish and process information from various contexts. Each module contains a linear layer (for generating keys, values, and queries, called KVQ) and a cross-attention mechanism to integrate information based on the current input and previously learned context states. Finally, the model inputs token embeddings for predicting future states, where the token embedding conversion expression is:
[0050] E = Embed(X);
[0051] Where X is the input text data; E is the tag embedding;
[0052] The cross attention mechanism expression is:
[0053]
[0054] The state space model integral expression is:
[0055] S t+1 =F·S t +G·A+H;
[0056] Among them, S t and S t+1 They represent the system states of the current and next time steps respectively; F, G, H are the state space model parameters; A is the output of the cross attention.
[0057] Therefore, the present invention adopts the above-mentioned Transformer agricultural pest intelligent prediction model based on text big data and state space, which has the following beneficial effects:
[0058] (1) Reduce the cost and operational complexity of pest and disease monitoring. By using existing agricultural data and simple sensing equipment, the cost and complexity of technology use are reduced, making it more suitable for ordinary farmers;
[0059] (2) The accuracy and real-time performance of predictions are improved. By proposing a model that combines state-space theory and the Transformer architecture, it is possible to more accurately simulate and predict the development trends of pests and diseases, and maintain efficient processing speed and high accuracy even when the amount of data is huge and changes rapidly.
[0060] (3) Analyze data in real time to predict the possibility and timing of pest and disease occurrence, realize the transition from passive response to active prevention, provide farmers with customer decision support, and greatly reduce the losses caused by pests and diseases.
[0061] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a flow chart of the Transformer network structure of the present invention;
[0063] Figure 2 This is a schematic diagram of the text data preprocessing process of the present invention;
[0064] Figure 3 A flowchart of the knowledge graph construction / update process of the present invention;
[0065] Figure 4 The overall working flow diagram of the state transformer model of the present invention;
[0066] Figure 5 It is a schematic diagram of the working process of the present invention;
[0067] Figure 6 A schematic diagram of the network structure of the state space transformer of the present invention;
[0068] Figure 7 The figure is a flow chart of the conversion process of the present invention. DETAILED DESCRIPTION
[0069] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] See also Figure 1-7 , a Transformer agricultural pest and disease intelligent prediction model based on text big data and state space, including the following steps:
[0071] S1. Collect a large amount of crop disease image data. In order to ensure the richness and diversity of the data, a variety of devices are used for data collection, including but not limited to high-resolution digital cameras, drones, mobile phones, etc. To ensure image clarity, the resolution of the collected images is guaranteed to be above 1280*720. In order to collect the diversity of image types and cover a wide geographical distribution, data images are collected from different natural environments, including grasslands, deserts, and farmlands; web crawler technology is used to collect image data to ensure the accuracy and consistency of the data, and specific crawler scripts are written to automatically extract case texts and related information from these databases. In the implementation of the crawler program, appropriate strategies are adopted to avoid unnecessary burdens on the target website, such as setting a reasonable request interval and using a proxy server. At the same time, in order to improve the efficiency and coverage of the crawler, dynamic crawler technology is used to process JavaScript-generated content, and distributed crawler technology is used to accelerate the data collection process.
[0072] After the data collection is completed, the next step is to accurately label the data. We use the open source labeling tool LabelImg for labeling. As a graphical image labeling tool, it has a good user interface and simple operation. It can easily draw bounding boxes on images and provide corresponding labels, thereby achieving accurate labeling of target objects in images. The latest version of LabelImg has been optimized in terms of function, performance and stability, which can better meet the needs of image labeling. Specifically, it includes:
[0073] To ensure the quality of annotation, first, establish an annotation guideline that contains annotation standards and procedures for various data types;
[0074] Secondly, in order to ensure the consistency and accuracy of annotation, a dual annotation strategy is adopted, which includes: preliminary annotation, which is independently annotated by several annotators, and several annotators regularly conduct cross-checks and discussions to solve complex problems in the annotation process and standardize the annotation standards; in addition, some automated tools can be used to assist the annotation process. For example, natural language processing technology is used to pre-analyze the text content and automatically identify and mark key information to reduce the burden of manual annotation. However, considering the potential errors of automated tools, all automatic annotation results are manually reviewed and corrected to ensure the accuracy of data annotation. Annotation review, in which the reviewer verifies and arbitrates the annotator's annotations. To further ensure the quality of data annotation, a comprehensive quality control mechanism is established, including regular training of annotators to familiarize them with annotation guidelines and standards, improve the accuracy and consistency of annotations, regularly check the annotated data to evaluate the quality of annotations, and adjust the annotation guidelines and procedures based on the inspection results, establish a feedback mechanism to encourage annotators to raise questions and suggestions during the annotation process, and continuously optimize the annotation process.
[0075] S2, preprocess the labeled data; use data enhancement techniques such as information extraction, knowledge integration, entity disambiguation and coreference resolution to preprocess the image. In NLP, the original image usually contains noise and redundant information, which may affect the learning efficiency and performance of the model. The task of preprocessing is to clean and format the data to improve the accuracy and efficiency of subsequent tasks. The data preprocessing in S2 specifically includes:
[0076] For images without clear boundaries, segmentation is more complicated and relies on a specific segmentation algorithm to identify word boundaries. The segmentation formula is:
[0077] W = Segment(T);
[0078] Where T represents the original data image after S1 processing; W represents the segmentation result; Segment(·) represents the segmentation function;
[0079] Cleaning involves removing noise and unnecessary information from the image. The cleaned image is more standardized, which is conducive to model processing. The formula is:
[0080] T′=Clean(T);
[0081] Where, T' is the cleaned image; Clean(·) is the cleaning operation;
[0082] Normalized, the formula is:
[0083] T″=Normalize(T′);
[0084] Where T” is the normalized image; Normalize(·) is the normalization function.
[0085] S3. The preprocessed data is converted into a form that can be effectively processed by a computer program. Subsequently, information extraction is performed to identify and extract entities, attributes, and relationships from the preprocessed data. Entity extraction is used to identify named entities in the text. Relation extraction determines the semantic connection between entities, which is crucial for building a knowledge graph. Entity extraction can identify named entities in the text, such as names of people, places, and organizations; relationship extraction determines the semantic connection between entities; attribute extraction focuses on descriptive information about entities; the extraction of entities and relationships is expressed as:
[0086] Unified Knowledge=
[0087] Knowledge Fusion (Extracted Information);
[0088] Building a knowledge graph is a process of converting scattered data information into interrelated structured knowledge. Using TF-IDF (TermFrequency-InverseDocumentFrequency) and WordVec to build a knowledge graph can significantly improve the information retrieval capability and semantic parsing efficiency of the knowledge graph. TF-IDF can effectively identify keywords or phrases, which usually carry key information connecting different entities and attributes. For example, when building an agricultural document knowledge graph, TF-IDF can help extract key agricultural terms and related definitions, which are essential for building entities and relationships. The application of WordVec further deepens this process; by analyzing the similarity between word vectors, words with similar semantics can be explored and identified, which is particularly important for building semantic links in the knowledge graph. TF-IDF evaluates the importance of a word by multiplying the term frequency TF of the word by the inverse document frequency IDF, thereby reducing the influence of common words in the document and emphasizing the importance of uncommon words. The term frequency TF is the number of times a word appears in a document divided by the total number of words in the document, while the inverse document frequency IDF is the total number of documents divided by the logarithm of the number of documents containing the word; it helps to identify the key points for filtering information from a large number of images; WordVec converts words into vector form through training and captures the contextual relationship between words, mainly through the Skip-gram model to predict the context and the CBOW model to predict the current word from the context.
[0089] S4, setting the state space transfer equation of the model, the state space based Transformer model and the state space loss function;
[0090] The state space transfer equation is introduced to enhance the prediction capability of complex systems. The state space transfer equation is a core concept of dynamic system theory, which describes the evolution of system state over time. The state space transfer equation of this embodiment is expressed as:
[0091]
[0092] y t =Ch t +Dx t ;
[0093] Among them, h t represents the hidden state at time t; x t represents the input at time t; y t Represents the output of the model; and ΔB represent the updates of the state transition matrix and the input control matrix, respectively; C and D represent the output matrix and the direct transmission matrix, respectively; the learning and updating of these matrices correspond to the linear layers in the model. The linear layer assumption of the state-space model allows the use of an exponential mapping to update the state transition matrix, capturing the continuous changes in the system dynamics, as follows:
[0094]
[0095] This index graph ensures that a , ΔB allows the model to adjust x according to the input t , while C and D transform the hidden state into the final output. The rationale for designing with state-space transfer equations is that it can simulate and predict dynamic changes in actual complex systems. Through the state-space model, the changes in system states over time can be captured more accurately, which is crucial for complex system predictions based on time series data. Transformers integrated with state-space models can be enhanced and optimized in the following ways:
[0096] 1. Enhanced temporal dynamic modeling: The state space transfer equation enables the model to capture the dynamic changes of time series data more finely and improve the prediction accuracy.
[0097] 2. Flexible parameter update: The parameter update mechanism in the state-space model enables the model to flexibly adapt to the characteristics of different complex systems, thereby achieving targeted optimization.
[0098] 3. Enhanced data processing capabilities: The introduction of the state space model enables the model to process data containing complex dynamics, such as financial data analysis or natural language text, which often embodies implicit time series characteristics. The model can not only process static semantic information, but also capture the changes in data over time, which is crucial for dynamic prediction.
[0099] 4. Optimized state change modeling: The state space transfer equation provides a mathematically rigorous method to describe continuous state changes. Compared with the traditional Transformer structure, this model can more accurately predict future states in time series forecasting tasks.
[0100] 5. Custom output equation: By combining the output equation y of the state space model t =Ch t +Dx t With the Transformer’s output layer, the model’s output contains information determined by the current state and takes into account the impact of direct input, providing a more comprehensive prediction.
[0101] These enhancements not only improve the accuracy of the model, but also enhance the model's ability to predict the behavior of complex systems and understand the deep semantics of text data. In addition, by integrating the concept of state space, the model is endowed with the ability to process and predict dynamic changes in the system, which is a capability that traditional Transformer models do not have. The state space model provides a mathematical framework for describing and predicting the evolution of system states over time, which is critical for monitoring, predicting, and analyzing crop pests and diseases. Combined with the state space model, our Transformer becomes not only a powerful tool for processing text sequences, but also a model that can deeply analyze and predict the behavior of complex systems.
[0102] The state-space based Transformer model includes:
[0103] The block consists of several blocks, each of which is a variant of the layer in the Transformer model, which integrates the characteristics of the state-space model. Each block contains a self-attention mechanism and a feed-forward neural network, and also embeds a state space for the representation layer to simulate the time evolution characteristics of the input data;
[0104] Inter-block connections,Blocks are connected via residual connections and layer normalization.,Residual connections allow information to flow directly from one block to another,,while layer normalization helps maintain stability during training.
[0105] The specific parameters are: Number of blocks: The standard version of the model in this embodiment includes 12 blocks, each of which corresponds to a layer in the Transformer. Size of each block: Each block has 12 heads in the self-attention mechanism, each with a dimension of 64, and a total dimension of 768 for each self-attention layer. The intermediate dimension of the feedforward network is 3072. State space representation layer: The state space representation layer of each block is designed to track and update state variables that change over time, and is usually set to match the input dimension, i.e. 768. Total number of parameters: Our model has a total of approximately 110 million parameters, including the parameters of the self-attention layer, the feedforward network, the state space representation layer, and other components in the model.
[0106] The state-space loss function is designed for a new prediction model that integrates the state-space model with the Transformer architecture. It is very different from the traditional Transformer model loss function, mainly in that it takes into account both the dynamic characteristics of the time series and the accuracy of the prediction. The state-space loss function not only measures the difference between the predicted output and the true value, but also considers the smoothness and coherence of the model state transition. Its mathematical expression is:
[0107] L(θ)=λ1L predict (θ)+λ2L smooth (θ);
[0108] Among them, L predict (θ) is the traditional prediction loss component; the mean square error (MSE) is usually used to quantify the difference between the model output and the true value, expressed as:
[0109]
[0110] Where N is the total number of data points; y t is the true value at time t; is the prediction output of the model;
[0111] L smooth (θ) is the smoothness loss component of the state space model, which is used to ensure the continuity and rationality of the state transfer. A typical choice is to use the Frobenius norm of the state transition matrix change as a smoothing term, as follows:
[0112]
[0113] Among them, ΔA represents the change of the state transition matrix; ||·|| FRepresents the Frobenius norm; λ1 and λ2 represent weighted parameters for balancing the contribution of the prediction loss component and the smoothness loss component to the total loss. The selection of these parameters depends on the specific task and data characteristics and needs to be determined by cross-validation or other model selection techniques. The design of the state space loss function takes into account not only the importance of prediction accuracy, but also the importance of the smooth evolution of the system state over time in the complex system prediction task. The state space loss function encourages the model to capture the dynamic characteristics of the system through terms. In addition, relying solely on prediction loss during model training usually leads to overfitting. By introducing smoothness loss, the state space loss function enhances the generalization ability of the model, making it more robust on unseen data. In practical applications, there may be a conflict between prediction accuracy and the smoothness of state transitions. By adjusting λ1 and λ2, the state space loss function allows the best balance to be found between the two according to actual needs. Therefore, the prediction task of the state space loss function in this embodiment has the following advantages:
[0114] 1. Consistency with dynamic prediction and loss functions: The state space loss function is closely related to the goal of dynamic system prediction, focusing on both the accuracy of the prediction and the smoothness of the state changes during the prediction process. This is critical for dynamic systems because dynamic systems usually change state smoothly over time, and any sudden changes may indicate system anomalies or data problems. smooth The (θ) component effectively reduces the likelihood of such abrupt changes, making the model’s predictions more credible.
[0115] 2. Prevent overfitting: Traditional Transformer models may perform poorly on future data due to overfitting historical data. The smoothness loss component in the state-space loss function helps improve the performance of the model on unseen data by forcing the model to learn more general state change rules rather than just remembering the patterns in the training dataset.
[0116] 3. Flexibility of parameter adjustment: By adjusting λ1 and λ2, researchers can flexibly control the weights of prediction loss and smoothness loss in the total loss function. For different complex system prediction tasks, these parameters can be optimized according to the specific characteristics of the system and task requirements to achieve the best prediction performance.
[0117] 4. Enhanced interpretability: The state space loss function not only optimizes the model’s predictive performance, but also enhances the model’s interpretability. Since the model needs to consider the coherence of state transitions, the trends learned by the model are more consistent with the physical or logical laws of the real world, enabling researchers to explain the model’s predictive behavior by analyzing the parameters of the state space model.
[0118] S5: Take the data processed by S3 as the input of the model, process it, and output token embedding to predict the future state. Specifically:
[0119] The processed text data is taken as input and converted into vector representation through an embedding layer (called input token embedding). These embeddings are designed to capture the semantic information of each word or character in the text. Learned context IDs are introduced in each module of the model to identify and learn different context states, enabling the model to distinguish and process information from various contexts. Each module contains a linear layer (for generating keys, values, and queries, called KVQ) and a cross-attention mechanism to integrate information based on the current input and previously learned context states. Finally, the model inputs token embeddings for predicting future states, where the token embedding conversion expression is:
[0120] E = Embed(X);
[0121] Where X is the input text data; E is the tag embedding;
[0122] The cross attention mechanism expression is:
[0123]
[0124] The state space model integral expression is:
[0125] S t+1 =F·S t +G·A+H;
[0126] Among them, S t and S t+1 They represent the system states of the current and next time steps respectively; F, G, H are the state space model parameters; A is the output of the cross attention.
[0127] In this embodiment, the performance results of this model and other comparison models are shown in Table 1:
[0128] Table 1 Experimental results
[0129] Model Precision Recall Accuracy Transformer 0.80 0.77 0.79 BERT 0.83 0.80 0.82 wwm-BERT 0.86 0.83 0.85 Finsformer 0.89 0.86 0.88 ProposedMethod 0.93 0.90 0.91
[0130] Although the traditional Transformer model has advantages in processing sequential data, its performance is not as good as the best BERT-based model due to the complexity and specificity of images and text. The BERT model has deep bidirectional contextual understanding capabilities, and has achieved significant improvements in both precision and recall due to a better grasp of the semantics in the text. As an improved version of BERT, wwm-BERT further improves performance by adopting a more sophisticated pre-training method to better understand the subtle differences between dictionaries, especially in scenarios such as monitoring, prediction and analysis of crop pests and diseases, which are crucial. Although Finsformer is optimized for the financial field and does not fully match the monitoring, prediction and analysis applications of crop pests and diseases, its ability to capture complex and professional texts and images also shows promising performance. This model achieved the best results in all indicators. Its success not only combines the deep contextual understanding of BERT, but also enhances the ability to capture dynamic changes in images by integrating state-space models. In scenarios such as monitoring, prediction and analysis of crop pests and diseases, temporal changes are shown, and state-space models perform well. Therefore, compared with the traditional Transformer structure, this model can capture the time series changes in text data more finely. In addition, the introduction of the state space loss function also promotes the model's ability to capture the logic of case development and improves prediction accuracy. By modeling these complex relationships more rigorously, the proposed model not only achieves mathematical accuracy, but also more comprehensively captures and understands the complexity of the text for monitoring, prediction and analysis of crop pests and diseases in practical applications. The characteristics of monitoring, prediction and analysis of crop pests and diseases require a high degree of attention to detail and strict logical coherence, which is an area where traditional models are insufficient.
[0131] As shown in Table 1, the accuracy of this model is 0.93, the recall rate is 0.90, and the precision rate is 0.91, which is significantly better than the traditional Transformer model (accuracy is 0.80, recall rate is 0.77, and precision rate is 0.79) and the performance of other comparative models; it not only expands the research boundaries of deep learning models in complex system prediction in theory, but also monitors, predicts and analyzes crop diseases and insect pests in complex and changing environments in practical applications, enhancing the adaptability and stability of the model. It meets the needs of practical applications.
[0132] Therefore, the present invention adopts the above-mentioned Transformer agricultural pest and disease intelligent prediction model based on text big data and state space, which can not only significantly improve the efficiency and accuracy of agricultural pest and disease monitoring, but also reduce agricultural production costs and improve the intelligence and automation level of agricultural production. It has important social and economic value. In addition, the promotion and application of this model will help promote the progress of agricultural science and technology and enhance the sustainable development capacity of agricultural production.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A Transformer agricultural pest and disease intelligent prediction model based on text big data and state space, characterized by: The following steps are involved: S1. Collect crop disease image data and annotate the collected image data; S2, preprocessing the labeled data; S3, extract information from the preprocessed data and then construct a knowledge graph; S4, setting the state space transfer equation of the model, the state space based Transformer model and the state space loss function; S5: The data processed by S3 is used as the input of the model. After being processed by the model, the token embedding is output to predict the future state.
2. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: In S1, web crawler technology is used to collect image data, and the data is annotated using the open source annotation tool LabelImg, which includes: First, establish an annotation guide that contains annotation standards and procedures for various data types; Secondly, preliminary annotation is performed, which is independently annotated by several annotators, and several annotators conduct cross-checking and discussion regularly; Finally, reviewers verify and arbitrate the annotators' annotations.
3. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: The data preprocessing in S2 specifically includes: Split, the formula is: W = Segment(T); Where T represents the original data image after S1 processing; W represents the segmentation result; Segment(·) represents the segmentation function; Clean up, the formula is: T′=Clean(T); Where T′ is the cleaned image; Clean(·) is the cleaning operation; Normalized, the formula is: T″=Normalize(T′); Where T″ is the normalized image; Normalize(·) is the normalization function.
4. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: Information extraction in S3 is to identify and extract entities, attributes and relations from preprocessed data. Entity extraction is used to identify named entities in text, relationship extraction determines the semantic connection between entities, and attribute extraction focuses on descriptive information about entities. The extraction of entities and relations is expressed as: Unified Knowledge= Knowledge Fusion(ExtractedInformation); Leveraging TF-IDF and WordVe c To construct the knowledge graph, TF-IDF evaluates the importance of a word by multiplying the term frequency TF of the word by the inverse document frequency IDF. The term frequency TF is the number of times a word appears in a document divided by the total number of words in the document, while the inverse document frequency IDF is the total number of documents divided by the logarithm of the number of documents containing the word. WordVec converts words into vector form through training, captures the contextual relationship between words, and predicts the context through the Skip-gram model and the CBOW model to predict the current word from the context.
5. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: The state space transfer equation in S4 is expressed as: y t =Ch t +Dx t ; Among them, h t represents the hidden state at time t; x t represents the input at time t; y t Represents the output of the model; and ΔB represent the updates of the state transition matrix and the input control matrix, respectively; C and D represent the output matrix and the direct transfer matrix, respectively; Use the exponential map to update the state transition matrix as follows:
6. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: The state-space-based Transformer model in S4 includes: The block consists of several blocks, each of which is a variant of the layer in the Transformer model, and each block contains a self-attention mechanism and a feedforward neural network, and also embeds a state space for the representation layer to simulate the time evolution characteristics of the input data; Inter-block connections, blocks are connected through residual connections and layer normalization.
7. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: The mathematical expression of the state space loss function in S4 is: L(θ)=λ1L predict (θ)+λ2L smooth (i); Among them, L predict (θ) is the traditional prediction loss component; it is expressed as: Where N is the total number of data points; y t is the true value at time t; is the prediction output of the model; L smooth (θ) is the smoothness loss component of the state space model, and the Frobenius norm of the state transition matrix change is used as the smoothing term, as follows: Among them, ΔA represents the change of the state transition matrix; ||·||| F represents the Frobenius norm; λ1 and λ2 represent weighting parameters used to balance the contribution of the prediction loss component and the smoothness loss component to the total loss.
8. The Transformer agricultural pest intelligent prediction model based on text big data and state space according to claim 1 is characterized in that: S5 specifically includes: The processed text data is taken as input and converted into vector representation through the embedding layer. The learned context ID is introduced in each module of the model to identify and learn different context states, and each module contains a linear layer and a cross attention mechanism to integrate information based on the current input and previously learned context states. Finally, the model inputs token embedding for predicting future states, where the token embedding conversion expression is: E = Embed(X); Where X is the input text data; E is the tag embedding; The cross attention mechanism expression is: The state space model integral expression is: S t+1 =F·S t +G·A+H; Among them, S t and S t+1 They represent the system states of the current and next time steps respectively; F, G, H are the state space model parameters; A is the output of the cross attention.