A rose planting medicine scheme decision method based on multi-modal feature fusion

By using multimodal data fusion and a closed-loop decision-making system, the problems of data isolation and static decision-making in the prevention and control of diseases and pests in rose cultivation have been solved, enabling dynamic and precise decision-making on pesticide application plans and improving the accuracy of disease and pest identification and environmental adaptability.

CN121093192BActive Publication Date: 2026-07-21ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2025-08-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, pest and disease monitoring systems in rose cultivation are difficult to adapt to complex planting environments due to their single data dimension and lack of dynamic optimization, resulting in problems such as static pest and disease control decisions, excessive pesticide use, and environmental pollution.

Method used

A decision-making method for pesticide application in rose cultivation based on multimodal feature fusion is constructed. By collecting multimodal data, a pre-trained basic network is used for feature extraction and relearning. The Transformer decoder and regression network are combined to make decisions on the types of pests and diseases and the amount of pesticides to be applied. Closed-loop control is achieved through pesticide mixing control and user interaction optimization.

Benefits of technology

It enables dynamic and precise control of pests and diseases in rose cultivation, improves the accuracy of pest and disease identification and the ability to optimize pesticide application in real time, reduces the risk of excessive pesticide use and environmental pollution, and adapts to the personalized needs of different planting environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093192B_ABST
    Figure CN121093192B_ABST
Patent Text Reader

Abstract

The application discloses a rose planting medicine scheme decision method based on multi-modal feature fusion. First, multi-modal data of a rose planting area is collected, and the obtained multi-modal data is input into a feature fusion module for multi-modal feature fusion. A double-branch decision network is used for disease and pest type identification and medicine type and dosage determination, medicine proportioning rules are introduced through a medicine liquid mixing control module to correct the medicine type and dosage. Finally, through a user interaction optimization mechanism, closed-loop control is achieved, and the medicine scheme decision is optimized. The application realizes dynamic closed-loop decision by combining user feedback to adapt to the individual needs of different growers and solve the medicine application problems caused by data isolation and static decision in traditional rose planting. The application deeply couples artificial intelligence technology and agronomic knowledge to promote the transformation of traditional agriculture to intelligence and automation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of crop pest and disease control decision-making, specifically involving a decision-making method for pesticide application schemes in rose cultivation based on multimodal feature fusion. Background Technology

[0002] Roses are a core economic crop for ornamental, medicinal, and essential oil extraction purposes. Diseases such as black spot and powdery mildew, as well as pests such as spider mites and aphids, can reduce yields by 30% to 50%. Global warming increases the number of pest generations by 1 to 2 per year, and cross-border trade exacerbates the risk of invasion by alien species such as western flower thrips. Manual inspection relies on experience, and the rate of missed detection of early lesions (such as chlorophyll degradation and larval activity) exceeds 40%. Misdiagnosis leads to delayed prevention and control, excessive use of pesticides, waste of water resources, and environmental pollution, all of which increase planting costs.

[0003] Existing technologies, such as single-sensor monitoring (e.g., soil moisture sensors) or image-based pest and disease recognition systems, suffer from limited data dimensions and a lack of learning capabilities from historical data. These limitations make them ill-suited to complex planting environments and unable to dynamically adjust strategies based on crop growth stages. Furthermore, traditional methods for managing rose pests and diseases lack real-time optimization of application plans and interactive feedback correction mechanisms, hindering closed-loop control. Therefore, this paper designs a method for calculating pesticide dosage in rose cultivation based on multimodal feature fusion. By integrating image, sensor, meteorological, and crop information data, and combining relearning methods to address modal imbalance and low computational efficiency, this method adds a pesticide mixing correction mechanism and a user-interactive optimization mechanism. This enables dynamic and precise pest and disease control and closed-loop pesticide application decisions, resolving the application problems caused by isolated data and static decision-making in traditional rose cultivation.

[0004] Patent CN202211654521.X proposes a pest and disease identification, prevention, and early warning system, including an image recognition module, a knowledge graph module, a prediction and early warning module, and an alarm notification module, which facilitates pest and disease control. However, this patent has a single data dimension, resulting in static decision-making, a lack of dynamic optimization capabilities, and deficiencies such as the absence of a decision-making layer and an execution layer, making closed-loop control impossible.

[0005] Patent CN202510811531.7 proposes a real-time identification and control decision-making method and system for crop pests and diseases based on multimodal edge computing, including multimodal data fusion, ant colony path optimization, and federated learning updates, which improves the accuracy of pest and disease identification and reduces pesticide use. However, the decision flow of this patent is a unidirectional open loop, which lacks real-time optimization and interactive feedback correction.

[0006] Therefore, it is urgent to build a multimodal intelligent system adapted to rose cultivation, break through the bottleneck of single-modal perception, realize cross-modal deep integration, real-time edge deployment and closed-loop decision-making, and solve the problem of pesticide application decision-making for pest and disease control. Summary of the Invention

[0007] The purpose of this invention is to provide a decision-making method for pesticide application in rose cultivation based on multimodal feature fusion. This method addresses the technical shortcomings of existing technologies, such as the difficulty in adapting to complex planting environments due to the limited data dimensions and lack of dynamic optimization in single-sensor monitoring or image recognition pest and disease systems. The invention constructs a dynamic closed-loop decision-making system.

[0008] The objective of this invention is achieved through the following technical solution:

[0009] In a first aspect, embodiments of this application provide a method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion, comprising the following steps:

[0010] Step 1: Collect multimodal data from the rose planting area.

[0011] Step 2: Input the obtained multimodal data into the feature fusion module for multimodal feature fusion.

[0012] Feature extraction is performed using different pre-trained base networks, i.e. encoders; a relearning method is used to detect whether the features of each modality are overfitted or underfitted, and the encoder parameters are updated and optimized; the features output by the updated encoders are concatenated and input into the cross-modal interaction layer for feature fusion to obtain multimodal fused features.

[0013] Step 3: Medication regimen decision-making and drug solution mixing control.

[0014] A two-branch decision network is used for pest and disease identification and for determining the type and dosage of pesticides. The two-branch decision network includes a Transformer decoder, a pesticide type decision engine, and a regression network. The Transformer decoder decodes the multimodal fusion features and outputs a pest and disease type label sequence. The pesticide type decision engine then outputs a preliminary pesticide type decision. Finally, the regression network determines the pesticide dosage.

[0015] The drug mixing control module introduces drug ratio rules to modify the types and dosages of drugs.

[0016] Step 4: Through user interaction optimization mechanisms, achieve closed-loop control and optimize medication plan decisions.

[0017] In one possible implementation, the specific operation of collecting multimodal data from the rose planting area is as follows:

[0018] Image data acquisition: Images including plant height, leaf morphology, and pest and disease patch characteristics are acquired using a visible light camera, while images including plant temperature distribution are acquired using an infrared camera. Time-series sensor data acquisition: Real-time environmental parameters are acquired using soil moisture sensors, ambient temperature and humidity sensors, and light sensors. Meteorological data acquisition: Rainfall and wind speed are acquired via a weather station API. Crop information acquisition: A database of rose plant variety disease resistance and phenological stage markers (budding, flowering, and dormancy stages) are acquired. Simultaneously, multimodal data from historical planting records and data on the actual types and amounts of pesticides applied at that time (i.e., historical data) are collected. The acquired multimodal dataset (containing historical and real-time data) is divided into training, validation, and test sets according to the corresponding modes. Historical data is used for encoder pre-training, relearning, and regression network training, as well as for constructing a rose plant pesticide knowledge base. Real-time data is used for real-time pest and disease identification and pesticide application decisions.

[0019] In one possible implementation, feature extraction using an encoder is specifically implemented as follows: For image data, a CNN is used to extract the spatial features of pests and diseases from infrared and visible light images. Specifically, for infrared images, preprocessing is required. After temperature matrix standardization, temperature threshold segmentation, and pseudo-color encoding, a single-channel pseudo-color image is obtained, and then a ResNet50 network structure is used for feature extraction. For meteorological data and time-series sensor data, LSTM is used for feature extraction, generating time-encoded vectors corresponding to each type of data. For crop information, a Large Language Model (LLM) based on cue word engineering is used to extract textual feature vectors of crop information. After quantifying the disease resistance of rose plants under different conditions through preset cue word template constraints, structured data is output.

[0020] In one possible implementation, the relearning method includes: first, diagnosing the learning state, i.e., extracting features for each modality, obtaining the purity of the dataset using the k-means clustering algorithm, i.e., the proportion of samples with the same label within the cluster of modality k dataset, and calculating the purity difference between the training set and the validation set respectively. Then, performing soft reinitialization, calculating the reinitialization intensity based on the purity difference, updating the encoder parameters of the feature extraction part, and using the updated encoder for feature extraction of real-time data.

[0021] In one possible implementation, the concatenated features are input into a cross-modal interaction layer, and feature fusion is performed based on a cross-attention mechanism. This mechanism includes: establishing a three-modal collaborative mechanism using a multi-head cross-attention layer, using image feature vectors as attention query benchmarks, constructing dynamic environment association keys using time-encoded vectors, and generating semantic constraint values ​​using text feature vectors, and dynamically adjusting feature association weights through three sets of trainable parameters.

[0022] In one possible implementation, the Transformer decoder operates as follows:

[0023] (1) Decoder input construction:

[0024] The multimodal fusion features are used as the initial hidden state H. enc ∈R d×n Add a sequence start character <sos>and position code E pos .

[0025] (2) Iterative tag generation: The decoder consists of an L-layer structure, and each layer performs the following:

[0026] Self-attention: This associates the current tag with historically generated tags to ensure semantic coherence. In the first round of generation, the historical tags are empty, and attention is only focused on... <sos>The attention computation is initialized with positional encoding, and new labels are appended to the historical sequence in each iteration for use in the next round of self-attention computation.

[0027] Encode-decode attention: Using fused features as key-value pairs and the vector embedded by the current label as query, it focuses on key regions of pests and diseases in the features.

[0028] Feedforward prediction: Outputs the probability distribution P of the next label. t ∈R V Candidate tags are selected using beam search.

[0029] (3) Category mapping and output:

[0030] When the decoder predicts <eos>Generation terminates when the probability of a label exceeds a threshold or the sequence reaches its maximum length (e.g., the number of disease label types exceeds a threshold). A rose-specific decision dictionary is used to map the labels to pest types: each label corresponds to a specific type, and the pest type with the highest confidence and its corresponding probability are output.

[0031] (4) Downstream interface:

[0032] The types of pests and diseases, time-encoded vectors, and text feature vectors are all input into the pesticide application decision engine.

[0033] In one possible implementation, the medication type decision engine is specifically as follows:

[0034] The input sources for the decision engine include: the probability distribution of disease or pest categories output by the Transformer decoder, the time encoding vector extracted by LSTM, the predefined crop phenological weighting coefficients in the regression network, and the text feature vector generated by LLM.

[0035] The core decision-making rules include: establishing a prohibition rule base: combining information on rose phenology and plant variety disease resistance and drug resistance to prohibit corresponding unsuitable agents; environmental adaptation mechanism: prohibiting agents that are unsuitable for use under special weather conditions; specificity mechanism: different types of agents are used for different types and severity of pests and diseases, as well as for plants in different phenological stages.

[0036] The medication type decision engine adopts a three-level hierarchical architecture, specifically including:

[0037] (1) Input preprocessing layer: Standardize the data format of each input source:

[0038] The probability distribution of pests and diseases is normalized to generate a confidence vector P∈[0,1]. n .

[0039] Temporal encoded vectors are fused with sensor data through linear projection to generate environmental feature embeddings E. t ∈R de .

[0040] Phenological sensitivity coefficient W is generated by mapping the basic coefficients. p .

[0041] (2) Rule execution layer: Based on the pest and disease confidence vector P and environmental features E obtained after input preprocessing. t and stage sensitivity coefficient W p The decision rules are executed according to the priority of the prohibition rule base - environmental adaptation mechanism - specific mechanism, and the decision on the type of drug to be applied is output. The specific steps include:

[0042] The stage sensitivity coefficient Wp is used for decision-making in the ban rule base section, reflecting information on rose phenology and plant variety disease resistance and drug resistance, and banning corresponding unsuitable drugs based on the knowledge base.

[0043] Environmental characteristics Et are used for decision-making in the environmental adaptation mechanism. They reflect environmental information including wind speed, light intensity, temperature, and humidity to determine whether the environment is suitable for drug use and to disable drugs that are not suitable for use under special weather conditions based on the knowledge base.

[0044] The pest and disease confidence vector P is used for decision-making in the specific mechanism part, reflecting the type and severity of pests and diseases, as well as the phenological stage of the rose plant. For rose plants suffering from different pests and diseases and in different phenological stages, different types of pesticides are output based on the knowledge base.

[0045] The rule enforcement layer outputs preliminary drug type decisions.

[0046] (3) Correlation Matrix Constraint Layer: Add an attention scoring step to score the pesticides output by the rule execution layer. Add the disease-pesticide correlation matrix M∈{0,1} during scoring. d×d To prevent agronomical errors, mandatory restrictions should be placed on the types of drugs used.

[0047] Where M ij =0 indicates that disease i and agent j are incompatible.

[0048] Through the above three-tiered architecture, the decision engine outputs the final drug type decision.

[0049] In one possible implementation, the regression network uses a fully connected neural network to predict the dose, and its specific calculation process is as follows:

[0050] (1) Input feature standardization: multimodal fusion features F in cross-modal interaction layers fusion Batch standardization is performed using the following formula:

[0051] Where, μ B , σ B γ represents the statistics for the current batch, and β are learnable parameters.

[0052] (2) Nonlinear feature mapping: Dose-related features are extracted through a two-layer fully connected network.

[0053] First layer: H1=ReLU(W1) + b1)

[0054] Second layer: H2 = Dropout(ReLU(W2H1+b2))

[0055] Where H is the feature tensor, W is the weight matrix, and b is the output layer bias term.

[0056] (3) Phenological stage feature modulation: Dynamically adjusting feature expression according to crop growth stages:

[0057] H p =H2⊙(1+αW p )

[0058] Where H2 is the tensor of the output features of the second hidden layer, ⊙ denotes element-wise multiplication, α is the modulation intensity coefficient, and W p It is a phenological option weight vector.

[0059] (4) Dosage prediction output: The final predicted value is restored to the actual dose after being constrained by Sigmoid:

[0060] y real =Q max ⋅σ(W3H p +b3)

[0061] Where W3 is the output layer weight matrix, b3 is the output layer bias term, and Q... max This is the preset maximum safe dosage of the drug.

[0062] In one possible implementation, the drug solution mixing control module operates as follows:

[0063] The specific rules for drug formulation are as follows:

[0064] First, construct a four-dimensional matching matrix R∈R K×M×P×E Where K represents the type of pesticide, M represents the phenological stage of roses, P represents the intensity of pests and diseases, and E represents environmental factors such as temperature and humidity.

[0065] Mixing ratio = R[k,m,p,e] × ×(1+αΔT)

[0066] Among them, C 标准 The standard concentration is given by ΔT, which is the temperature difference between the reagent and the environment to compensate for the error caused by the thermal expansion and contraction of the reagent volume. α is the temperature compensation coefficient.

[0067] For medication type adjustments, whether a substitution mechanism is triggered is determined based on whether the mixing ratio exceeds a threshold; for dosage adjustments, a baseline coefficient for concentration adjustment is calculated based on the mixing ratio. Details are as follows:

[0068] The alternative drug matching mechanism is set up to modify the types of drugs used: First, the alternative drug schemes are dynamically generated based on the drug mechanism database (MOA database). Then, resistance accumulation detection is carried out and the historical drug use of the same plot is recorded. If it is detected that the drug with the same mechanism of action has been used repeatedly, the drug type is changed to avoid the accumulation of resistance.

[0069] Based on real-time collected environmental factors, the system determines whether there are abnormal environmental conditions and adjusts the dosage when abnormal conditions are detected. Specifically, this includes stopping medication if the wind speed exceeds a threshold and is unsuitable for medication; reducing the dosage if the temperature or light intensity exceeds a threshold; and adding an anti-drift agent to the medication if the humidity exceeds a threshold.

[0070] In one possible implementation, the user interaction optimization mechanism is as follows:

[0071] Users monitor the effects of medication over a certain period and can send correction commands to the system's recommended dosage via a mobile app.

[0072] The reinforcement learning reward function parameters are updated based on the corrected data. The specific update formula is as follows:

[0073]

[0074] in Set the dosage for the user. λ represents the system's recommended dosage, and λ is the learning rate.

[0075] In one possible implementation, the encoder is pre-trained using historical data, with focus loss used for CNN, mean absolute error loss used for LSTM, and label smoothing cross-entropy loss used for LLM, to reduce overfitting during feature extraction.

[0076] In one possible implementation, after the collected multimodal historical data is processed by an encoder updated through relearning, features are extracted, and then feature concatenation and fusion are performed. The fused data is then input into a regression network for training.

[0077] The loss function for the regression network is the phenological period constraint mean squared error (Phenology-MSE), which is related to the phenological period weighting coefficient, as the loss function.

[0078]

[0079] Where N is the batch sample size, ω pheno The phenological weighting coefficient is determined based on the phenological stage of the rose. pred To predict drug dosage using the model, y true This represents the actual dosage used.

[0080] Secondly, embodiments of this application provide a decision-making device for rose planting pesticide application schemes based on multimodal feature fusion, which consists of a multimodal data acquisition unit, a feature fusion module, a pesticide application scheme decision-making module, and a user interaction feedback system.

[0081] The multimodal data acquisition unit is used to collect multimodal data from the rose planting area, including image data, sensor data, meteorological data, and crop information.

[0082] The feature fusion module is used for feature extraction, feature concatenation, and feature fusion of multimodal data. Different base networks, i.e., encoders, are used for feature extraction; a relearning method is employed to detect overfitting or underfitting of each modality's features, and the encoder parameters are updated and optimized accordingly; the updated encoder output features are concatenated and input into a cross-modal interaction layer for feature fusion to obtain multimodal fused features.

[0083] The medication application decision module utilizes a dual-branch decision network to identify pests and diseases and determine the type and dosage of medication. This dual-branch decision network includes a Transformer decoder, a medication type decision engine, and a regression network. The Transformer decoder decodes the multimodal fusion features, outputting a pest and disease type label sequence. The medication type decision engine then outputs a preliminary medication type decision. The regression network determines the dosage. Finally, a medication mixing control module introduces medication ratio rules to refine the medication type and dosage.

[0084] The user interaction feedback system is used to receive user correction instructions for the system's recommended medication dosage via a mobile APP after the user observes the effect of the implementation of the plan, thereby achieving closed-loop control and optimizing medication plan decisions.

[0085] The beneficial effects of this invention are as follows:

[0086] 1. Deeply integrate multimodal data: Deeply integrate multimodal data such as image data, sensor data, meteorological data, and crop information to solve the problem of correlation between local image features and global environmental parameters, improve the accuracy of pest and disease identification, and then output preliminary pesticide application decisions through the pesticide application type decision engine.

[0087] 2. User Feedback and Dynamic Decision-Making: Regression networks are used to predict pesticide dosage and make corresponding decisions, outputting the optimal dosage. This is then combined with user feedback to achieve dynamic closed-loop decision-making, adapting to the personalized needs of different growers and solving the pesticide application problems caused by isolated data and static decision-making in traditional rose cultivation. This deeply couples artificial intelligence technology with agronomic knowledge, driving the transformation of traditional agriculture towards intelligence and automation.

[0088] 3. Improve data utilization: Make full use of the acquired multimodal data, extract key plant features, determine plant status and pesticide dosage, and improve the utilization of acquired data. Attached Figure Description

[0089] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0090] Figure 1 This is a flowchart illustrating the technical solution of an embodiment of the present invention.

[0091] Figure 2 This is a flowchart illustrating the use of a feature fusion module to achieve multimodal data feature fusion in an embodiment of the present invention.

[0092] Figure 3 This is a flowchart illustrating the use of a bi-branch decision network for medication regimen decision-making in an embodiment of the present invention. Detailed Implementation

[0093] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0094] The objective of this invention is achieved through the following technical solutions: collecting multimodal data from rose-growing areas, multimodal feature fusion using a feature fusion module, decision-making for pesticide application and control of pesticide mixing via a two-branch decision network, and closed-loop control through a user interaction optimization mechanism, as shown in the appendix. Figure 1 As shown.

[0095] In one possible implementation, the specific operation of collecting multimodal data from the rose planting area is as follows:

[0096] Image data acquisition: Images including plant height, leaf morphology, and pest and disease patch characteristics are acquired using a visible light camera, while images including plant temperature distribution are acquired using an infrared camera. Time-series sensor data acquisition: Real-time environmental parameters are acquired using soil moisture sensors, ambient temperature and humidity sensors, and light sensors. Meteorological data acquisition: Rainfall and wind speed are acquired via a weather station API. Crop information acquisition: A database of rose plant variety disease resistance and phenological stage markers (budding, flowering, and dormancy stages) are acquired. Simultaneously, multimodal data from historical planting records and data on the actual types and amounts of pesticides applied at that time (i.e., historical data) are collected. The acquired multimodal dataset (containing historical and real-time data) is divided into training, validation, and test sets according to the corresponding modes. Historical data is used for encoder pre-training, relearning, and regression network training, as well as for constructing a rose plant pesticide knowledge base. Real-time data is used for real-time pest and disease identification and pesticide application decisions.

[0097] In one possible implementation, the obtained multimodal data is input into a feature fusion module for multimodal feature fusion, such as... Figure 2 As shown.

[0098] Feature extraction is performed using different pre-trained base networks, i.e., encoders: for image data, CNN is used to extract spatial features of pests and diseases from infrared and visible light images; for meteorological and time-series data, LSTM is used for feature extraction to generate time-encoded vectors; for crop information, LLM is used for feature extraction, and structured data is output by constraining the output with preset prompt word templates. The relearning method first diagnoses the learning state by calculating the purity difference between the training and validation sets for each feature extracted from each modality. Then, soft re-initialization is performed, calculating the re-initialization intensity based on the purity difference and updating the encoder parameters. Finally, feature updates are performed, outputting the re-initialized feature vector. The feature concatenation method concatenates the image features, time-series features, and text features obtained after relearning initialization into a single feature vector. The concatenated feature is then input into a cross-modal interaction layer for feature fusion based on a cross-attention mechanism.

[0099] In one possible implementation, a CNN (Convolutional Neural Network) is used to extract spatial features of pests and diseases from infrared and RGB images. For example, an infrared image might show an area of ​​rose bush leaves with an abnormally higher temperature than the surrounding area, or an RGB image might show an area covered with white powder or radial black spots. For the infrared image, preprocessing is required. This involves temperature matrix normalization, temperature threshold segmentation, and pseudo-color encoding to obtain a single-channel pseudo-color image. Then, a ResNet50 network structure is used for feature extraction.

[0100] Input data specifications:

[0101] Data type: RGB three-channel image data.

[0102] Input dimensions: 224×224×3 (standard ResNet50 input size).

[0103] Preprocessing flow:

[0104] Normalization: Pixel values ​​are scaled to [0,1].

[0105] Data enhancements: rotation (±30°), color dithering (brightness / contrast adjustment ±20%).

[0106] ResNet50 core architecture:

[0107] Residual module formula:

[0108] Output = F(x, {W_i}) + x

[0109] Where F(x) is the convolutional layer mapping and x is the skip input.

[0110] Hierarchical configuration (key components):

[0111] stage Module type Kernel size / stride Output Dimension Conv1 7×7 convolution 7×7, stride=2 112×112×64 Stage 2 Residual block ×3 [1×1,64]→[3×3,64]→[1×1,256] 56×56×256 Stage 3 Residual blocks ×4 [1×1,128]→[3×3,128]→[1×1,512] 28×28×512 Stage 4 6 residual blocks [1×1,256]→[3×3,256]→[1×1,1024] 14×14×1024 Stage 5 Residual block ×3 [1×1,512]→[3×3,512]→[1×1,2048] 7×7×2048

[0112] Each residual block contains 3 convolutional layers, and gradient vanishing is solved by skipping connections.

[0113] Attention Enhancement (CBAM module):

[0114] An attention enhancement module is introduced in Stages 2-5 to strengthen the response in pest and disease areas.

[0115] Feature output:

[0116] Final layer: Global Average Pooling (GAP), 7×7×2048 → 2048-dimensional vector.

[0117] This implementation method processes image data by using shallow convolutional kernels to capture temperature anomaly edges in infrared images, while deep networks can identify the topological relationship between lesions and healthy areas. Using a CNN model trained on infrared and RGB multimodal data, the overall recognition accuracy exceeds 90%. The technical advantage of this implementation method lies in automatically learning the joint distribution of insect morphology and temperature characteristics through CNN, overcoming the problem of low sensitivity in detecting small pests using traditional methods.

[0118] In one possible implementation, LSTM is used to process meteorological data and time-series sensor data to generate time-encoded vectors corresponding to each type of data. Specifically:

[0119] The input consists of meteorological data including rainfall and wind speed, and sensor data including soil moisture, air temperature and humidity, and light intensity. Different meteorological and planting environment changes correspond to different diseases that are prone to occur, such as continuous rainfall leading to gray mold and excessive humidity leading to downy mildew. LSTM is used to process various types of data in the meteorological data and time series data respectively, and corresponding time encoding vectors are generated.

[0120] Input data specifications:

[0121] Data type: Time series meteorological data (temperature, humidity, wind speed, etc.).

[0122] Input dimensions: (n_timesteps, n_features).

[0123] For example, if you use data from the previous 24 hours (at hourly intervals) to predict precipitation in the next 3 hours, then n_timesteps=24, n_features=5 (temperature, dew point, humidity, air pressure, wind speed).

[0124] Feature engineering:

[0125] Lag characteristics: x(t-1), x(t-2), x(t-3).

[0126] Rolling statistics: 3-hour mean / standard deviation, 7-hour mean / standard deviation.

[0127] Time-coded vector generation:

[0128] Output dimension: The hidden state h_t of the last layer of LSTM (dimension=2048) is used as the temporal encoding vector.

[0129] Network configuration (taking weather forecasting as an example):

[0130]

[0131] The Dropout layer prevents overfitting, and the last layer outputs the predicted value linearly.

[0132] Time-coded vector generation:

[0133] Output dimension: The hidden state h_t of the last layer of LSTM (dimension=2048) is used as the temporal encoding vector.

[0134] Integration Application: Concatenates with CNN feature vectors (2048 dimensions) and inputs into a fully connected layer for joint prediction.

[0135] Data processing using this method can handle complex time-series patterns that are nonlinear, non-stationary, and multi-scale. It achieves technical effects such as long-term dependency modeling and dynamic adaptation in the feature extraction of meteorological and time-series data. As a joint prediction method for environmental factor variables, it can be input into medication decision-making to enhance the effectiveness of medication.

[0136] In one possible implementation, a large language model (LLM) based on cue word engineering is used to extract textual feature vectors of crop information:

[0137] Input a database of rose variety disease resistance and descriptive text representing phenological stages, such as a rose variety exhibiting high resistance to powdery mildew but low resistance to black spot, or different disease resistance at different phenological stages (budding, flowering), etc. Using preset prompt templates, quantify the disease resistance of rose plants under different conditions, and then output structured data.

[0138] Example of a prompt word template:

[0139] Extract the disease resistance description for {variety name}:

[0140] [Disease Type]: Powdery mildew [Resistance Level]: High resistance;

[0141] [Disease Type]: Gray mold [Resistance Level]: Moderately resistant

[0142] Loss function optimization: Adopting label-smoothed cross-entropy loss to reduce overfitting.

[0143]

[0144] Where yc is the true label, pc is the predicted probability, and ϵ is the smoothing coefficient (set to 0.1). The final output dimension is also a 2048-dimensional vector.

[0145] This implementation method utilizes Large Language Modeling (LLM) to extract textual feature vectors of crop information. Preset prompts support complex decision-making, simulate expert reasoning processes, and handle multi-step decision-making problems. When applied to a user interaction feedback system, the preset prompts guide and optimize the user interaction feedback mechanism. Based on the current distribution of pests and diseases, it helps users understand the reasoning process for pesticide dosage decisions, facilitates observation of pesticide effects, and allows for comparison between the implementation plan and the results.

[0146] In one possible implementation, a relearning method is used to detect whether each modal feature is overfitted or underfitted, and the encoder parameters are updated and optimized accordingly.

[0147] After feature extraction from historical data, it is input into the relearning module. The relearning method includes: first, diagnosing the learning state: for each modality, extracting features, obtaining the purity of the dataset using the k-means clustering algorithm, i.e., the proportion of samples with the same label within the cluster of modality k dataset, and calculating the purity difference g between the training set and the validation set (these two datasets are sourced from: historical rose planting data (including multimodal samples) collected from multiple growing seasons, labeled by experts according to agricultural standards as the pest and disease types, used as the training set, and then a certain amount of data randomly sampled proportionally from the training set as the validation set). k :

[0148] g k =∣P D k - P V k |

[0149] Among them, P D k P represents the purity of the training set for mode k. V k Let g be the purity of the validation set for mode k. k The larger the value, the worse the generalization ability of the feature extractor corresponding to mode k.

[0150] Then, soft re-initialization is performed: the re-initialization intensity is calculated based on the purity difference, and the encoder parameters of the feature extraction part are updated (the encoder is either CNN, LSTM, or LLM; if overfitting occurs in the data of a certain modality, i.e., a large purity difference, the re-initialization intensity of its encoder is increased). The formula for updating the parameters is:

[0151] θ new =(1−α k )⋅θ current +α k ⋅θ init

[0152] Where, θ init Reset the intensity α for the initial training weights of the encoder. k =tanh(λ⋅g k ), positively correlated with purity, combined with the updated encoder parameters θ new The updated encoder is obtained and used for feature extraction from real-time data.

[0153] This implementation method achieves multimodal data fusion through relearning, enabling cross-modal feature alignment and improving the accuracy of pest and disease diagnosis; enhancing feature robustness to cope with the complexity and noise of rose pest and disease data; and enabling incremental learning capabilities to support dynamic updates of pest and disease models.

[0154] In one possible implementation, the feature concatenation is as follows:

[0155] The text feature vector extracted by the updated encoder Image feature vectors With time-coded vector The concatenation is performed, M=[T,I,S], with dimensions (3,b).

[0156] In one possible implementation, cross-modal feature fusion is as follows:

[0157] The concatenated features are input into the cross-modal interaction layer, and feature fusion is performed based on the cross-attention mechanism. This mechanism includes: establishing a three-modal collaborative mechanism using a multi-head cross-attention layer, using image feature vectors as attention query benchmarks, constructing dynamic environment association keys using time-encoded vectors, and generating semantic constraint values ​​using text feature vectors, and dynamically adjusting feature association weights through three sets of trainable parameters.

[0158] In one possible implementation, such as Figure 3 As shown, the specific details of medication regimen decision-making and drug mixing control using a two-branch decision network are as follows:

[0159] (1) Identification of pest and disease types and decision-making on pesticide types:

[0160] The Transformer decoder is used to decode the multimodal fusion features and output the pest and disease type label sequence. The specific steps are as follows:

[0161] 1. Decoder input construction:

[0162] The multimodal fusion features output from the cross-modal interaction layer are used as the initial hidden state H. enc ∈R d×n Add a sequence start character <sos>and position code E pos .

[0163] 2. Iterative Tag Generation: The decoder consists of an L-layer structure, with each layer performing the following:

[0164] Self-attention: Associates the current tag with historically generated tags to ensure semantic coherence (where, in the first round of generation, the historical tags are empty, and only through...). <sos>Attention computation is initialized with positional encoding, and new labels are appended to the history sequence in each iteration for use in the next round of self-attention computation.

[0165] Encoder-decoder attention: using fused features as key-value pairs and the vector embedded by the current label as the query, it focuses on key regions of pests and diseases in the features;

[0166] Feedforward prediction: Outputs the probability distribution P of the next label. t ∈R V Candidate tags are selected using beam search.

[0167] 3. Category Mapping and Output:

[0168] When the decoder predicts <eos>Generation terminates when the probability of a label exceeds a threshold or the sequence reaches its maximum length (e.g., the number of disease label types exceeds a threshold). A rose-specific decision dictionary is used to map the labels to pest types: each label corresponds to a specific type, and the pest type with the highest confidence and its corresponding probability are output.

[0169] Basic tags Drug type [Azoxystrobin] [Mancozeb] Composite Marker Dosage range [0.2L / mu] [0.5L / mu] Environmental constraint markers Phenological Period / Meteorological Constraints {Flowering period disabled} {Rainy weather delay} Termination mark End of sequence

[0170] 4. Downstream interface:

[0171] The pesticide application decision engine is jointly inputted with the pest and disease type and environmental characteristics (LSTM temporal coding) and the crop disease resistance textual characteristics (LLM coding).

[0172] The specific drug type decision engine is as follows:

[0173] The input sources for the decision engine include: the probability distribution of disease or pest categories output by the Transformer decoder, the time encoding vector extracted by LSTM, the predefined crop phenological weighting coefficients in the regression network, and the text feature vector generated by LLM.

[0174] The core decision-making rules include: establishing a prohibition rule base: combining information on rose phenology and plant variety disease resistance and drug resistance to prohibit corresponding unsuitable agents; environmental adaptation mechanism: prohibiting agents that are unsuitable for use under special weather conditions; specificity mechanism: different types of agents are used for different types and severity of pests and diseases, as well as for plants in different phenological stages.

[0175] The medication type decision engine adopts a three-level hierarchical architecture, specifically including:

[0176] 1. Input preprocessing layer: Standardizes the data format of each input source:

[0177] The probability distribution of pests and diseases is normalized to generate a confidence vector P∈[0,1]. n .

[0178] Temporal encoded vectors are fused with sensor data through linear projection to generate environmental feature embeddings E. t ∈R de .

[0179] Phenological sensitivity coefficient W is generated by mapping the basic coefficients. p .

[0180] 2. Rule Execution Layer: Based on the pest and disease confidence vector P and environmental features E obtained after preprocessing the input. t and stage sensitivity coefficient W p The decision rules are executed according to the priority of the prohibition rule base - environmental adaptation mechanism - specific mechanism, and the decision on the type of drug to be applied is output. The specific steps include:

[0181] The stage sensitivity coefficient Wp is used for decision-making in the ban rule base section, reflecting information on rose phenology and plant variety disease resistance and drug resistance, and banning corresponding unsuitable drugs based on the knowledge base.

[0182] The environmental characteristic Et is used for decision-making in the environmental adaptation mechanism. It reflects environmental information including wind speed, light intensity, temperature, and humidity to determine whether the environment is suitable for drug use and to disable drugs that are not suitable for use under special weather conditions based on the knowledge base. For example, powdered drugs are disabled when the wind speed is greater than the threshold, and easily photodegradable drugs are disabled when the light intensity is greater than the threshold.

[0183] The pest and disease confidence vector P is used for decision-making in the specific mechanism part, reflecting the type and severity of pests and diseases, as well as the phenological stage of the rose plant. For rose plants suffering from different pests and diseases and in different phenological stages, different types of pesticides are output based on the knowledge base.

[0184] The rule enforcement layer outputs preliminary drug type decisions.

[0185] 3. Association Matrix Constraint Layer: Add an attention scoring step to score the pesticides output by the rule execution layer. During scoring, add the disease-pesticide association matrix M∈{0,1}. d×d To prevent agronomical errors, mandatory restrictions should be placed on the types of drugs used.

[0186] Where M ij =0 indicates that disease i and agent j are incompatible.

[0187] Through the above three-tiered architecture, the decision engine outputs the final drug type decision.

[0188] (2) Using regression networks for medication dosage decisions:

[0189] The regression network uses a fully connected neural network to predict dose, and its specific calculation process is as follows:

[0190] (1) Input feature standardization: multimodal fusion features F in cross-modal interaction layers fusion Batch standardization is performed using the following formula:

[0191] Where, μ B , σ B γ represents the statistics for the current batch, and β are learnable parameters.

[0192] (2) Nonlinear feature mapping: Dose-related features are extracted through a two-layer fully connected network.

[0193] First layer: H1=ReLU(W1) + b1)

[0194] Second layer: H2 = Dropout(ReLU(W2H1+b2))

[0195] Where H is the feature tensor, W is the weight matrix, and b is the output layer bias term.

[0196] (3) Phenological stage feature modulation: Dynamically adjusting feature expression according to crop growth stages:

[0197] H p =H2⊙(1+αW p )

[0198] Where H2 is the tensor of the output features of the second hidden layer, ⊙ denotes element-wise multiplication, α is the modulation intensity coefficient, and W p It is a phenological option weight vector.

[0199] (4) Dosage prediction output: The final predicted value is restored to the actual dose after being constrained by Sigmoid:

[0200] y real =Q max ⋅σ(W3H p +b3)

[0201] Where W3 is the output layer weight matrix, b3 is the output layer bias term, and Q... max This is the preset maximum safe dosage of the drug.

[0202] (3) The drug ratio rules are introduced through the drug solution mixing control module to modify the types and amounts of drugs used.

[0203] The specific rules for drug formulation are as follows:

[0204] First, construct a four-dimensional matching matrix R∈R K×M×P×E Where K represents the type of pesticide, M represents the phenological stage of roses, P represents the intensity of pests and diseases, and E represents environmental factors such as temperature and humidity.

[0205] Mixing ratio = R[k,m,p,e] × ×(1+αΔT)

[0206] Among them, C 标准 The standard concentration is given by ΔT, which is the temperature difference between the reagent and the environment to compensate for the error caused by the thermal expansion and contraction of the reagent volume. α is the temperature compensation coefficient.

[0207] For medication type adjustments, whether a substitution mechanism is triggered is determined based on whether the mixing ratio exceeds a threshold; for dosage adjustments, a baseline coefficient for concentration adjustment is calculated based on the mixing ratio. Details are as follows:

[0208] The alternative drug matching mechanism is set up to modify the types of drugs used: First, the alternative drug schemes are dynamically generated based on the drug mechanism database (MOA database). Then, resistance accumulation detection is carried out and the historical drug use of the same plot is recorded. If it is detected that the drug with the same mechanism of action has been used repeatedly, the drug type is changed to avoid the accumulation of resistance.

[0209] Based on real-time collected environmental factors (including wind speed, temperature, humidity, and light intensity), the system determines whether there are abnormal environmental conditions and adjusts the dosage when abnormal conditions are detected. Specifically, this includes: stopping medication if wind speed > 0.5 m / s is unsuitable; reducing the dosage to 80% when temperature > 30℃; reducing the dosage to 70% when light intensity > 80 klux; and adding an anti-drift agent to the medication when humidity > 75% RH.

[0210] By using environmental factors as a condition for adjusting pesticide dosage decisions in this implementation method, the effectiveness of pesticide application can be significantly enhanced, thus improving pest control. For example, accurately identifying the effects of temperature, light, and humidity can reduce pesticide dosage, decrease pesticide residues, and save on pesticide costs; at the same time, it enhances the sustainability of the agricultural ecosystem.

[0211] In one possible implementation, a user interaction optimization mechanism is used to achieve closed-loop control and optimize medication regimen decisions.

[0212] Users monitor the effects of medication over a certain period and can send correction commands to the system's recommended dosage via a mobile app.

[0213] The reinforcement learning reward function parameters are updated based on the corrected data. The specific update formula is as follows:

[0214]

[0215] in Set the dosage for the user. λ represents the system's recommended dosage, and λ is the learning rate.

[0216] In one possible implementation, during encoder training, focus loss is used for CNN, mean absolute error loss is used for LSTM, and label smoothing cross-entropy loss is used for LLM to reduce overfitting during feature extraction.

[0217] In one possible implementation, features are extracted from the collected multimodal historical data by an encoder updated through relearning. These features are then concatenated and fused, and the fused data is input into a regression network for training.

[0218] The loss function for the regression network is the phenological period constraint mean squared error (Phenology-MSE), which is related to the phenological period weighting coefficient, as the loss function.

[0219]

[0220] Where N is the batch sample size, ω pheno The phenological weighting coefficient is determined based on the phenological stage of the rose (the budding stage has the highest weighting coefficient (2.0 in this example), followed by the bud stage (1.2 in this example), and the dormant stage has the lowest weighting coefficient (0.5 in this example)). pred To predict drug dosage using the model, y true This represents the actual dosage used.

[0221] This application also provides a decision-making device for rose planting pesticide application based on multimodal feature fusion, which consists of a multimodal data acquisition unit, a feature fusion module, a pesticide application decision-making module, and a user interaction feedback system.

[0222] The multimodal data acquisition unit is used to collect multimodal data from the rose planting area, including image data, sensor data, meteorological data, and crop information.

[0223] The feature fusion module is used for feature extraction, feature concatenation, and feature fusion of multimodal data. Different base networks, i.e., encoders, are used for feature extraction; a relearning method is employed to detect overfitting or underfitting of each modality's features, and the encoder parameters are updated and optimized accordingly; the updated encoder output features are concatenated and input into a cross-modal interaction layer for feature fusion to obtain multimodal fused features.

[0224] The medication application decision module utilizes a dual-branch decision network to identify pests and diseases and determine the type and dosage of medication. This dual-branch decision network includes a Transformer decoder, a medication type decision engine, and a regression network. The Transformer decoder decodes the multimodal fusion features, outputting a pest and disease type label sequence. The medication type decision engine then outputs a preliminary medication type decision. The regression network determines the dosage. Finally, a medication mixing control module introduces medication ratio rules to refine the medication type and dosage.

[0225] The user interaction feedback system is used to receive user correction instructions for the system's recommended medication dosage via a mobile APP after the user observes the effect of the implementation of the plan, thereby achieving closed-loop control and optimizing medication plan decisions.

[0226] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0227] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.

[0228] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.< / eos> < / sos> < / sos> < / eos> < / sos> < / sos>

Claims

1. A decision-making method for pesticide application in rose cultivation based on multimodal feature fusion, characterized in that, The steps include the following: Step 1: Collect multimodal data from the rose planting area; Step 2: Input the obtained multimodal data into the feature fusion module for multimodal feature fusion; Feature extraction is performed using different pre-trained base networks, i.e. encoders; a relearning method is used to detect whether the features of each modality are overfitted or underfitted, and the encoder parameters are updated and optimized; the features output by the updated encoders are concatenated and input into the cross-modal interaction layer for feature fusion to obtain multimodal fused features; Step 3: Medication regimen decision-making and drug mixing control; A dual-branch decision network is used to identify pests and diseases and determine the type and dosage of pesticides. The dual-branch decision network includes a Transformer decoder, a pesticide type decision engine, and a regression network. The Transformer decoder decodes the multimodal fusion features and outputs a pest and disease type label sequence. The pesticide type decision engine then outputs a preliminary pesticide type decision. Finally, the regression network is used to make a pesticide dosage decision. The drug mixing control module introduces drug ratio rules to modify the types and dosages of drugs. Step 4: Through user interaction optimization mechanisms, achieve closed-loop control and optimize medication plan decisions; The Transformer decoder operates as follows: (1) Decoder input construction: The multimodal fusion features are used as the initial hidden state H. enc ∈R d×n Add a sequence start character <sos>and position code E pos ;< / sos> (2) Iterative tag generation: The decoder consists of an L-layer structure, and each layer performs the following: Self-attention: This involves associating the current tag with historically generated tags to ensure semantic coherence; specifically, in the first round of generation, the historical tags are empty, and attention is only focused on... <sos> The attention computation is initialized with positional encoding, and new labels are appended to the historical sequence in each iteration for use in the next round of self-attention computation;< / sos> Encoder-decoder attention: using fused features as key-value pairs and the vector embedded by the current label as the query, it focuses on key regions of pests and diseases in the features; Feedforward prediction: Outputs the probability distribution P of the next label. t ∈R V Candidate tags are selected using beam search; (3) Category mapping and output: When the decoder predicts <eos> The generation process terminates when the probability of a label exceeds a threshold or the sequence reaches its maximum length. A rose-specific decision dictionary is used to map the labels to pest and disease types: each label corresponds to a specific type, and the pest and disease type with the highest confidence and its corresponding probability are output.< / eos> (4) Downstream interface: The pest and disease type, time encoding vector, and text feature vector are all input into the pesticide application decision engine; The specific drug type decision engine is as follows: The input sources for the decision engine include: the probability distribution of disease or pest categories output by the Transformer decoder, the time encoding vector extracted by LSTM, the predefined crop phenological weighting coefficients in the regression network, and the text feature vectors generated by LLM; The core decision-making rules include: establishing a prohibition rule base: combining information on rose phenology and plant variety disease resistance and drug resistance to prohibit the use of corresponding unsuitable agents; environmental adaptation mechanism: prohibiting the use of agents that are unsuitable under special weather conditions; specificity mechanism: different types of agents are used for different types and severity of pests and diseases, as well as for plants in different phenological stages. The medication type decision engine adopts a three-level hierarchical architecture, specifically including: (1) Input preprocessing layer: Standardize the data format of each input source: The probability distribution of pests and diseases is normalized to generate a confidence vector P∈[0,1]. n ; Temporal encoded vectors are fused with sensor data through linear projection to generate environmental feature embeddings E. t ∈R de ; Phenological sensitivity coefficient W is generated by mapping the basic coefficients. p ; (2) Rule execution layer: Based on the pest and disease confidence vector P and environmental features E obtained after input preprocessing. t and stage sensitivity coefficient W p The decision rules are executed according to the priority of the prohibition rule base - environmental adaptation mechanism - specific mechanism, and the decision on the type of drug to be applied is output. The specific steps include: The stage sensitivity coefficient Wp is used for decision-making in the prohibition rule base section, reflecting information on rose phenology and plant variety disease resistance and drug resistance, and prohibiting corresponding unsuitable drugs based on the knowledge base; Environmental characteristics Et are used for decision-making in the environmental adaptation mechanism. They reflect environmental information including wind speed, light, temperature and humidity, determine whether the environment is suitable for drug use, and prohibit the use of drugs that are not suitable for use under special weather conditions based on the knowledge base. The pest and disease confidence vector P is used for decision-making in the specific mechanism part, reflecting the type and severity of pests and diseases, as well as the phenological stage of the rose plant. For rose plants suffering from different pests and diseases and in different phenological stages, different types of pesticides are output based on the knowledge base. The rule enforcement layer outputs preliminary medication type decisions; (3) Correlation Matrix Constraint Layer: Add an attention scoring step to score the pesticides output by the rule execution layer. Add the disease-pesticide correlation matrix M∈{0,1} during scoring. d×d To prevent agronomical errors, mandatory restrictions should be placed on the types of drugs used. Where M ij =0 indicates that disease i and agent j are incompatible; Through the above three-level hierarchical architecture, the decision engine outputs the final drug type decision.

2. The method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion according to claim 1, characterized in that, The specific steps for collecting multimodal data from the rose planting area are as follows: Image data acquisition: Images including plant height, leaf morphology, and pest and disease patch characteristics are acquired using a visible light camera, and images including plant temperature distribution are acquired using an infrared camera; Time-series data acquisition from sensors: Real-time environmental parameters are acquired using soil moisture sensors, ambient temperature and humidity sensors, and light sensors; Meteorological data acquisition: Rainfall and wind speed are acquired using a weather station API; Crop information acquisition: A database of rose plant variety disease resistance and phenological stage identifiers are acquired; Simultaneously, multimodal data from historical planting records and data on the actual types and amounts of pesticides applied at that time are acquired, i.e., historical data; The acquired multimodal dataset containing historical and real-time data is divided into training, validation, and test sets according to the corresponding modes. The historical data is used for pre-training, relearning, and training of the regression network for the encoder, as well as for constructing a knowledge base for rose plant pesticides. The real-time data is used for real-time identification of pest and disease types and pesticide application decisions.

3. The method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion according to claim 2, characterized in that, The specific implementation of feature extraction using an encoder is as follows: For image data, CNN is used to extract the spatial features of pests and diseases from infrared and visible light images; for infrared images, preprocessing is required first, which involves temperature matrix standardization, temperature threshold segmentation, and pseudo-color encoding to obtain a single-channel pseudo-color image, and then feature extraction is performed using a ResNet50 network structure; for meteorological data and time-series data from sensors, LSTM is used for feature extraction to generate time-encoded vectors for each type of data; for crop information, a large language model (LLM) based on prompt word engineering is used to extract text feature vectors of crop information, and the disease resistance of rose plants under different conditions is quantified by pre-set prompt word template constraints before outputting structured data.

4. A method for decision-making on pesticide application in rose cultivation based on multimodal feature fusion as described in claim 1 or 3, characterized in that, The relearning method includes: first, diagnosing the learning state, i.e., extracting features for each modality, obtaining the purity of the dataset through the k-means clustering algorithm, i.e., the proportion of samples with the same label within the cluster of modality k dataset, and calculating the purity difference between the training set and the validation set respectively; then, performing soft reinitialization, calculating the reinitialization intensity based on the purity difference, updating the encoder parameters of the feature extraction part, and using the updated encoder for feature extraction of real-time data.

5. The method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion according to claim 1, characterized in that, The regression network uses a fully connected neural network to predict dose, and its specific calculation process is as follows: (1) Input feature standardization: multimodal fusion features F in cross-modal interaction layers fusion Batch standardization is performed using the following formula: Where, μ B , σ B Here are the statistics for the current batch, and γ and β are learnable parameters. (2) Nonlinear feature mapping: Dose-related features are extracted through a two-layer fully connected network. First layer: H1=ReLU(W1) + b1) Second layer: H2 = Dropout(ReLU(W2H1+b2)) Where H is the feature tensor, W is the weight matrix, and b is the output layer bias term; (3) Phenological stage feature modulation: Dynamically adjusting feature expression according to crop growth stages: H p =H2⊙(1+αW p ) Where H2 is the tensor of the output features of the second hidden layer, ⊙ denotes element-wise multiplication, α is the modulation intensity coefficient, and W p For phenological options, the weight vector is... (4) Dosage prediction output: The final predicted value is restored to the actual dose after being constrained by Sigmoid: y real =Q max ⋅σ(W3H p +b3) Where W3 is the output layer weight matrix, b3 is the output layer bias term, and Q... max This is the preset maximum safe dosage of the drug.

6. The method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion according to claim 1, characterized in that, The specific operation of the drug solution mixing control module is as follows: The specific rules for drug formulation are as follows: First, construct a four-dimensional matching matrix R∈R K×M’×P×E Where K is the type of pesticide, M' is the phenological stage of roses, P is the intensity of pests and diseases, and E is the environmental factor; Mixing ratio = R[k,m,p,e] × ×(1+αΔT) Among them, C 标准 The standard concentration is given by ΔT, which is the temperature difference between the agent and the environment to compensate for the error caused by the thermal expansion and contraction of the drug solution volume, and α is the temperature compensation coefficient. For medication type adjustments, the triggering of a substitution mechanism is determined based on whether the mixing ratio exceeds a threshold; for dosage adjustments, a baseline coefficient for concentration adjustment is calculated based on the mixing ratio; details are as follows: The types of drugs used are modified by setting up an alternative drug matching mechanism: First, alternative drug schemes are dynamically generated based on the drug mechanism database. Then, resistance accumulation detection is carried out and the historical drug use in the same plot is recorded. If it is detected that the same drug mechanism has been used repeatedly, the drug type is changed to avoid the accumulation of resistance. Based on real-time collected environmental factors, it is determined whether there are abnormal environmental conditions. When the environmental conditions are abnormal, the dosage is adjusted. Specifically, if the wind speed is higher than the threshold and it is not suitable for medication, medication is stopped; if the temperature or light intensity is higher than the threshold, the dosage is reduced; if the humidity is higher than the threshold, an anti-drift adjuvant is added to the medication.

7. The method for decision-making on pesticide application schemes for rose cultivation based on multimodal feature fusion according to claim 5, characterized in that, After the collected multimodal historical data is processed by the encoder updated through relearning, features are extracted, and then feature concatenation and fusion are performed. The fused data is then input into the regression network for training. The loss function for the regression network is the mean squared error of the phenological period constraint, which is related to the phenological weighting coefficient, as the loss function. Where N is the batch sample size, ω pheno The phenological weighting coefficient is determined based on the phenological stage of the rose. pred To predict drug dosage using the model, y true This represents the actual dosage used.

8. A decision-making device for pesticide application in rose cultivation based on multimodal feature fusion, used to implement the method described in claim 1, characterized in that, It consists of a multimodal data acquisition unit, a feature fusion module, a medication plan decision module, and a user interaction feedback system; The multimodal data acquisition unit is used to collect multimodal data from the rose planting area, including image data, sensor data, meteorological data, and crop information. The feature fusion module is used to extract, concatenate, and fuse features from multimodal data. It uses different base networks, i.e., encoders, to extract features. It uses a relearning method to detect whether the features of each modality are overfitted or underfitted and updates and optimizes the encoder parameters. The updated encoder output features are concatenated and input into the cross-modal interaction layer for feature fusion to obtain multimodal fused features. The medication application decision module uses a dual-branch decision network to identify pests and diseases and determine the type and dosage of medications. The dual-branch decision network includes a Transformer decoder, a medication type decision engine, and a regression network. The Transformer decoder decodes the multimodal fusion features and outputs a pest and disease type label sequence. Then, the medication type decision engine outputs a preliminary medication type decision. Regression networks are used to make medication dosage decisions; then, medication ratio rules are introduced through a drug-liquid mixing control module to correct the types and dosages of medications. The user interaction feedback system is used to receive user correction instructions for the system's recommended medication dosage via a mobile APP after the user observes the effect of the implementation of the plan, thereby achieving closed-loop control and optimizing medication plan decisions.