Multimodal open field vegetable management decision method and device, electronic equipment and storage medium

By using multimodal data fusion and knowledge graph decision-making, the problems of low operational accuracy and efficiency in open-field vegetable production have been solved, enabling precise control and dynamic optimization of open-field vegetable production, and improving production efficiency and decision-making accuracy.

CN121436497BActive Publication Date: 2026-06-02BEIJING RES CENT FOR INFORMATION TECH & AGRI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING RES CENT FOR INFORMATION TECH & AGRI
Filing Date
2025-10-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing open-field vegetable production, the lack of in-depth analysis and intelligent support based on big data leads to low operational accuracy and efficiency. Especially in the middle and late stages of production, changes in crop growth status and production environment have a significant impact on decision-making, and the response speed and accuracy of existing methods cannot meet the needs.

Method used

By integrating multi-source heterogeneous data, constructing a knowledge graph, and employing a dynamic collaborative control algorithm, multimodal perception data is acquired, modal features and feature fusion weights are determined, and decisions are made based on the agronomic knowledge graph, thereby achieving precise regulation and dynamic optimization of the entire process of open-field vegetable production.

Benefits of technology

It has improved the production efficiency of open-field vegetable production, ensured the real-time and accurate nature of decision-making, optimized agricultural management measures such as fertilization, irrigation and pest and disease control, and improved the precision of operations and the efficiency of resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436497B_ABST
    Figure CN121436497B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal decision-making method, device, electronic device, and storage medium for open-field vegetable management. The method includes: acquiring modal perception data corresponding to a target open-field vegetable field; determining modal features corresponding to each modal perception data point and determining feature fusion weights corresponding to the modal features at each time anchor point; determining global multimodal fusion features based on the feature fusion weights and modal features; determining knowledge representation features from an agronomic knowledge graph based on the modal perception data; and determining a target decision scheme based on the knowledge representation features and the global multimodal fusion features. By integrating multiple data sources such as meteorology, soil, crop status, and agricultural machinery operating parameters, and employing spatiotemporal alignment and adaptive weighting mechanisms, combined with reinforcement learning technology, the method dynamically adjusts agricultural machinery operation strategies based on real-time data, optimizes agricultural management measures such as fertilization, irrigation, and pest and disease control, improves production efficiency, and ensures the real-time nature and accuracy of decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a multimodal decision-making method, apparatus, electronic device, and storage medium for open-field vegetable management. Background Technology

[0002] Open-field vegetable production involves multiple stages, encompassing the entire process from land preparation, planting, field management to harvesting. Each stage is influenced by numerous factors, including weather conditions, agronomic management, and the status of machinery operations. Current agricultural informatization primarily relies on single sensors or remote sensing imagery for localized data collection and monitoring. However, in open-field vegetable production, factors such as climate change, soil water and fertilizer conditions, and pest and disease transmission are highly uncertain. Monitoring methods using single data sources cannot promptly capture the complex dynamic relationship between the environment and crops, nor can they adequately meet the demands for precise decision-making in agricultural production.

[0003] However, traditional decision-making methods largely rely on experience-based judgment and lack in-depth analysis and intelligent support based on big data, resulting in low operational accuracy and efficiency. This is especially true during the middle and late stages of production, when changes in crop growth and the production environment have a significant impact on decision-making, and existing methods cannot meet the required response speed and accuracy. Furthermore, mechanized and automated operations in agricultural production also face technological limitations. Most existing agricultural mechanization operates independently, making it difficult to effectively integrate with environmental changes, crop growth status, and agronomic management, leading to low resource utilization efficiency and poor operational results. Summary of the Invention

[0004] This invention provides a multimodal open-field vegetable management decision-making method, device, electronic device and storage medium. By integrating multi-source heterogeneous data, constructing a knowledge graph and adopting a dynamic collaborative control algorithm, it achieves precise control and dynamic optimization of the entire open-field vegetable production process.

[0005] According to one aspect of the present invention, a multimodal open-field vegetable management decision-making method is provided, comprising:

[0006] Acquire modal sensing data corresponding to the target open-field vegetable field, wherein the modal sensing data includes segment sensing data, crop sensing data, agricultural machinery sensing data, and agronomic operation parameters;

[0007] Determine the modal features corresponding to each modal sensing data, and determine the feature fusion weights corresponding to the modal features at each time anchor point. Based on the feature fusion weights and the modal features, determine the global multimodal fusion features.

[0008] Based on the modal perception data, knowledge representation features are determined from the agronomic knowledge graph, and a target decision scheme is determined based on the knowledge representation features and the global multimodal fusion features.

[0009] According to another aspect of the present invention, a decision-making device for open-field vegetable management based on multimodal perception is provided, comprising:

[0010] The data acquisition module is used to acquire modal sensing data corresponding to the target open-field vegetable field, wherein the modal sensing data includes segment sensing data, crop sensing data, agricultural machinery sensing data, and agronomic operation parameters;

[0011] The feature fusion module is used to determine the modal features corresponding to each modal sensing data, and to determine the feature fusion weights corresponding to the modal features at each time anchor point, and to determine the global multimodal fusion features based on the feature fusion weights and the modal features;

[0012] The decision module is used to determine knowledge representation features from the agronomic knowledge graph based on the modal perception data, and to determine the target decision scheme based on the knowledge representation features and the global multimodal fusion features.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the multimodal open-field vegetable management decision-making method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the multimodal open-field vegetable management decision-making method according to any embodiment of the present invention.

[0018] The technical solution of this invention acquires modal perception data corresponding to a target open-field vegetable field, determines modal features corresponding to each modal perception data point, and determines feature fusion weights corresponding to the modal features at each time anchor point. Based on the feature fusion weights and the modal features, a global multimodal fusion feature is determined. Then, based on the modal perception data, knowledge representation features are determined from an agronomic knowledge graph. Finally, a target decision-making scheme is determined based on the knowledge representation features and the global multimodal fusion feature. Based on this technical solution, by integrating multiple data sources such as meteorology, soil, crop status, and agricultural machinery operating parameters, and employing spatiotemporal alignment and adaptive weighting mechanisms, combined with reinforcement learning techniques, agricultural machinery operation strategies can be dynamically adjusted based on real-time data. This optimizes agricultural management measures such as fertilization, irrigation, and pest and disease control, thereby improving production efficiency and ensuring the real-time nature and accuracy of decision-making.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a multimodal open-field vegetable management decision-making method provided in an embodiment of the present invention;

[0022] Figure 2 This is a flowchart of the multimodal open-field vegetable management decision-making method provided in this embodiment of the invention;

[0023] Figure 3 This is a flowchart of a multimodal open-field vegetable management decision-making method provided in an embodiment of the present invention;

[0024] Figure 4 This is a schematic diagram of the structure of an application control device for a wearable device provided in an embodiment of the present invention;

[0025] Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Example 1

[0029] Figure 1 This is a flowchart illustrating a multimodal open-field vegetable management decision-making method provided in an embodiment of the present invention. This embodiment is applicable to situations where a management plan corresponding to open-field vegetables is determined based on collected multimodal data. This method can be executed by an open-field vegetable management decision-making device based on multimodal perception. This device can be implemented in hardware and / or software and can be configured in an electronic device, such as a PC or a server. Figure 1 As shown, the method includes:

[0030] S110. Acquire modal sensing data corresponding to the target open-field vegetable field.

[0031] The modal sensing data includes process sensing data, crop sensing data, agricultural machinery sensing data, and agronomic operation parameters.

[0032] Specifically, high-precision weather stations are deployed in key areas of the vegetable fields to collect environmental parameters such as temperature, humidity, light intensity, wind speed, and rainfall in real time, and the data is synchronized to a central processing unit via wireless transmission modules. Simultaneously, soil sensors are embedded at different soil depths to monitor soil temperature, humidity, pH value, and nutrient content such as nitrogen, phosphorus, and potassium, ensuring comprehensive soil environmental data. Regarding crop status, high-definition cameras and multispectral cameras are used to periodically capture crop images, and deep learning algorithms are used to extract crop height, leaf area index, and pest and disease characteristics. Furthermore, sensors built into agricultural machinery record parameters such as operating speed, tillage depth, and posture in real time, and these are structurally integrated with manually set agronomic operation records (such as fertilizer application rate and irrigation time).

[0033] For example, by deploying weather stations, soil sensors, and cameras in open-field vegetable fields, real-time collection of multi-source data such as temperature and humidity, precipitation, soil moisture, crop growth status, and agricultural machinery operation trajectories can be achieved, providing basic support for subsequent decision-making. The data collected in multimodal sensing is defined as: environmental sensing data. Data collected by weather stations and soil sensors includes temperature and humidity, light intensity, wind speed, rainfall, soil temperature and humidity, pH, and nutrient content. Crop sensing data. Agricultural machinery sensing data includes crop height, leaf area index, and images of pests and diseases acquired using cameras, multispectral cameras, and depth sensors. Real-time operating parameters of agricultural machinery are collected, including operating speed, operating path, tillage depth, posture, power output, and fuel consumption. Management behavior records are also included. Recorded, manually set agronomic operation parameters, including planting density, operation time, fertilization and irrigation, are fused with sensor data in a structured manner.

[0034] S120. Determine the modal features corresponding to each modal sensing data, and determine the feature fusion weights corresponding to the modal features at each time anchor point. Determine the global multimodal fusion features based on the feature fusion weights and the modal features.

[0035] Modal features can be data features corresponding to each modality of data. Global multimodal fusion features can be understood as global feature data obtained by weighted fusion of features from multiple modalities.

[0036] Specifically, for environmental perception data, a Transformer encoder extracts temporal dependency features and utilizes a self-attention mechanism to capture long-range meteorological change patterns. For crop perception data, a convolutional neural network extracts spatial texture features and combines them with depth sensor data to generate three-dimensional morphological features. Agricultural machinery operating parameters are modeled using a temporal convolutional network to model dynamic operating modes. After the signal-to-noise ratio of each modal feature is evaluated by a modality quality scorer, an anchor point alignment mechanism is introduced to achieve spatiotemporal synchronization, and then an adaptive weighted Transformer dynamically allocates fusion weights. Finally, the weighted multimodal features are input into a graph convolutional neural network, and relational reasoning is performed based on the constructed agronomy-environment-machinery knowledge graph to generate global fusion features containing cross-modal causal dependencies, providing semantically consistent data support for decision-making.

[0037] Based on the above technical solution, determining the modal features corresponding to each modal sensing data includes: normalizing each modal sensing data to obtain standardized data corresponding to each modal sensing data; and extracting temporal features from the standardized data based on the encoder to determine the modal features corresponding to each standardized data.

[0038] The standardized data can be the data obtained by normalizing the collected raw data. The encoder can be a pre-set Transformer encoder.

[0039] Specifically, Z-Score normalization is performed on environmental perception data, crop perception data, and agricultural machinery operating parameters: the mean of each modality is subtracted and then divided by the standard deviation to map the data to a standard normal distribution interval, eliminating dimensional differences. Subsequently, based on the Transformer encoder structure, feature extraction is performed on the normalized temporal data: long-range temporal dependencies are captured through a multi-head self-attention mechanism, and temporal order information is preserved by combining position encoding. For example, preprocessing, spatiotemporal alignment, and normalization are performed on each modality to ensure that the input data is represented within a unified time window and coordinate system. For instance, the original input data is... ,(in Each corresponds to different modal data from the first step, including environmental data, crop data, agricultural machinery parameter data, and agronomic operation parameters. For dimension The matrix, Indicates the number of time steps. This represents the feature dimension of the modality.

[0040] Based on the above technical solution, the step of extracting temporal features from the standardized data using an encoder to determine the modal features corresponding to each standardized data includes: for each standardized data, performing a linear transformation on the standardized data to generate a query matrix, a key matrix, and a numerical matrix corresponding to the standardized data; calculating the similarity matrix between the query matrix and the key matrix; and determining the modal features corresponding to the standardized data based on the similarity matrix and the numerical matrix.

[0041] The query matrix represents the information needed for the currently processed element, while the key and numerical matrices store query-related information. The similarity matrix can be understood as an attention score matrix.

[0042] Specifically, for normalized environmental sensing data (such as temperature and humidity sequences), crop sensing data (such as multispectral image feature sequences), and agricultural machinery operating parameters (such as speed time-series data), linear projection operations are performed respectively: through three independent fully connected layers, the standardized data of each modality is mapped to three low-dimensional matrix spaces: query (Q), key (K), and value (V), where the matrix dimensions are determined by the data time-series length and feature dimensions. The dot product similarity between the query matrix and the key matrix is ​​calculated, and a scaling factor is introduced to stabilize the gradient, generating a similarity matrix.

[0043] The attention weight distribution at each time step is obtained through softmax normalization. Finally, the normalized weight matrix is ​​multiplied by the numerical matrix to complete the weighted aggregation of temporal information, outputting a high-dimensional modal feature vector that incorporates the global context. By dynamically adjusting the contribution of different time steps, long-range dependencies in each modality are effectively captured. To improve the robustness of the feature representation, a multi-head attention mechanism is used to process multiple subspaces in parallel. The outputs of each head are concatenated and linearly transformed to generate the final modal features, providing a deep representation with temporal semantics for subsequent multimodal fusion.

[0044] For example, a deep fusion algorithm combining adaptive weighted Transformer and Graph Convolutional Neural Network (GCN) is used to efficiently fuse various types of data. This eliminates the differences between different modalities of data, unifies data representation, and thus improves the accuracy and real-time performance of decision-making.

[0045] To address the temporal characteristics of each modality's data, a Transformer model is employed, employing a self-attention mechanism to capture long-range dependencies and global information. To cope with differences in data quality and noise, an adaptive weighting mechanism is introduced, enabling the model to automatically adjust the weight of each modality in the overall representation based on its signal-to-noise ratio and information contribution.

[0046] A Transformer encoder is used to extract temporal features from each modal data, outputting a high-dimensional feature vector for each time step. For each modal data... Generate queries through linear transformation ,key Sum of values : ,in, The input data matrix; The weight matrix for the query and the key; The weight matrix is ​​a numerical value.

[0047] The output of self-attention is: ,in, This represents the similarity matrix between the query and the key. This is a scaling factor used to stabilize the gradient; Normalize each row to generate attention weights.

[0048] Based on the above technical solution, the step of determining the global multimodal fusion feature according to the modal features includes: aligning the modal features in the time dimension based on a preset set of time anchors; determining the evaluation score corresponding to the modal features of each time anchor based on a modal quality scorer for the time-dimensional aligned modal features; determining the feature fusion weight corresponding to the modal features of each time anchor based on the evaluation score; performing weighted fusion of the modal features of each time anchor based on the feature fusion weight to obtain a multimodal temporal representation; and determining the global multimodal fusion feature based on the multimodal temporal representation.

[0049] The preset time anchor set can be a collection of anchor points with a fixed pre-defined time length, and the preset time anchor set includes at least two time anchor points. The modality quality scorer can be understood as a network structure used to score the modality corresponding to each anchor point. The feature fusion weights can be weight values ​​used for weighted fusion.

[0050] Specifically, a discrete-time anchor point set is predefined, and the feature sequences of each modality are unified to the anchor point time window through linear interpolation or average sampling to achieve cross-modal temporal alignment. Subsequently, a modal quality scorer based on the Transformer architecture is deployed. This scorer takes the multimodal features at each anchor point as input, calculates the intermodal correlation through a self-attention mechanism, and generates an evaluation score by combining the signal-to-noise ratio (SNR) metric. For the global representation of each modal feature at anchor point τ, a multilayer perceptron is used to predict its quality score, and after Softmax normalization, weight coefficients in the [0,1] interval are obtained. Finally, the features of each modal anchor point are weighted and summed according to the weight coefficients to generate a multimodal temporal representation: the feature vectors of the environment, crop, and agricultural machinery modalities at anchor point τ are multiplied by their corresponding weights and then summed to form a fused feature. A graph convolutional network is used to further aggregate the spatial correlation information in the multimodal temporal representation, ultimately outputting a global fused feature that combines temporal dynamics and cross-modal causality, providing semantically consistent data support for decision-making.

[0051] Based on the above technical solution, the step of aligning the modal features in time dimension according to a preset time anchor set includes: determining the anchor time length corresponding to the preset time anchor set, and determining the feature time length corresponding to the modal feature; when the feature time length is greater than the anchor time length, performing sliding window sampling on the modal feature according to the preset time anchor set; when the feature time length is less than the anchor time length, performing interpolation processing on the modal feature according to the preset time anchor set.

[0052] The anchor point time length can be a time length value corresponding to the time anchor point set. Users can set it according to their needs and then divide the anchor point time length according to the preset step size to obtain the preset time anchor point set. The feature time length can be understood as the time length value corresponding to the modal feature.

[0053] Specifically, the anchor time length corresponding to the preset time anchor set is determined. This length is jointly determined by the preset time interval and the number of anchors, reflecting the desired time scale for alignment. Simultaneously, the feature time length corresponding to the modal feature is determined, i.e., the actual span of the modal feature on the time axis. Then, a time length comparison is performed. If the feature time length is greater than the anchor time length, it indicates that the modal feature is too long in time and needs compression. Based on the preset time anchor set, a sliding window sampling method is used to slide a window across the modal feature at a certain step size, extracting feature information within the window to generate a feature sequence matching the anchor time length. If the feature time length is less than the anchor time length, it indicates that the modal feature is too sparse in time and needs expansion. Based on the preset time anchor set, an interpolation method is used to insert new feature points on the time axis of the modal feature, and reasonable estimation is performed based on the information of adjacent feature points to generate a feature sequence that meets the anchor time length requirements.

[0054] For example, to achieve time alignment between different modalities, a unified set of anchor points is introduced. Its length The set unified fusion time number. For each mode i, construct the anchor-level feature sequence. If Then for Using a sliding window average sampling method, if The original sequence is then interpolated and expanded to ensure that all modalities have the same length in the anchor domain, thus aligning subsequent fusion operations in the time dimension.

[0055] Based on the above technical solution, the step of determining the feature fusion weights corresponding to each anchor feature according to the evaluation score includes: for the modal feature, determining the evaluation score corresponding to each anchor feature corresponding to the modal feature; and normalizing the evaluation scores corresponding to each anchor feature based on the normalized fusion function to determine the feature fusion weights corresponding to each anchor feature in the modal feature.

[0056] The evaluation score can be the numerical value obtained by the modal quality scorer after scoring the features. It should be noted that the modal quality scorer can be a pre-defined function used to evaluate the features. Designing a network with shared parameters is also important. As a "modal quality scorer", it evaluates the features of each mode at each anchor point τ. Generate score : ,in As a weighting factor, As a bias term, the above scorers are shared across all modalities and anchor points.

[0057] Specifically, the evaluation scores corresponding to each anchor feature in the modal features are determined. It should be noted that the evaluation scores reflect the relative importance of each anchor feature in a specific task. To convert the evaluation scores into weight values ​​that can be used for feature fusion, a normalized fusion function is used to normalize the evaluation scores corresponding to each anchor feature. Normalization maps evaluation scores with different dimensions and ranges to a unified numerical range [0,1], thereby eliminating the influence of dimensional differences on weight allocation. The feature fusion weights corresponding to each anchor feature in the modal features are determined through normalization. The feature fusion weights reflect the relative contribution of each anchor feature in the fusion process; the larger the weight, the higher the importance of that anchor feature during fusion.

[0058] For example, weights are generated through Softmax normalization to calculate all modality scores. Input Softmax: , This process guarantees that for each anchor point τ, ,and This allows high-quality, high-signal-to-noise ratio modes to automatically acquire greater fusion weights. Finally, weighted spatiotemporal fusion is performed. Based on the feature fusion weights, at each anchor point τ, the aligned features of the four modes are summed according to their corresponding weights to generate the fused features. The fused sequence This will be used as a unified multimodal timing representation and fed into the subsequent transformer decoding module, where... The number of fusion times is preset, and D is the dimension of the feature. This feature dimension corresponds to multiple preset feature data, that is, to different modal data such as environmental data, crop data, agricultural machinery parameter data and agronomic operation parameters.

[0059] Based on the above technical solution, the step of determining the global multimodal fusion feature based on the multimodal temporal representation includes: constructing graph structure nodes based on the multimodal temporal representation, and calculating the similarity between modal features and the similarity between small nodes under the same modal feature according to cosine similarity; determining edge weights based on the similarity, and constructing an initial weighted graph based on the graph structure nodes and the edge weights, and performing feature aggregation based on the initial weighted graph to determine the global multimodal fusion feature.

[0060] The graph structure nodes include global nodes and local sub-nodes; the global nodes are used to represent the global features of a single modality of data; and the local sub-nodes are used to represent the features of key data types in the current modality.

[0061] Specifically, graph structure nodes are generated based on multimodal temporal representation: cross-modal features (environment, crops, agricultural machinery) at each time anchor point are concatenated into the initial feature vector of the node. Simultaneously, different sub-features within the same modality (such as temperature, humidity, and light intensity in the environmental modality) are individually constructed as sub-nodes, forming a hierarchical node structure. Node similarity is calculated to determine edge weights: for cross-modal node pairs, cosine similarity is used to measure the consistency of feature space direction; for sub-nodes within the same modality, dynamic cosine similarity with a time decay factor is introduced to strengthen recent temporal correlation. An initial weighted graph is generated based on the similarity matrix: only edges with similarity higher than a threshold are retained, and edge weights are normalized to the [0,1] interval. Finally, feature aggregation is performed through a graph attention network: when each node aggregates the features of its neighboring nodes, attention weights are dynamically learned to highlight the contributions of highly similar neighbors. After multi-layer graph convolutional propagation, the output of the root node (full-cycle aggregation node) contains globally fused features with cross-modal spatiotemporal dependencies, providing a structured semantic expression for subsequent decision-making tasks.

[0062] For example, the global features output by the adaptive Transformer are treated as nodes. Edge weights are constructed by calculating the data similarity between and within modalities, forming a multimodal data relationship graph. This facilitates subsequent use of Graph Convolutional Networks (GCNs) to perform information transfer and aggregation on the graph structure, thereby capturing environmental data. Crop data Agricultural machinery parameter data and agronomic operating parameters The deep causal relationships and dependencies between data sources.

[0063] It should be noted that the nodes in the multimodal data relationship graph of this invention are divided into two levels, including global nodes and local small nodes. Global nodes represent the global features of a modality of data (X1~X4) or its aggregated representation in different time periods, while local small nodes are used to characterize the features of key data types under the current modality.

[0064] For global nodes, let the feature representation of the i-th modality in the fused feature matrix Y be as follows: The corresponding features are obtained using time-series average pooling (MeanPool): While using large nodes, local small nodes are constructed for key data types (such as temperature) in each modality. Here, d corresponds to the characteristics of key data types in different modal data, for example, temperature corresponds to the environmental data modality.

[0065] For local small nodes, let mode i have a total of Data types If each time step is t, then the features formed by one-dimensional convolution (Conv1D) and average pooling for each data type are represented as follows: .

[0066] Based on the global nodes and local small nodes obtained after the above processing, construct the node set V in the graph, including all large nodes and their corresponding small nodes: To construct the edge weights between nodes, cosine similarity is used to calculate the similarity between modal features (between large nodes) and between small nodes of the same modality. The edge weight between any two nodes i and j is defined as: ;in, These represent the node characteristics. Represents a node With nodes Similarity; , denoted as the Euclidean norm corresponding to the node feature.

[0067] It should be noted that by setting a threshold in advance... When the calculated edge weights At the node With nodes Establish an edge between them, that is, for different edge weights have: Furthermore, each modality's major node is directly connected to its subordinate minor nodes, and a connection is set... Through the above steps, a weighted graph with a clearly defined hierarchical structure is finally constructed. The node set V includes all large and small nodes, and the adjacency matrix A contains all the edges between nodes and their weights.

[0068] Furthermore, based on the constructed graph structure, a Graph Convolutional Network (GCN) is used to aggregate and update node features. The update formula for each layer of the graph convolution is: ;in, Indicates the first The "intermediate representation" of a layer is both the input to this layer's convolution and the output of the previous layer's convolution. To add the adjacency matrix after adding the self-loop, Represents the adjacency matrix of the original graph. It is the identity matrix; for The degree matrix is ​​defined as ; For the first The learnable weight matrix of the layer; This is a non-linear activation function. Initial layer. Then take the initial representation of each node. The matrix formed.

[0069] The node features aggregated by GCN are used as the final output of multimodal data fusion, providing a high-quality representation with causal dependency information for subsequent growth state prediction. After a series of graph convolutions, the final representation of each node is obtained. ,in Let be the number of layers in the graph convolutional neural network. All node representations are concatenated to form the final multimodal fusion feature. This feature encompasses the spatiotemporal information and causal dependencies of each modality of data, providing strong data support for subsequent accurate decision-making. After information aggregation and representation learning in the multi-layer graph convolutional network, the final feature representation of each node is obtained: Where M is the total number of nodes in the graph (in this invention, there are 4 nodes, which represent the four modes of environment, crop, agricultural machinery and management respectively). The graph convolutional neural network represents the first... The output of the first layer Each node's feature vector represents a mode. Global context-aware feature representation in multimodal fusion process. It not only includes the semantic information of the modality itself, but also the edge weights in the graph structure. It guides cross-modal information propagation and fusion. In each layer of a graph convolutional network, the information transfer and aggregation mechanism between nodes is influenced by edge weights. The effect of this is that the larger the edge weight, the stronger the semantic similarity or dependency between modalities, and the higher the corresponding information transmission weight. This mechanism makes the final representation... It offers greater causal interpretability, reflecting potential causal relationships such as the impact of environmental changes on crop status and the feedback of agronomic operations on agricultural machinery operations. The feature vectors of each node are directly concatenated along the feature dimension in a predetermined order to form a continuous vector, seamlessly connecting the features of each node in the column direction. Finally, the representations of all nodes are concatenated to obtain the global multimodal fusion feature. ;

[0070] The fusion feature Z generated through the above steps not only integrates the temporal evolution patterns and spatial structural relationships of various modal data, but also encodes the dependency paths obtained from intermodal causal reasoning, thus realizing the efficient expression of multimodal data in a unified vector space.

[0071] S130. Based on the modal perception data, determine the knowledge representation features from the agronomic knowledge graph, and determine the target decision scheme according to the knowledge representation features and the global multimodal fusion features.

[0072] The agronomic knowledge graph can be a pre-constructed knowledge graph associated with open-field vegetables. Knowledge representation features can be understood as the feature representations of knowledge extracted from the knowledge graph. The target decision scheme can be the final determined decision scheme for treating open-field vegetables.

[0073] Specifically, based on a pre-constructed agronomic knowledge graph (containing triple relationships such as crop growth cycle, environmental-agronomic response rules, and agricultural machinery operation specifications), entities (such as "tomato flowering period") and relationships (such as "increased water requirement") in the graph are mapped into low-dimensional knowledge representation vectors using graph embedding algorithms (such as TransE), forming a computable agronomic prior knowledge base. Subsequently, global multimodal fusion features (including dynamic information on environment, crops, and agricultural machinery) are aligned with knowledge representation features across modalities: semantic similarity between multimodal features and knowledge entities is calculated through an attention mechanism, dynamically activating relevant agronomic rules (such as activating the "increased risk of tomato flower drop" rule under high temperature weather). Finally, the decision-making module integrates the activated knowledge rules with real-time multimodal states, and outputs the target decision using a reinforcement learning framework: with agronomic knowledge as constraints, parameters such as irrigation amount and fertilizer ratio are optimized through a policy network to ensure that the decision simultaneously meets crop growth needs and sustainable production goals, forming a closed-loop decision-making system of "data-driven - knowledge constraint".

[0074] like Figure 2 As shown, this invention utilizes a multimodal perception technology that integrates adaptive weighted Transformer and graph convolutional neural networks, combined with multi-task learning and adaptive optimization methods, to achieve comprehensive optimization of all elements—human, machinery, environment, and agronomy—in open-field vegetable production. It can accurately perceive the crop growth environment, intelligently analyze water and fertilizer requirements, and optimize agricultural machinery operation scheduling, thereby improving the level of intelligent agricultural production. Compared to traditional experience-based planting methods, this invention significantly improves operational accuracy, water and fertilizer utilization efficiency, and pest and disease control capabilities, while reducing reliance on manual labor and resource waste, providing an integrated solution for large-scale, intelligent open-field vegetable production.

[0075] For example, after fusing multimodal perception information and completing agronomic-environment-machinery correlation reasoning through knowledge graphs, not only are real-time perception results of the field environment and crop status obtained, but also structured knowledge such as operational specifications, agricultural machinery configuration requirements, and operational condition constraints derived from expert knowledge and historical experience are acquired. Based on this, the agricultural machinery collaborative control module fully utilizes the following two types of information sources when making decisions: on the one hand, real-time fusion features This invention provides the current perceived state of the environment, crops, and machinery operation. On the other hand, the knowledge graph reasoning results provide a scientific basis for agricultural machinery operation tasks, including recommended operation time windows, selection of suitable machinery types, suggestions for adjusting operation parameters (such as sowing depth, fertilizer application rate, and pesticide dosage), and rule constraints for multi-machine collaborative operation. To achieve efficient collaborative operation of multiple agricultural implements in open-field vegetable production scenarios, this invention introduces a multi-task learning and reinforcement learning mechanism to achieve intelligent optimization of task allocation. In this stage, the input data includes not only environmental data... Crop data Agricultural machinery parameter data and agronomic operating parameters This information, along with the global feature representation Z output from the GCN in the previous module (multimodal data fusion), is incorporated. This feature Z is obtained through... Causal relationship modeling of multimodal data provides environmental semantic information with global context awareness, offering strong knowledge support for agricultural machinery operation scheduling tasks.

[0076] It should be noted that, based on the needs of open-field vegetable production, agricultural machinery operations are categorized into tasks such as sowing, fertilizing, irrigating, weeding, and harvesting, each with different operational requirements. Multi-task learning is used to train the operational tasks. By sharing features among operational tasks, including crop variety, soil type, and weather information, a multi-task optimization model is trained. This model can share the learning process across multiple tasks, enabling automatic allocation of the most suitable operational task when multiple agricultural machines are operating in parallel. For each task… Design a dedicated fully connected layer (FC) network branch for the task. Predicted value It can be represented as: ,in For the task The weight matrix, The input feature vector contains fused environmental perception information and task prior features; For the task The weight matrix; For the task The bias; The activation function. Weight matrix. It is automatically learned during training using the gradient descent optimization algorithm. The model is based on historical task execution data and continuously adjusts itself using the backpropagation algorithm by minimizing the error between predicted values ​​and actual performance. The value of optimizes the model's predictive ability. In multi-task learning scenarios, all task branches share the input features. Each has its own independent parameter matrix By using a joint loss function for overall training, the goal of information sharing among tasks while maintaining appropriate differentiation is achieved. Loss function settings: Each task branch has an independent loss function to calculate its prediction error. A weighted average is used to combine the losses from each task to form the total loss, which guides the backpropagation optimization of the entire network. For each task... The mean squared error (MSE) is used as the loss function: ;in, Indicates the number of samples; Indicates the first The true value of each sample; This represents the predicted value.

[0077] The total loss is obtained by weighting the losses from each task: ,in, Total number of tasks; For the task The weighting coefficients satisfy and Weighting coefficients The method for obtaining it is: [to obtain each] These are set as learnable network parameters, updated along with the model parameters during training. A softmax transformation is introduced to ensure that they satisfy probability constraints. ;in, For the task The corresponding learnable parameters. During training, the model automatically adjusts the importance of each task, making the training process more robust and adaptive.

[0078] Based on the above technical solution, the step of determining knowledge representation features from the agronomic knowledge graph based on the modal sensing data includes: determining a state vector based on the modal sensing data, determining a knowledge query template based on the state vector, obtaining query results from the agronomic knowledge graph based on the knowledge query template, and determining the knowledge representation features based on the query results.

[0079] Here, the state vector can be understood as a vector used for data querying. The knowledge query template can be a query template obtained by filling in a preset query template based on the state vector.

[0080] Specifically, a dynamic state vector is constructed based on multimodal sensing data. Environmental modalities (temperature, humidity, light, etc.), crop modalities (growth stage, pest and disease characteristics), and agricultural machinery modalities (operational parameters, equipment status) are time-aligned and normalized, then encoded into a fixed-dimensional state vector using an LSTM network. This vector contains real-time state information of the current field, such as numerical representations of key elements like [temperature 28℃, humidity 75%, tomato third flowering stage, moderate soil nitrogen content]. Secondly, a hierarchical knowledge query template is designed. Based on the crop growth stage (e.g., flowering stage) and environmental conditions (e.g., high temperature and high humidity) in the state vector, a structured query statement is matched from a predefined template library. For example, when "tomato flowering stage + humidity > 80% for 3 consecutive days" is detected, a SPARQL query is automatically generated: "SELECT Pest and Disease Type FROM Agronomic Knowledge Graph WHERE Crop = 'Tomato' AND Growth Stage = 'Flowering Stage' AND Environmental Conditions = 'Humidity > 80%' AND Associated Risk = 'Pest and Disease'". Finally, knowledge graph query and feature fusion are performed. Queries are executed using the Neo4j graph database engine to obtain agronomic rules (e.g., "high humidity during tomato flowering period easily leads to gray mold") and treatment suggestions (e.g., "reduce humidity to below 60% and spray with iprodione") that match the current state. The query results are encoded into knowledge representation vectors and fused with the original state vectors using attention-weighted methods to generate enhanced decision features containing prior agronomic knowledge, providing a scientific basis for subsequent precision farming operations.

[0081] For example, to achieve adaptive scheduling and agronomic compliance control of multiple machines in open-field vegetable scenarios, this invention introduces a reinforcement learning mechanism, defines the job state space, action space, and reward function, and deeply integrates the knowledge graph reasoning results into multiple key stages of the decision-making process, thereby constructing a hybrid optimization strategy of "data-driven + knowledge-guided". The core components of the reinforcement learning model include three parts: state space definition, action space design, and reward function construction. The generation, recall, and usage of knowledge graph reasoning results are explained in detail below. First, at the knowledge graph level, an agronomic-environment-machinery graph ontology is pre-constructed. Offline reasoning is performed using rule engines such as SWRL and graph neural network embedding models to produce structured knowledge related to job specifications. For example, under the combined conditions of "current weather conditions + crop growth stage + soil parameters", a triplet recommendation result of "recommended sowing time is mid-April, use a certain model of seeder, and sowing depth is 30mm" is inferred. The inference results are stored in a graph database (such as Neo4j) and a semantic endpoint (such as Fuseki) in a structure of <task, preconditions, recommended configuration>, supporting online querying and retrieval. During the online decision-making phase, at each time step, the current state vector is set. for: Where T represents the current temperature (°C); H represents the current relative humidity (%); R represents the current precipitation (mm); M represents the current soil moisture (%); and G represents the crop growth status index. The SPARQL query template is automatically populated based on the current state vector to retrieve knowledge graph inference results matching the current operating environment conditions. Subsequently, the recall results are ranked based on rule confidence and entity similarity, and the optimal k results are selected and converted into vector representations using an embedding model. .

[0082] Based on the above technical solution, the step of determining the target decision scheme according to the knowledge representation features and the global multimodal fusion features includes: concatenating the knowledge representation features and the global multimodal fusion features to determine the target state features; and determining the target decision scheme based on the target state features and a preset reward function.

[0083] The target state feature is obtained by concatenating existing features. The preset reward function can be understood as a pre-set reward function used for iterative calculations.

[0084] Specifically, knowledge representation features (including structured knowledge such as agronomic rules and crop growth constraints) are fused with global multimodal fusion features (temporal features integrating dynamic data of the environment, crops, and agricultural machinery) through a gating concatenation mechanism. This is implemented by using a learnable weight matrix to linearly transform the two types of features, generating dynamic gating coefficients through the Sigmoid function to achieve adaptive selection and weighted summation of feature dimensions. For example, when a high-temperature drought condition is detected, the weights of knowledge features such as "crop water requirement patterns" are automatically increased to generate a target state feature vector containing real-time environmental conditions and agronomic constraints. A decision generation model based on deep reinforcement learning is then designed. Using the fused target state features as input, a policy network is constructed using the PPO algorithm, and high-level decision features are extracted through multiple fully connected layers and the ReLU activation function. A preset reward function comprehensively considers dimensions such as crop yield prediction, resource utilization efficiency, and agronomic compliance; for example, reward value = α × yield increment + β × water saving rate - γ × penalty for violations. During training, the strategy network parameters are optimized by comparing and learning with decision-making data from agronomic experts, so that the generated decision schemes (such as irrigation amount and fertilizer ratio) gradually approach the optimal solution. The output includes the target decision scheme containing specific agricultural operation parameters, and the decision basis (such as activated agronomic rules and real-time environmental risk warnings) is displayed through a visual interface, realizing interpretable closed-loop control of "data-knowledge-decision".

[0085] For example, a unified knowledge representation can be obtained by averaging all inference embedding vectors. This is further concatenated with the environmental semantic vector Z generated by multimodal perception fusion in the previous module to form the final state input for reinforcement learning: ;,in This indicates a splicing operation. As the state input for reinforcement learning, it includes not only environmental sensing data but also explicitly incorporates structured knowledge reasoning results, thus providing a more semantically relevant decision-making context. As the decision outcome of reinforcement learning, the action space... This represents the various executable agricultural management operations, with each action corresponding to a decision-making option. Define the action. For a vector: Where: I represents irrigation amount (L / m²); F represents fertilizer amount (kg / mu); P represents pesticide spraying amount (L / mu). Each action directly affects crop growth and resource consumption, and is an important output of decision optimization. Reward function Used to evaluate the state The immediate benefit obtained after taking action 'a' is used as the optimization objective for feedback. The following multi-objective reward function is designed: ;in, This indicates crop yield, in kg / mu (unit: 0.067 hectares). This indicates a waste of resources, such as excessive use of water and fertilizer. The environmental burden, pesticide residues, and soil pollution levels are expressed using quantitative indicators. Production costs are expressed in yuan per mu. The reward is for compliance with knowledge graphs and is a fixed value matched according to the recall ranking. The weights for each objective depend on the specific requirements and priorities of agricultural production. To achieve optimal decision-making, this invention employs the Q-learning algorithm as a reinforcement learning strategy. Let the current time be t, and its core update formula is: ,in, Indicates the state Next action The current value estimate (Q value), in units of expected cumulative reward; Represents the learning rate, ranging from 0 to 1. ≤1 controls the degree of influence of new information on updating the old Q value; Indicates at time Take action The reward obtained immediately afterwards is calculated based on the reward function described above; Represents the discount factor, in the range 0 ≤ <1, reflecting the importance of future rewards; This represents the maximum Q-value among possible future actions. This update formula iteratively updates the Q-value, allowing the model to gradually learn the optimal action strategy for each state.

[0086] The technical solution of this invention acquires modal perception data corresponding to a target open-field vegetable field, determines modal features corresponding to each modal perception data point, and determines feature fusion weights corresponding to the modal features at each time anchor point. Based on the feature fusion weights and the modal features, a global multimodal fusion feature is determined. Then, based on the modal perception data, knowledge representation features are determined from an agronomic knowledge graph. Finally, a target decision-making scheme is determined based on the knowledge representation features and the global multimodal fusion feature. Based on this technical solution, by integrating multiple data sources such as meteorology, soil, crop status, and agricultural machinery operating parameters, and employing spatiotemporal alignment and adaptive weighting mechanisms, combined with reinforcement learning techniques, agricultural machinery operation strategies can be dynamically adjusted based on real-time data. This optimizes agricultural management measures such as fertilization, irrigation, and pest and disease control, thereby improving production efficiency and ensuring the real-time nature and accuracy of decision-making.

[0087] Example 2

[0088] Figure 3 This is a flowchart illustrating a multimodal open-field vegetable management decision-making method provided by an embodiment of the present invention. This embodiment further optimizes the technical solution for determining the knowledge graph based on the above-mentioned technical solution. For example... Figure 3 As shown, it includes:

[0089] To address reasoning problems involving multiple knowledge domains in agricultural production, this paper constructs an agronomy-environment-machinery knowledge graph. This graph integrates knowledge from agricultural management, environmental control, and mechanized operations, thereby understanding the interrelationships between various elements and providing knowledge-based reasoning support for decision-making.

[0090] The knowledge graph construction framework consists of: An agronomic knowledge module, including agricultural management knowledge such as crop varieties, cultivation methods, fertilization and irrigation strategies, and pest and disease control; an environmental knowledge module, including meteorological conditions such as temperature, humidity, and precipitation, and soil conditions such as water, nutrients, and pH; and a mechanical knowledge module, including the technical parameters, operating efficiency, and operating modes of different agricultural machinery such as seeders and harvesters.

[0091] This knowledge graph enables knowledge reasoning based on existing data, automatically generating agronomic management plans, and providing decision support for subsequent agricultural production.

[0092] Knowledge Graph Construction: Knowledge Extraction: First, relevant knowledge is extracted from various sources such as agricultural literature, standard operating procedures, and sensor data using both manual and automated methods. Knowledge Representation: The extracted knowledge is transformed into nodes and edges in a knowledge graph. Nodes represent knowledge entities (crops, agronomic measures, meteorological factors, agricultural machinery, etc.), and edges represent the relationships between these entities ("fitness" relationships, "dependency" relationships, etc.). Knowledge Reasoning: Based on the knowledge graph, reasoning techniques, combined with environmental data and crop growth information, can be used to deduce management recommendations for the current stage of crop growth. For example, based on soil moisture and meteorological data, it can be deduced whether irrigation or fertilization is necessary.

[0093] This invention focuses on collaborative decision-making across all factors in open-field vegetable production. It constructs a collaborative decision-making system by introducing multimodal data perception and adaptive optimization techniques. Adaptive weighted Transformer and GCN technologies are introduced to deeply fuse data from different sources and at different granularities in open-field vegetable production, achieving spatiotemporal calibration and efficient fusion of multimodal data. By constructing an agronomic-environmental-mechanical knowledge graph, it achieves the integration and reasoning of knowledge from different domains, helping to understand the complex interactions between crops, the environment, and agricultural machinery. Through multi-task learning technology, it optimizes multiple tasks in agricultural production, ensuring efficient collaborative operation of agricultural machinery when performing multiple tasks. Through real-time interaction with the environment, it integrates reinforcement learning algorithms and knowledge graphs to automatically adjust agricultural machinery operation strategies, optimizing fertilization, irrigation, pesticide management, and other measures, ensuring continuous improvement in operational accuracy and resource utilization efficiency in changing production environments.

[0094] Example 3

[0095] Figure 4 This is a schematic diagram of a multimodal perception-based decision-making device for open-field vegetable management, provided as an embodiment of the present invention. Figure 4 As shown, the device includes: a data acquisition module 410, a feature fusion module 420, and a decision module 430; wherein,

[0096] The data acquisition module 410 is used to acquire modal sensing data corresponding to the target open-field vegetable field, wherein the modal sensing data includes segment sensing data, crop sensing data, agricultural machinery sensing data and agronomic operation parameters;

[0097] The feature fusion module 420 is used to determine the modal features corresponding to each modal sensing data, and to determine the feature fusion weights corresponding to the modal features at each time anchor point, and to determine the global multimodal fusion features based on the feature fusion weights and the modal features;

[0098] The decision module 430 is used to determine knowledge representation features from the agronomic knowledge graph based on the modal perception data, and to determine the target decision scheme based on the knowledge representation features and the global multimodal fusion features.

[0099] Based on the above technical solution, the feature fusion module is used to normalize each modal sensing data to obtain standardized data corresponding to each modal sensing data; and to extract temporal features from the standardized data based on the encoder to determine the modal features corresponding to each standardized data.

[0100] Based on the above technical solution, the feature fusion module is used to perform a linear transformation on each standardized data to generate a query matrix, a key matrix, and a numerical matrix corresponding to the standardized data; calculate the similarity matrix between the query matrix and the key matrix; and determine the modal features corresponding to the standardized data based on the similarity matrix and the numerical matrix.

[0101] Based on the above technical solution, the feature fusion module is used to align the modal features in terms of time dimension based on a preset time anchor set, wherein the preset time anchor set includes at least two time anchors; for the modal features after time dimension alignment, an evaluation score corresponding to the modal features of each time anchor is determined according to a modal quality scorer; a feature fusion weight corresponding to the modal features of each time anchor is determined according to the evaluation score; the modal features of each time anchor are weighted and fused according to the feature fusion weight to obtain a multimodal temporal representation; and a fused feature with the global multimodal representation is determined based on the multimodal temporal representation.

[0102] Based on the above technical solution, the feature fusion module is used to determine the anchor time length corresponding to the preset time anchor set and to determine the feature time length corresponding to the modal feature; when the feature time length is greater than the anchor time length, the modal feature is sampled by a sliding window according to the preset time anchor set; when the feature time length is less than the anchor time length, the modal feature is interpolated according to the preset time anchor set.

[0103] Based on the above technical solution, the feature fusion module is used to determine the evaluation score corresponding to each anchor feature corresponding to the modal feature; and to normalize the evaluation score corresponding to each anchor feature based on the normalized fusion function to determine the feature fusion weight corresponding to each anchor feature in the modal feature.

[0104] Based on the above technical solution, the feature fusion module is used to construct graph structure nodes based on the multimodal temporal representation, and calculate the similarity between modal features and the similarity between small nodes under the same modal feature according to cosine similarity; wherein, the graph structure nodes include global nodes and local small nodes; the global nodes are used to represent the global features of a single modality data; the local small nodes are used to represent the features of key data types under the current modality; the edge weights are determined based on the similarity, and an initial weighted graph is constructed based on the graph structure nodes and the edge weights, and the global multimodal fusion features are determined by feature aggregation based on the initial weighted graph.

[0105] Based on the above technical solution, the decision module is used to determine a state vector based on the modal perception data, determine a knowledge query template based on the state vector, obtain query results from the agronomic knowledge graph based on the knowledge query template, and determine the knowledge representation features based on the query results.

[0106] Based on the above technical solution, the decision module is used to concatenate the knowledge representation features and the global multimodal fusion features to determine the target state features; and to determine the target decision scheme based on the target state features and the preset reward function.

[0107] The technical solution of this invention acquires modal perception data corresponding to a target open-field vegetable field, determines modal features corresponding to each modal perception data point, and determines feature fusion weights corresponding to the modal features at each time anchor point. Based on the feature fusion weights and the modal features, a global multimodal fusion feature is determined. Then, based on the modal perception data, knowledge representation features are determined from an agronomic knowledge graph. Finally, a target decision-making scheme is determined based on the knowledge representation features and the global multimodal fusion feature. Based on this technical solution, by integrating multiple data sources such as meteorology, soil, crop status, and agricultural machinery operating parameters, and employing spatiotemporal alignment and adaptive weighting mechanisms, combined with reinforcement learning techniques, agricultural machinery operation strategies can be dynamically adjusted based on real-time data. This optimizes agricultural management measures such as fertilization, irrigation, and pest and disease control, thereby improving production efficiency and ensuring the real-time nature and accuracy of decision-making.

[0108] The open-field vegetable management decision-making device based on multimodal perception provided in this invention can execute the multimodal open-field vegetable management decision-making method provided in any embodiment of this invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0109] Example 4

[0110] Figure 5A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0111] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0112] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0113] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as multimodal open-field vegetable management decision-making methods.

[0114] In some embodiments, the multimodal open-field vegetable management decision-making method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the multimodal open-field vegetable management decision-making method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the multimodal open-field vegetable management decision-making method by any other suitable means (e.g., by means of firmware).

[0115] The various embodiments of the techniques described above and applied herein can be implemented in digital electronic circuits, integrated circuits, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable device including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from memory, at least one input device, and at least one output device, and transferring data and instructions to the memory, the at least one input device, and the at least one output device.

[0116] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with instruction execution, means or apparatus. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor means or apparatus, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0118] To provide interaction with a user, the techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The technologies described herein can be implemented in computing that includes backend components (e.g., as a data server), or middleware components (e.g., an application server), or frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with the embodiments of the technologies described herein), or any combination of such backend, middleware, or frontend components. The components can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0120] Computation can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A multimodal decision-making method for open-field vegetable management, characterized in that, include: Acquire modal sensing data corresponding to the target open-field vegetable field, wherein the modal sensing data includes segment sensing data, crop sensing data, agricultural machinery sensing data, and agronomic operation parameters; Determine the modal features corresponding to each modal sensing data, and determine the feature fusion weights corresponding to the modal features at each time anchor point. Based on the feature fusion weights and the modal features, determine the global multimodal fusion features. Based on the modal perception data, knowledge representation features are determined from the agronomic knowledge graph, and a target decision scheme is determined based on the knowledge representation features and the global multimodal fusion features. The process of determining the feature fusion weights corresponding to the modal features at each time anchor point, and determining global multimodal fusion features based on the feature fusion weights and the modal features, includes: The modal features are aligned in time dimension based on a preset set of time anchor points, wherein the preset set of time anchor points includes at least two time anchor points; For modal features aligned along the time dimension, an evaluation score corresponding to the modal features at each time anchor point is determined based on the modal quality scorer. Based on the evaluation scores, feature fusion weights corresponding to the modal features of each time anchor point are determined. Based on the feature fusion weights, the modal features of each time anchor point are weighted and fused to obtain a multimodal temporal representation. Based on the multimodal temporal representation, the fusion features with the global multimodal representation are determined; The determination of the fusion features based on the multimodal temporal representation and the global multimodal representation includes: A graph structure node is constructed based on the multimodal temporal representation, and the similarity between modal features and the similarity between small nodes under the same modal feature are calculated according to cosine similarity. The graph structure node includes global nodes and local small nodes. The global nodes are used to represent the global features of a single modality data. The local small nodes are used to represent the features of key data types under the current modality. Edge weights are determined based on the similarity, and an initial weighted graph is constructed based on the graph structure nodes and the edge weights. Feature aggregation is then performed based on the initial weighted graph to determine the global multimodal fusion features. The step of determining the target decision scheme based on the knowledge representation features and the global multimodal fusion features includes: The target state features are determined by concatenating the knowledge representation features and the global multimodal fusion features. The target decision scheme is determined based on the target state characteristics and the preset reward function.

2. The method according to claim 1, characterized in that, The temporal alignment of the modal features based on a preset set of time anchors includes: Determine the anchor time length corresponding to the preset time anchor set, and determine the feature time length corresponding to the modal feature; When the feature time length is greater than the anchor point time length, the modal feature is sampled using a sliding window based on the preset time anchor point set; If the feature time length is less than the anchor time length, the modal feature is interpolated according to the preset time anchor set.

3. The method according to claim 1, characterized in that, The step of determining the feature fusion weights corresponding to the features of each anchor point based on the evaluation scores includes: For the modal features, determine the evaluation scores corresponding to each anchor point feature corresponding to the modal features; The evaluation scores corresponding to each anchor feature are normalized based on the normalized fusion function to determine the feature fusion weights corresponding to each anchor feature in the modal features.

4. The method according to claim 1, characterized in that, The process of determining knowledge representation features from the agronomic knowledge graph based on the modality-aware data includes: A state vector is determined based on the modal perception data, and a knowledge query template is determined based on the state vector. Based on the knowledge query template, query results are obtained from the agronomic knowledge graph, and the knowledge representation features are determined based on the query results.

5. A decision-making device for open-field vegetable management based on multimodal perception, characterized in that, include: The data acquisition module is used to acquire modal sensing data corresponding to the target open-field vegetable field, wherein the modal sensing data includes segment sensing data, crop sensing data, agricultural machinery sensing data, and agronomic operation parameters; The feature fusion module is used to determine the modal features corresponding to each modal sensing data, and to determine the feature fusion weights corresponding to the modal features at each time anchor point, and to determine the global multimodal fusion features based on the feature fusion weights and the modal features; The decision module is used to determine knowledge representation features from the agronomic knowledge graph based on the modal perception data, and to determine the target decision scheme based on the knowledge representation features and the global multimodal fusion features. The feature fusion module is used to align the modal features in terms of time dimension based on a preset time anchor set, wherein the preset time anchor set includes at least two time anchors; for the time-dimensional aligned modal features, an evaluation score corresponding to the modal features of each time anchor is determined according to a modal quality scorer; a feature fusion weight corresponding to the modal features of each time anchor is determined according to the evaluation score; the modal features of each time anchor are weighted and fused according to the feature fusion weight to obtain a multimodal temporal representation; and a fused feature with the global multimodal representation is determined based on the multimodal temporal representation. The feature fusion module is used to construct graph structure nodes based on the multimodal temporal representation, and calculate the similarity between modal features and the similarity between small nodes under the same modal feature according to cosine similarity; wherein, the graph structure nodes include global nodes and local small nodes; the global nodes are used to represent the global features of a single modality data; the local small nodes are used to represent the features of key data types under the current modality; the edge weights are determined based on the similarity, and an initial weighted graph is constructed based on the graph structure nodes and the edge weights; feature aggregation is performed based on the initial weighted graph to determine the global multimodal fusion feature; The decision module is used to concatenate the knowledge representation features and the global multimodal fusion features to determine the target state features; and to determine the target decision scheme based on the target state features and the preset reward function.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal open-field vegetable management decision-making method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multimodal open-field vegetable management decision-making method according to any one of claims 1-4.