Method, device, storage medium and electronic equipment for agricultural autonomous decision making
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF COMPUTING TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]1、缺乏全流程集成能力:现有的技术大多聚焦于农业生产的某一特定环节,不同技术之间相互独立,数据难以互通,无法实现从“耕、种、管、收”全生产过程的有机统一与闭环管控
[0222](1)决策精度实现代际跳跃:通过动态最优矩阵算法,本系统在实际大田测试中的决策准确率达到89.7%,较传统单一通用大模型提升了36.7%。
Smart Images

Figure CN122525906A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated technology of smart agriculture and generative artificial intelligence, and in particular to a method, apparatus, storage medium and electronic device for autonomous agricultural decision-making. Background Technology
[0002] The people are the foundation of the nation, and grain is the lifeblood of the people. Food security is fundamental to human survival and development. With the increasing challenges of global population growth and climate change, developing smart agriculture has become the only way to ensure global food security and achieve sustainable agricultural development.
[0003] However, in long-term agricultural production practices, agricultural operation decisions have primarily relied on human experience. Due to limitations in individual knowledge, accumulated experience, and information asymmetry, human experience-based decisions often exhibit lag and bias, making it difficult to cope with complex and ever-changing natural environments. This imprecise decision-making model not only limits further yield increases but also leads to excessive inputs such as "excessive use of fertilizers and pesticides," resulting in serious resource waste and environmental pressure.
[0004] Currently, although some single-point smart agriculture technologies have emerged, such as localized water and fertilizer management systems, remote sensing monitoring technology, and pest and disease identification models, the following significant problems still exist in practical applications:
[0005] 1. Lack of end-to-end integration capabilities: Most existing technologies focus on a specific link in agricultural production. Different technologies are independent of each other, and data is difficult to share, making it impossible to achieve organic unity and closed-loop management of the entire production process from "cultivation, planting, management, and harvesting".
[0006] 2. Low degree of autonomous decision-making: Most existing systems are still in the decision-making support stage, that is, they only provide data analysis results, and the final agricultural instructions still require human intervention to make judgments. They lack the ability to make comprehensive autonomous reasoning and dynamic decisions based on real-time environmental changes.
[0007] 3. Difficulty in fusion of multi-source data: Agricultural production involves massive amounts of heterogeneous data from multiple sources, such as meteorology, soil, crop phenotypes, and agricultural machinery status. Traditional models often show insufficient robustness when processing this type of "sky-air-ground-human-machine" multimodal data, making it difficult to achieve deep semantic alignment and comprehensive analysis.
[0008] 4. Insufficient precision in addressing specific ecological problems: For example, in the black soil region of China, there are serious problems of soil thinning, becoming less fertile, and hardening. Existing general agricultural technologies often lack deeply coupled decision-making solutions for such special protection and utilization scenarios.
[0009] Based on the above issues, the existing smart agriculture system urgently needs an intelligent decision-making center that can deeply integrate "perception-cognition-decision-execution", has high robustness, and can cover the entire agricultural life cycle.
[0010] Especially when dealing with complex and ever-changing field environments, a system is urgently needed that can:
[0011] 1. Break down data silos: Semantically align multi-source heterogeneous data from "sky-air-ground-human-machine" that span time, space, and dimensions to form a deep understanding of farmland conditions.
[0012] 2. Eliminate model illusion: Introduce professional agricultural knowledge graphs to enhance logic and ensure that the decision-making instructions generated by AI conform to the rigorous scientific logic of agricultural production.
[0013] 3. Achieve dynamic evolution: It has the ability to coordinate multiple agents and can dynamically schedule the optimal algorithm resources according to different agricultural tasks (such as precision fertilization, pest and disease early warning, yield prediction, etc.) to achieve self-game and improvement of decision accuracy.
[0014] Therefore, developing an agricultural brain system based on a large artificial intelligence model that can make autonomous decisions throughout the entire process is of great practical significance and technical value for improving agricultural production efficiency, ensuring national food security, and protecting key natural resources such as black soil.
[0015] In conclusion, the existing technology obviously has inconveniences and defects in practical use, so it is necessary to improve it. Summary of the Invention
[0016] To address the aforementioned shortcomings, the present invention aims to provide a method, apparatus, storage medium, and electronic device for autonomous agricultural decision-making, which possesses the capability for autonomous and precise decision-making throughout the entire agricultural lifecycle and closed-loop control of the entire process, thereby significantly improving economic benefits.
[0017] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0018] In a first aspect, embodiments of the present invention provide a method for autonomous agricultural decision-making, comprising:
[0019] The data acquisition process involves collecting multi-source, heterogeneous data related to agricultural production.
[0020] The semantic alignment step involves semantically aligning the multi-source heterogeneous data within a shared embedding space.
[0021] The task decomposition and optimization step involves the brain's decision layer decomposing agricultural tasks into at least one sub-task based on intent recognition, and using a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and aggregating the reasoning results of each sub-task to generate autonomous decision instructions.
[0022] The equipment management and control steps involve outputting the autonomous decision-making instructions to the intelligent agricultural machinery equipment for closed-loop management and control.
[0023] According to the aforementioned method for autonomous agricultural decision-making, the data collection step further includes:
[0024] The multi-source heterogeneous data related to agricultural production can be acquired in real time or semi-real time through satellite remote sensing, UAV aerial surveying, sensor feedback, agricultural machinery operation and / or manual field inspection feedback.
[0025] According to the aforementioned agricultural autonomous decision-making method, the semantic alignment step further includes:
[0026] The multimodal rotational position embedding method is used to encode and align the multimodal heterogeneous data, including images, text, audio, and video, in the temporal and spatial dimensions. Pre-defined visual encoders, text encoders, audio encoders, and video encoders are used to extract features from the data of different modalities, so as to minimize the distance of semantically similar content in the shared embedding space.
[0027] According to the aforementioned agricultural autonomous decision-making method, the multi-source heterogeneous data includes: image sequences. Text sequence audio sequence Video sequence ;
[0028] The semantic alignment step further includes: performing semantic alignment of the multi-source heterogeneous data in a shared embedding space;
[0029] Its goal is to learn four mapping functions, including: image mapping function Text mapping function Audio mapping function Video mapping function This makes the shared embedded space Minimize the distance between semantically similar content from different modalities:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] Formula 9;
[0035] The model aligns inputs from different modalities into the same semantic space;
[0036] Feature extraction is performed using a dedicated encoder for each modality:
[0037] For the image modality, image I is processed by a predetermined visual encoder, which segments image I into image blocks and converts them into embedded representations, as shown in Equation 10. It is the embedding of the image patch;
[0038] Formula 10;
[0039] For the text modality, the text T is processed by a predetermined text encoder and encoded as shown in Formula 11, where ti′∈Rd is the embedding of the word in the text T;
[0040] Formula 11;
[0041] For the audio modality, audio A is processed by a predetermined audio encoder, as shown in Formula 12, to extract the audio features of audio A and convert them into time series embeddings;
[0042] Formula 12;
[0043] For the video modality, the video V is processed by a predetermined video encoder, as shown in Equation 13, treating each frame of the video V as an image and adding an embedding in the time dimension, where vit is the visual block embedding in each frame t.
[0044] Formula 13;
[0045] A multimodal rotational position embedding method is adopted. In the 1D position embedding of the text T, as shown in Formula 14, 1D rotational position embedding is used to capture word order information, where i is the position of the word in the text sequence, j is the embedding dimension index, and d is the embedding dimension.
[0046] Formula 14;
[0047] In the 2D position embedding of image I, the two-dimensional spatial position (h, w) of the image is rotated and encoded as shown in Formula 15, where h and w represent the positions of words in the image, and d h and dw For the corresponding embedding dimension;
[0048] Formula 15;
[0049] In the 3D position embedding of the video V, as shown in Formula 16, corresponding position information is generated for each video word in both time and space;
[0050] Formula 16;
[0051] The audio A is mainly time series data, and 1D rotation embedding is used to represent time information, as shown in Formula 17, where t represents the time step of each feature in the audio signal and j is the embedding dimension.
[0052] Formula 17.
[0053] According to the aforementioned agricultural autonomous decision-making method, the brain decision-making layer adopts a multi-agent collaborative architecture, including a management agent and at least one executive agent; the model combination is selected from a hybrid multi-model architecture, including a general large model, an intent recognition model, and multiple agricultural-specific models;
[0054] The task decomposition and optimization steps further include:
[0055] The task decomposition sub-step involves the management agent decomposing the agricultural task into at least one sub-task using the intent recognition model, with different sub-tasks being executed by different execution agents.
[0056] The model matching sub-step employs the dynamic optimal matrix algorithm to perform reasoning for each model combination dynamically matched and optimized by the executing agent based on the real-time characteristics of each sub-task, and completes decision reasoning through multi-agent game collaboration.
[0057] According to the aforementioned agricultural autonomous decision-making method, the task decomposition sub-step further includes:
[0058] The management agent performs task segmentation analysis on the agricultural task I using an intent recognition model. The output intent and task segmentation are shown in Formula 1. Simultaneously, an optimal allocation matrix is designed within the hybrid multi-model architecture for each subtask S. j For each ∈P, an optimal allocation matrix M is established, as shown in Formula 2;
[0059] (1)
[0060] Formula 2;
[0061] Where P represents the task segmentation and type, Mi Let S represent the i-th model. j This represents the j-th task;
[0062] The model matching sub-step further includes:
[0063] For each of the subtasks S j Use the function Select Model(S) j The optimal model is dynamically selected using K models M1, M2, ..., Mn. K Each model M i For the subtask S j An adaptation score is calculated by evaluating the model's historical performance, current environment, and / or task characteristics; therefore, for each model M... i and each of the subtasks S j Define an fitness score A ij As shown in Formula 3, f is an evaluation function used to measure the model M. i For the subtask S j Adaptability;
[0064] Formula 3;
[0065] The model M i The selection objective is to select the fitness score A. ij The highest model is shown in Equation 4;
[0066] Formula 4;
[0067] The design constraints are shown in Equation 5, and the model M... i The resource consumption is R i R max This is the maximum allowed resource consumption, ensuring that the model M... i Choose to proceed within the acceptable resource limits;
[0068] Formula 5;
[0069] The dynamic model selection process of the hybrid multi-model architecture is shown in Equation 6, where θ is a predetermined threshold used to determine the model M. i The lowest fitness score if the selected model M i If the score is below the threshold, the default model is returned, and M is in this case. selected (S) j () is the optimal model selected based on the allocation matrix;
[0070] Formula 6;
[0071] During the reasoning process, for each selected model M selected (S) j As shown in Formula 7, perform reasoning and output the reasoning result;
[0072] Formula 7;
[0073] The reasoning results of each of the sub-tasks are aggregated to generate autonomous decision-making instructions, as shown in Formula 8;
[0074] Formula 8.
[0075] According to the aforementioned agricultural autonomous decision-making method, the task decomposition and optimization step further includes:
[0076] In the knowledge enhancement step, based on the forced response mechanism of the general large model, the brain decision layer compares the reasoning results and / or the autonomous decision instructions with the professional knowledge patterns in the predetermined agricultural knowledge graph, and the brain decision layer forcibly corrects the autonomous decision instructions according to the feedback.
[0077] Secondly, embodiments of the present invention provide an agricultural autonomous decision-making device constructed based on any one of the methods described above, the device comprising:
[0078] The data acquisition module is used to collect multi-source heterogeneous data related to agricultural production;
[0079] A semantic alignment module is used to perform semantic alignment of the multi-source heterogeneous data in a shared embedding space.
[0080] The task decomposition and optimization module is used to decompose agricultural tasks into at least one sub-task based on intent recognition by the brain decision layer, and to use a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and to aggregate the reasoning results of each sub-task to generate autonomous decision instructions.
[0081] The equipment management module is used to output the autonomous decision-making instructions to the intelligent agricultural machinery equipment for closed-loop management.
[0082] Thirdly, embodiments of the present invention provide a storage medium for storing a computer program for performing any of the methods described herein.
[0083] Fourthly, embodiments of the present invention provide an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement any of the methods described above.
[0084] Therefore, this invention is proposed. Attached Figure Description
[0085] Figure 1 This is a flowchart illustrating the agricultural autonomous decision-making method provided in Embodiment 1 of the present invention;
[0086] Figure 2 This is an architecture diagram of the four-level artificial intelligence decision-making system provided in Embodiment 2 of the present invention;
[0087] Figure 3 This is a schematic diagram of the brain decision-making layer provided in Embodiment 3 of the present invention;
[0088] Figure 4 This is a schematic diagram of the full-modal alignment training provided in Embodiment 4 of the present invention;
[0089] Figure 5 This is a schematic diagram of the agricultural autonomous decision-making device provided in Embodiment 5 of the present invention;
[0090] Figure 6 This is a schematic diagram of the structure of the electronic device provided in Embodiment Six of the present invention. Detailed Implementation
[0091] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0092] It should be noted that references to "an embodiment," "embodiment," "example embodiment," etc., in this specification refer to the described embodiment including specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. Furthermore, such expressions do not refer to the same embodiment. Moreover, when describing specific features, structures, or characteristics in conjunction with embodiments, whether or not explicitly described, it is indicated that incorporating such features, structures, or characteristics into other embodiments is within the knowledge of those skilled in the art.
[0093] Furthermore, certain terms are used in the specification and subsequent claims to refer to specific components or parts. Those skilled in the art will understand that manufacturers may use different names or terms to refer to the same component or part. This specification and subsequent claims do not distinguish components or parts by differences in name, but rather by differences in function. The terms "comprising" and "including" used throughout the specification and subsequent claims are open-ended and should be interpreted as "including but not limited to." Additionally, the term "connection" here includes any direct and indirect electrical connection means. Indirect electrical connection means include connections made through other means.
[0094] The method for autonomous agricultural decision-making provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0095] This invention relates to an agricultural intelligent decision-making method based on agricultural IoT, driven by a general large-scale model, and possessing autonomous and precise decision-making capabilities throughout the entire lifecycle. Specifically, it covers the following sub-technical fields:
[0096] 1. Cross-modal agricultural big data integration and sensing technology: This involves a five-in-one full-space scale data acquisition system integrating "sky-air-ground-human-machine" to monitor multi-source heterogeneous data such as crop remote sensing, UAV aerial survey, sensor feedback, and agricultural machinery operations in real time and semi-real time.
[0097] 2. Multi-Agent Collaboration Architecture in Agriculture: This involves a task decomposition and execution mechanism built in the brain's decision-making layer. Through the dynamic interaction between the management agent and multiple execution agents, it enables automated analysis and response to complex agricultural logic.
[0098] 3. Hybrid Multi-Model Dynamic Optimization Decision Algorithm: This involves a dynamic optimal matrix algorithm, the core of which is to match the optimal model combination in real time in a hybrid model architecture (including a general large model and multiple types of agricultural-specific models) based on the sub-task characteristics of the intention recognition model decomposition, so as to achieve high-precision decision output.
[0099] 4. Alignment and Fusion Training of Multimodal Agricultural Data: This involves a multimodal alignment training method that utilizes multimodal rotational position embedding technology to semantically align heterogeneous data from multiple sources, such as images, text, audio, and video, in a shared embedding space, thereby enhancing the system's robustness in complex field environments.
[0100] 5. End-to-end autonomous agricultural management and equipment linkage: This involves a closed-loop management technology that transforms digital decision-making instructions into physical operation actions. Through real-time linkage with intelligent agricultural machinery equipment (such as pure electric unmanned agricultural machinery), it enables variable precision operation covering the entire process of "plowing, planting, management, and harvesting".
[0101] This invention aims to overcome the shortcomings of existing agricultural decision-making systems, such as data silos, model illusions, broken decision-making links, and a lack of full lifecycle autonomy. It proposes an intelligent agricultural decision-making method based on a large model and fully autonomous decision-making throughout the entire process. This invention solves core challenges such as deep semantic alignment of multi-source heterogeneous data, logical decomposition of complex agricultural tasks, and dynamic optimal resource scheduling by constructing a multi-layered coupled architecture.
[0102] Figure 1 This is a flowchart illustrating the agricultural autonomous decision-making method provided in Embodiment 1 of the present invention, the method comprising:
[0103] Step S101, data acquisition step, collecting multi-source heterogeneous data related to agricultural production.
[0104] Preferably, the data acquisition step further includes:
[0105] Multi-source heterogeneous data related to agricultural production can be acquired in real time or semi-real time through satellite remote sensing, UAV aerial surveying, sensor feedback, agricultural machinery operations and / or manual field inspection feedback.
[0106] Step S102, semantic alignment step, performs semantic alignment of multi-source heterogeneous data in a shared embedding space.
[0107] Preferably, the semantic alignment step further includes:
[0108] By using a multimodal rotational position embedding method, positional encoding and alignment are performed on multi-source heterogeneous data including images, text, audio, and video in the temporal and spatial dimensions. Pre-defined visual encoders, text encoders, audio encoders, and video encoders are used to extract features from data of different modalities, so as to minimize the distance of semantically similar content in a shared embedding space.
[0109] Preferably, the multi-source heterogeneous data includes: image sequences. Text sequence audio sequence Video sequence .
[0110] Preferably, the semantic alignment step further includes: performing semantic alignment of multi-source heterogeneous data in a shared embedding space.
[0111] Its goal is to learn four mapping functions, including: image mapping function Text mapping function Audio mapping function Video mapping function This enables shared embedded spaces Minimize the distance between semantically similar content from different modalities:
[0112] .
[0113] .
[0114] .
[0115] .
[0116] Formula 9.
[0117] The model aligns inputs from different modalities into the same semantic space.
[0118] Feature extraction is performed using a dedicated encoder for each modality:
[0119] For the image modality, image I is processed by a predefined visual encoder, which segments image I into image patches and converts them into embedded representations, as shown in Equation 10. It is the embedding of image patches.
[0120] Formula 10.
[0121] For the text modality, text T is processed by a predefined text encoder and encoded as shown in Equation 11, where ti′∈Rd is the embedding of the word in text T.
[0122] Formula 11.
[0123] For the audio modality, audio A is processed by a predefined audio encoder, as shown in Equation 12, to extract the audio features of audio A and convert them into time-series embeddings.
[0124] Formula 12.
[0125] For the video modality, the video V is processed by a predefined video encoder, as shown in Equation 13. Each frame of the video V is treated as an image, and an embedding is added in the time dimension, where vit is the visual block embedding in each frame t.
[0126] Formula 13.
[0127] A multimodal rotational position embedding method is adopted. In the 1D position embedding of text T, as shown in Equation 14, 1D rotational position embedding is used to capture word order information, where i is the position of the word in the text sequence, j is the embedding dimension index, and d is the embedding dimension.
[0128] Formula 14.
[0129] In the 2D position embedding of image I, the two-dimensional spatial position (h, w) of the image is rotated and encoded as shown in Equation 15, where h and w represent the positions of the words in the image, and d h and d w For the corresponding embedding dimension.
[0130] Formula 15.
[0131] In the 3D position embedding of video V, as shown in Equation 16, corresponding position information is generated for each video lexical unit in both time and space.
[0132] Formula 16.
[0133] Audio A is mainly time series data, and 1D rotation embedding is used to represent time information, as shown in Equation 17, where t represents the time step of each feature in the audio signal and j is the embedding dimension.
[0134] Formula 17.
[0135] Step S103, Task Decomposition and Optimization Step: The brain decision layer decomposes the agricultural task into at least one sub-task based on intent recognition, and uses a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and aggregates the reasoning results of each sub-task to generate autonomous decision instructions.
[0136] Preferably, the task decomposition and optimization step further includes:
[0137] The task decomposition sub-step involves the management agent breaking down agricultural tasks into at least one sub-task using an intent recognition model. Different sub-tasks are executed by different execution agents.
[0138] The model matching sub-step employs a dynamic optimal matrix algorithm. Based on the real-time characteristics of each sub-task, it performs reasoning on the dynamically matched and optimized model combination for each executing agent, and completes decision reasoning through multi-agent game collaboration.
[0139] Preferably, the task decomposition sub-step further includes:
[0140] The management agent performs task segmentation analysis on agricultural task I using an intent recognition model. The output intent and task segmentation are shown in Equation 1. Simultaneously, an optimal allocation matrix is designed within the hybrid multi-model architecture for each subtask S. j For each ∈P, an optimal allocation matrix M is established, as shown in Formula 2.
[0141] (1)
[0142] Formula 2.
[0143] Where P represents the task segmentation and type, M i Let S represent the i-th model. j This represents the j-th task.
[0144] Preferably, the model matching sub-step further includes:
[0145] For each subtask S j Use the function Select Model(S) j The optimal model is dynamically selected using K models M1, M2, ..., Mn. K Each model M i For subtask S j An adaptive score is calculated by evaluating the model's historical performance, current environment, and / or task characteristics; therefore, for each model M... i and each subtask S j Define an fitness score A ij As shown in Equation 3, f is an evaluation function used to measure model M. i For subtask S j Adaptability.
[0146] Formula 3.
[0147] Model M i The selection objective is to select the fitness score A. ij The highest model is shown in Equation 4.
[0148] Formula 4.
[0149] Design constraints are shown in Equation 5, Model M i The resource consumption is R i R max It is the maximum allowed resource consumption, ensuring that model M i Choose to proceed within the acceptable resource limits.
[0150] Formula 5.
[0151] Preferably, the dynamic model selection process of the hybrid multi-model architecture is as shown in Formula 6, where θ is a predetermined threshold used to determine model M. i The lowest fitness score if the selected model M i If the score is below the threshold, the default model is returned, and M is in this case. selected (S) j () is the optimal model selected based on the allocation matrix.
[0152] Formula 6.
[0153] During the reasoning process, for each selected model M selected (S) j As shown in Formula 7, perform reasoning and output the reasoning results.
[0154] Formula 7.
[0155] The reasoning results of each subtask are aggregated to generate autonomous decision-making instructions, as shown in Formula 8.
[0156] Formula 8.
[0157] Step S104, Equipment Control Step, outputs autonomous decision-making instructions to intelligent agricultural machinery equipment for closed-loop control.
[0158] Preferably, the brain decision-making layer adopts a multi-agent collaborative architecture, including a management agent and at least one executive agent. The model combination is selected from a hybrid multi-model architecture, including a general large model, an intent recognition model, and multiple agriculture-specific models.
[0159] Preferably, after the task decomposition and optimization step and before the equipment management and control step, the following steps are further included:
[0160] The knowledge enhancement step, based on the mandatory response mechanism of a general large model, involves the brain's decision-making layer comparing the reasoning results and / or autonomous decision-making instructions with the professional knowledge patterns in the predetermined agricultural knowledge graph. The brain's decision-making layer then forcibly corrects the autonomous decision-making instructions based on the feedback.
[0161] The key technical points of this invention include:
[0162] Key Point 1: Four-Tier AI Decision-Making System Architecture. This invention constructs a highly collaborative four-tier architecture from the bottom up. Figure 2 This is an architecture diagram of the four-level artificial intelligence decision-making system provided in Embodiment 2 of the present invention.
[0163] 1. Comprehensive Data Acquisition Layer: Construct a five-in-one perception system encompassing "space, air, ground, human, and machine." Acquire real-time / semi-real-time heterogeneous data streams covering the entire spatial scale through satellite remote sensing, UAV hyperspectral imaging, ground sensor arrays, and agricultural machinery operation feedback.
[0164] 2. Small Algorithm Model Enhancement Layer: Integrates 101 professional and lightweight agricultural-specific models in 14 categories, including image phenotypes, crop pests, and soil nutrients, which are responsible for feature extraction and preliminary quantitative analysis of the underlying raw data.
[0165] 3. Knowledge Enhancement Layer Based on Agricultural Knowledge Graph: The agricultural knowledge graph is self-expanded through a multi-round data collection strategy. Structured knowledge is extracted using the forced response mechanism of the Large Language Model (LLM) and injected into the decision-making process to eliminate the logical illusion of the general large model.
[0166] 4. Brain Decision Layer: The core central hub of the system. Based on a Hybrid Multi-Model architecture, it achieves full automation from task perception to policy generation through the game and cooperation between the management agent (C-Agent) and multiple execution agents.
[0167] Key Point 2: Core Algorithm: Dynamic Optimal Matrix Algorithm (DOMA).
[0168] To achieve accurate decomposition and model matching of complex agricultural tasks, this invention proposes a dynamic optimal matrix algorithm.
[0169] Figure 3 This is a schematic diagram of the brain decision-making layer provided in Embodiment 3 of the present invention. In the brain decision-making layer, a hybrid multi-model architecture and a multi-agent collaborative architecture are designed. The large model base adopts a hybrid multi-model architecture, including a full-modality general large model and an agricultural-specific model. It can dynamically select the most suitable model for inference based on the characteristics of the task, combining various types of models to achieve more comprehensive and accurate functions, including but not limited to: an intent recognition model, a general large model, a pest and disease identification model, a planting pattern model, a weather forecasting model, a seed recommendation model, and a crop phenotypic analysis model. The intent recognition model analyzes the tasks input to the large model base, realizing task segmentation. Let the input be I, then the output intent and task segmentation are as shown in Formula 1. Simultaneously, an optimal allocation matrix is designed in the hybrid multi-model architecture for each subtask S. j For each ∈P, an optimal allocation matrix M is established, as shown in Formula 2.
[0170] Formula 2
[0171] Formula 2
[0172] Where P represents the task segment and type. M i Let S represent the i-th model. j This represents the j-th task.
[0173] In dynamic model selection, the total number of subtasks is N, and each subtask can be assigned to at least one model. For each subtask, this invention uses the function Select Model(S) j The optimal model is dynamically selected using K models M1, M2, ..., Mn. K Each model receives an adaptive score for a subtask. This score is calculated by evaluating the model's historical performance, current environment, and task characteristics; therefore, for each model M... i and each subtask S j Define an fitness score A ij The calculation method is shown in Formula 3, where f is an evaluation function used to measure model M. i For subtask S j The adaptability of this function is assessed by considering the characteristics of the model, the relevance of the training data, and the requirements of the current environment.
[0174] Formula 3
[0175] In the model selection process, the goal of this invention is to select the model with the highest fitness score, as shown in Equation 4. This means that a model M will be found. i* This makes it possible for subtask S j Adaptability score A ij maximum.
[0176] Formula 4
[0177] Meanwhile, this invention designs constraints, as shown in Equation 5, where the resource consumption of the model is R. i R max This represents the maximum allowed resource consumption, ensuring that model selection is performed within an acceptable resource range.
[0178] Formula 5
[0179] Based on the above conditions, the dynamic model selection process of the hybrid multi-model architecture is shown in Equation 6, where θ is a pre-defined threshold used to determine the model's minimum fitness score. If the selected model's score is lower than the threshold, the default model is returned, and M... selected (S) j() is the optimal model selected based on the allocation matrix.
[0180] (6)
[0181] During the reasoning process, for each selected model M selected (S) j The reasoning output is shown in Formula 7. Finally, Formula 8 combines the reasoning output of all subtasks into a single overall result.
[0182] Formula 7
[0183] Formula 8
[0184] Key Point 3: Full Modality Alignment Training Method
[0185] This invention proposes a full-modal alignment training method to improve the fusion performance of the proposed decision model on multi-source heterogeneous data. By learning a mapping method from different modalities to a shared embedding space, it minimizes the distance between semantically similar content in that space. Figure 4 This is a schematic diagram of the full-modal alignment training provided in Embodiment 4 of the present invention.
[0186] The inputs of this invention include: image sequences Text sequence audio sequence Video sequence The goal is to learn four mapping functions, image mapping. Text mapping Audio mapping Video mapping This enables shared embedded spaces Minimize the distance between semantically similar content from different modalities:
[0187]
[0188]
[0189]
[0190]
[0191] Formula 9
[0192] In this way, the model can align inputs from different modalities to the same semantic space. Simultaneously, a dedicated encoder is used for feature extraction in each modality. In the visual encoder, a pre-trained VisionTransformer is used to process the image, as shown in Equation 10, segmenting image I into patches and transforming it into an embedding representation. It is an embedding of image patches.
[0193] Formula 10
[0194] In the text encoder, as shown in Equation 11, the text T is encoded using LLama3 (an open-source large language model developed by Meta), where ti′∈Rd is the embedding of the text token.
[0195] Formula 11
[0196] In the audio encoder, as shown in Equation 12, Whisper-large-v3 is introduced to extract audio features and convert audio A into a time-series embedding.
[0197] Formula 12
[0198] In the video encoder, as shown in Equation 13, each frame of the video is treated as an image, a visual encoder is used, and an embedding is added in the time dimension, where vit is the visual patch embedding in each frame t.
[0199] Formula 13
[0200] M-RoPE can align inputs of different modalities in time and space dimensions. This invention adopts the multimodal rotational position embedding M-RoPE method. In the 1D position embedding of text, as shown in Equation 14, 1D rotational position embedding is used to capture word order information, where i is the position of the word in the text sequence, j is the embedding dimension index, and d is the embedding dimension.
[0201] Formula 14
[0202] In 2D image embedding, the two-dimensional spatial position (h, w) of the image is rotated and encoded, as shown in Equation 15, where h and w represent the position of the token in the image, and d h and d w For the corresponding embedding dimension.
[0203] Formula 15
[0204] In the 3D location embedding of video, the token of each frame of the video not only contains two-dimensional spatial location information, but also needs to be embedded in the time dimension, as shown in formula (16). In this way, unique location information can be generated for each video token in both time and space.
[0205] (16)
[0206] Audio data is mainly time series data, so 1D rotation embedding is used to represent time information, as shown in Equation 17, where t represents the time step of each feature in the audio signal and j is the embedding dimension.
[0207] Formula 17
[0208] (III) Experimental Results
[0209] To further verify the performance of the proposed decision-making model in the field of agricultural decision-making, this invention was compared with GPT-4 on multiple agricultural benchmark datasets. The results are as follows: Figure 4 As shown in the figure, these datasets cover the entire life cycle of crops, agricultural operations, logical reasoning, pest and disease management, and land protection. In 15 tasks, the method proposed in this invention outperforms GPT-4 in 10 tasks, and performs very close to it in the remaining 5. The method proposed in this invention shows significant advantages in tasks such as fertilization, planting, land preparation, irrigation, weeding, deep loosening, topdressing, black soil protection, yield estimation, and disease control. Although GPT-4 performs strongly in logical reasoning tasks, the method proposed in this invention also performs excellently in this area, scoring only one point lower than GPT-4. Figure 4 Scoring of agricultural benchmark datasets.
[0210] A comparison of the brain's decision-making process with other models.
[0211] Given that the proposed method is a composite architecture composed of multiple models, it can dynamically select the optimal model combination during each inference process to achieve collaborative task execution. Table 1 shows the score comparison between the proposed method and five current mainstream models, covering generation quality score, response accuracy score, context understanding score, and overall score. In this invention, we compared GPT-4, Claude3.5, LLaMA70B, Qwen, ChatGLM, and the proposed method. The results show that the proposed method outperforms other single models in all evaluation metrics, with its overall score being 36.7 percentage points higher than ChatGLM. This achievement significantly demonstrates its advancement and superiority in the field of intelligent agricultural decision-making.
[0212]
[0213] (2) Comparison of the proposed full-modal alignment method with ordinary methods
[0214] To evaluate the performance improvement of the proposed full-modal alignment method compared to traditional methods, this section conducts extensive experiments on multimodal benchmark datasets, providing an in-depth analysis of the model's performance in various modal tasks, including image, video, and audio. In the experiments, the proposed method was compared with current leading open-source state-of-the-art models, GPT-4, and other mainstream models. The results show that the proposed method exhibits significant advantages in multiple tasks. Specifically, Table 2 demonstrates the superior performance of the proposed method on image understanding benchmarks such as MMNIST, OCRbench, and VCR, significantly outperforming models like GPT-4 and Claude-3.5. Particularly on OCRbench, the proposed method achieved a score of 879, a significant improvement over GPT-4's 736, highlighting its superior performance in text extraction and recognition. In video understanding tasks, as shown in Table 3, the method proposed in this invention achieves a score of 74.8 on the Video-MME benchmark, surpassing GPT-4's 72.0 score. Furthermore, it outperforms models such as Gemini and GPT-4 in multiple video tasks, establishing its leading position in the field of video modal processing. For audio understanding tasks, as shown in Table 4, the method proposed in this invention achieves a score of 5.73 on the SpeechSound benchmark, surpassing Gemini-1.5-pro's 5.27 score, demonstrating its leading capabilities in audio recognition and sound understanding. In summary, the leading advantages demonstrated by the method proposed in this invention in complex multimodal understanding and cross-modal tasks fully validate its advanced nature in full-modal alignment and dynamic data interaction technologies.
[0215] Table 2 Image Evaluation Scores
[0216]
[0217] Table 3 Video Evaluation Scores
[0218]
[0219] Table 4 Audio Ability Assessment Scores
[0220]
[0221] (III) Technical Effects
[0222] (1) Decision accuracy achieves generational leap: Through the dynamic optimal matrix algorithm, the decision accuracy of this system in actual field test reaches 89.7%, which is 36.7% higher than the traditional single general large model.
[0223] (2) Closed-loop control throughout the entire process: For the first time, autonomous decision-making throughout the entire life cycle has been realized, from pre-production planning (variety selection, sowing strategy) to production management (precise water and fertilizer, pesticide prescription) to post-production yield estimation, without the need for high-frequency human intervention.
[0224] (3) Significantly improved economic benefits: Experiments have shown that in Hulunbuir City, 1,367 mu of experimental fields fully adopted the decision-making suggestions of the method proposed in this invention and implemented them in conjunction with intelligent agricultural machinery equipment for unmanned variable operations. Compared with traditional planting methods, the experimental fields achieved an average cost saving of 39.405 yuan / mu, an efficiency increase of 6 yuan / mu, and a yield increase of 58.55 yuan / mu, for a total cost saving, efficiency increase, and yield increase of 103.955 yuan / mu. Detailed Implementation
[0225] (I) System Implementation Environment and Sensing Network Construction.
[0226] This embodiment uses a large-scale farm in the Hulunbuir region of China as the application scenario for the system of this invention. During implementation, the system first constructs a five-in-one perception network integrating "sky-air-ground-human-machine" through a comprehensive data acquisition layer. Satellite remote sensing periodically acquires macroscopic NDVI (NDVI, vegetation index), drones equipped with hyperspectral cameras perform centimeter-level resolution field inspections, and ground sensor arrays monitor soil temperature, humidity, and nutrient content in real time. This is combined with real-time load feedback during agricultural machinery operations and unstructured notes generated by manual field inspections. These multi-source, heterogeneous raw data streams from different sources and with varying frequencies are aggregated in real time to the system backend, forming a massive agricultural native database as the foundation for the system's autonomous processing.
[0227] (ii) The active scheduling mechanism of the brain center for quantization algorithms.
[0228] The core feature of this invention lies in the fact that the 101 specialized algorithms in the small algorithm model enhancement layer are not part of a pre-defined, fixed pipeline, but rather exist as a "toolbox" for the brain's decision-making layer. During implementation, upon receiving raw perceptual data or vague user instructions, the brain's decision-making layer first utilizes its intent recognition and logical reasoning functions to autonomously determine the feature dimensions required to complete the task. Subsequently, the brain's decision-making layer issues a call command, actively retrieving and activating the corresponding quantification algorithm from the toolbox. For example, when the brain "captures" abnormal fluctuations in the color of a plot of land in a massive amount of raw satellite imagery, it autonomously determines, based on the current agricultural season, that this could be due to pests or nutrient deficiencies, and then proactively and concurrently calls the "pest and disease visual feature extraction algorithm" and the "leaf nutrient spectral analysis model" for parallel verification. This "on-demand calling, autonomous quantification" mechanism transforms the system from a traditional data-driven approach to a task-goal-driven one.
[0229] (iii) Knowledge graph-driven logic verification and mandatory response.
[0230] After the algorithm is actively invoked and returns quantified features (such as chlorophyll content and soil nitrogen concentration), the brain's decision-making layer generates a decision, which is then placed in a knowledge enhancement layer based on an agricultural knowledge graph for "scientific verification." This process is driven by a "forced response extraction" strategy initiated by a general large model: the brain compares the quantified data with specialized schemas in the agricultural knowledge graph. For example, when the brain calculates that a certain plot of land needs a large dose of fertilizer, the knowledge enhancement layer will immediately provide feedback that the plot is located within the policy red line of the "core area for black soil protection" and the probability of heavy rainfall in the next 24 hours. At this point, the brain will forcibly revise the initial decision based on this feedback, changing "immediate fertilization" to "variable fertilization after rainfall," thereby ensuring that all decisions operate within scientific logic and regulatory red lines.
[0231] (iv) Dynamic optimal matrix-driven multi-agent collaborative decision-making.
[0232] For complex agricultural decisions involving multiple dimensions (such as integrated pre-production planning), the brain-based decision-making layer decomposes them using a multi-agent collaborative architecture. In implementation, the brain-based decision-making layer acts as the management agent (C-Agent), breaking down the overall goal into multiple execution agent tasks, such as seed selection, soil improvement, and weather mitigation. Subsequently, the system utilizes a dynamic optimal matrix algorithm to dynamically match the best-performing combination of specialized large-scale models to each execution agent based on the real-time characteristics of each sub-task. By calculating the model adaptability score matrix $S_{ij}$, the system can autonomously decide whether to use a general-purpose large-scale model or a high-precision agricultural-specific model. This dynamic allocation mechanism ensures the flexibility of resource scheduling and avoids the limitations of a single model when handling interdisciplinary tasks.
[0233] It should be noted that the agricultural autonomous decision-making method provided in this embodiment of the invention can be executed by an electronic device, a apparatus, or a control module within that apparatus for executing the method. This embodiment of the invention uses an apparatus executing the method as an example to illustrate the agricultural autonomous decision-making apparatus provided in this embodiment of the invention.
[0234] Figure 5 This is a schematic diagram of the agricultural autonomous decision-making device provided in Embodiment 5 of the present invention. The device 100 includes a data acquisition module 10, a semantic alignment module 20, a task decomposition and optimization module 30, and an equipment management and control module 40, wherein:
[0235] The data acquisition module 10 is used to collect multi-source heterogeneous data related to agricultural production.
[0236] The semantic alignment module 20 is used to perform semantic alignment of multi-source heterogeneous data in a shared embedding space.
[0237] The task decomposition and optimization module 30 is used to decompose agricultural tasks into at least one sub-task based on intent recognition by the brain decision layer, and to use a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and to aggregate the reasoning results of each sub-task to generate autonomous decision instructions.
[0238] The equipment control module 40 is used to output autonomous decision-making instructions to intelligent agricultural machinery equipment for closed-loop control.
[0239] Preferably, the data acquisition module 10 is further used to: acquire multi-source heterogeneous data related to agricultural production in real time or semi-real time through satellite remote sensing, UAV aerial survey, sensor feedback, agricultural machinery operation and / or manual field inspection feedback.
[0240] Preferably, the semantic alignment module 20 is further configured to: use a multimodal rotational position embedding method to perform position encoding and alignment of multimodal, multi-source heterogeneous data including images, text, audio, and video in the time and space dimensions, and use pre-defined visual encoders, text encoders, audio encoders, and video encoders to extract features from data of different modalities, so as to minimize the distance between semantically similar content in the shared embedding space.
[0241] Preferably, the multi-source heterogeneous data includes: image sequences. Text sequence audio sequence Video sequence .
[0242] Preferably, the semantic alignment module 20 is further used to: perform semantic alignment of multi-source heterogeneous data in a shared embedding space.
[0243] Its goal is to learn four mapping functions, including: image mapping function Text mapping function Audio mapping function Video mapping function This enables shared embedded spaces Minimize the distance between semantically similar content from different modalities:
[0244] .
[0245] .
[0246] .
[0247] .
[0248] Formula 9.
[0249] The model aligns inputs from different modalities into the same semantic space.
[0250] Feature extraction is performed using a dedicated encoder for each modality:
[0251] For the image modality, image I is processed by a predefined visual encoder, which segments image I into image patches and converts them into embedded representations, as shown in Equation 10. It is the embedding of image patches.
[0252] Formula 10.
[0253] For the text modality, text T is processed by a predefined text encoder and encoded as shown in Equation 11, where ti′∈Rd is the embedding of the word in text T.
[0254] Formula 11.
[0255] For the audio modality, audio A is processed by a predefined audio encoder, as shown in Equation 12, to extract the audio features of audio A and convert them into time-series embeddings.
[0256] Formula 12.
[0257] For the video modality, the video V is processed by a predefined video encoder, as shown in Equation 13. Each frame of the video V is treated as an image, and an embedding is added in the time dimension, where vit is the visual block embedding in each frame t.
[0258] Formula 13.
[0259] A multimodal rotational position embedding method is adopted. In the 1D position embedding of text T, as shown in Equation 14, 1D rotational position embedding is used to capture word order information, where i is the position of the word in the text sequence, j is the embedding dimension index, and d is the embedding dimension.
[0260] Formula 14.
[0261] In the 2D position embedding of image I, the two-dimensional spatial position (h, w) of the image is rotated and encoded as shown in Equation 15, where h and w represent the positions of the words in the image, and d h and d w For the corresponding embedding dimension.
[0262] Formula 15.
[0263] In the 3D position embedding of video V, as shown in Equation 16, corresponding position information is generated for each video lexical unit in both time and space.
[0264] Formula 16.
[0265] Audio A is mainly time series data, and 1D rotation embedding is used to represent time information, as shown in Equation 17, where t represents the time step of each feature in the audio signal and j is the embedding dimension.
[0266] Formula 17.
[0267] Preferably, the brain decision-making layer adopts a multi-agent collaborative architecture, including a management agent and at least one executive agent. The model combination is selected from a hybrid multi-model architecture, including a general large model, an intent recognition model, and multiple agriculture-specific models.
[0268] Task decomposition and optimization module 30 further includes:
[0269] The task decomposition submodule is used by the management agent to decompose agricultural tasks into at least one subtask through an intent recognition model, and different subtasks are executed by different execution agents.
[0270] The model matching submodule is used to perform reasoning on the dynamic optimal matrix algorithm for each agent, based on the real-time characteristics of each subtask, and to complete decision reasoning through multi-agent game collaboration.
[0271] Preferably, the task decomposition submodule further includes:
[0272] The management agent performs task segmentation analysis on agricultural task I using an intent recognition model. The output intent and task segmentation are shown in Equation 1. Simultaneously, an optimal allocation matrix is designed within the hybrid multi-model architecture for each subtask S. j For each ∈P, an optimal allocation matrix M is established, as shown in Formula 2.
[0273] (1)
[0274] Formula 2.
[0275] Where P represents the task segmentation and type, M i Let S represent the i-th model. j This represents the j-th task.
[0276] Preferably, the model matching sub-step further includes:
[0277] For each subtask S j Use the function Select Model(S) j The optimal model is dynamically selected using K models M1, M2, ..., Mn. K Each model M i For subtask S j An adaptive score is calculated by evaluating the model's historical performance, current environment, and / or task characteristics; therefore, for each model M... i and each subtask S j Define an fitness score A ij As shown in Equation 3, f is an evaluation function used to measure model M. i For subtask S j Adaptability.
[0278] Formula 3.
[0279] Model M i The selection objective is to select the fitness score A. ij The highest model is shown in Equation 4.
[0280] Formula 4.
[0281] Design constraints are shown in Equation 5, Model M i The resource consumption is R i R max It is the maximum allowed resource consumption, ensuring that model M i Choose to proceed within the acceptable resource limits.
[0282] Formula 5.
[0283] The dynamic model selection process of the hybrid multi-model architecture is shown in Equation 6, where θ is a predetermined threshold used to determine model M. i The lowest fitness score if the selected model M i If the score is below the threshold, the default model is returned, and M is in this case. selected (S) j () is the optimal model selected based on the allocation matrix.
[0284] Formula 6.
[0285] During the reasoning process, for each selected model M selected (S) j As shown in Formula 7, perform reasoning and output the reasoning results.
[0286] Formula 7.
[0287] The reasoning results of each subtask are aggregated to generate autonomous decision-making instructions, as shown in Formula 8.
[0288] Formula 8.
[0289] Preferably, the agricultural autonomous decision-making device 100 further includes:
[0290] The knowledge enhancement module is used after the task decomposition and optimization module 30 is executed and before the equipment control module 40 is executed. It is used for a forced response mechanism based on a general large model. The brain decision layer compares the reasoning results and / or autonomous decision instructions with the professional knowledge patterns in the predetermined agricultural knowledge graph. The brain decision layer forcibly corrects the autonomous decision instructions based on the feedback.
[0291] The agricultural autonomous decision-making device provided in this embodiment of the invention can achieve Figures 1-4 The various processes implemented in the illustrated method embodiment for agricultural autonomous decision-making will not be described again here to avoid repetition.
[0292] The present invention also provides a storage medium for storing, for example, Figures 1-4 A computer program for any of the methods of autonomous agricultural decision-making. For example, computer program instructions, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operation of the computer, achieving the same technical effect; to avoid repetition, these will not be elaborated further here. The program instructions for invoking the methods of the present invention may be stored in a fixed or removable storage medium, and / or transmitted via data streams in broadcast or other signal carrying media, and / or stored in the storage medium of a computer device operating according to the program instructions.
[0293] According to one embodiment of the present invention, the present invention also provides such a Figure 6 The illustrated electronic device 400 may optionally include a storage medium 200 for storing a computer program and a processor 300 for executing the computer program. When the computer program is executed by the processor 300, it implements any of the aforementioned methods for autonomous agricultural decision-making, triggering the electronic device 400 to execute methods and / or technical solutions based on the foregoing embodiments, achieving the same technical effect. To avoid repetition, these will not be elaborated further here. It should be noted that the electronic devices in this embodiment include mobile electronic devices and non-mobile electronic devices. For example, mobile electronic devices may be mobile phones, tablets, laptops, handheld computers, in-vehicle electronic devices, wearable devices, super mobile personal computers, netbooks, or personal digital assistants, etc., while non-mobile electronic devices may be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This embodiment does not specifically limit the scope of the invention.
[0294] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including associated data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of the present invention can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.
[0295] This invention can be implemented on a computer as a computer-based method, or in dedicated hardware, or a combination of both. Executable code or portions thereof for the method according to the invention can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Optionally, the computer program product includes non-transitory program code components stored on a computer-readable medium so as to execute the method according to the invention when the program product is executed on a computer.
[0296] In an optional embodiment, the computer program includes computer program code components adapted to perform all the steps of the method according to the invention when the computer program is run on a computer. Optionally, the computer program is embodied on a computer-readable medium.
[0297] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0298] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A method for autonomous agricultural decision-making, characterized in that, include: The data acquisition process involves collecting multi-source, heterogeneous data related to agricultural production. The semantic alignment step involves semantically aligning the multi-source heterogeneous data within a shared embedding space. The task decomposition and optimization step involves the brain's decision layer decomposing agricultural tasks into at least one sub-task based on intent recognition, and using a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and aggregating the reasoning results of each sub-task to generate autonomous decision instructions. The equipment management and control steps involve outputting the autonomous decision-making instructions to the intelligent agricultural machinery equipment for closed-loop management and control.
2. The method for agricultural autonomous decision-making according to claim 1, characterized in that, The data acquisition step further includes: The multi-source heterogeneous data related to agricultural production can be acquired in real time or semi-real time through satellite remote sensing, UAV aerial surveying, sensor feedback, agricultural machinery operation and / or manual field inspection feedback.
3. The method for agricultural autonomous decision-making according to claim 1, characterized in that, The semantic alignment step further includes: The multimodal rotational position embedding method is used to encode and align the multimodal heterogeneous data, including images, text, audio, and video, in the temporal and spatial dimensions. Pre-defined visual encoders, text encoders, audio encoders, and video encoders are used to extract features from the data of different modalities, so as to minimize the distance of semantically similar content in the shared embedding space.
4. The method for agricultural autonomous decision-making according to claim 3, characterized in that, The multi-source heterogeneous data includes: image sequences. Text sequence audio sequence Video sequence ; The semantic alignment step further includes: performing semantic alignment of the multi-source heterogeneous data in a shared embedding space; Its goal is to learn four mapping functions, including: image mapping function Text mapping function Audio mapping function Video mapping function This makes the shared embedded space Minimize the distance between semantically similar content from different modalities: ; ; ; ; Official 9; The model aligns inputs from different modalities into the same semantic space; Feature extraction is performed using a dedicated encoder for each modality: For the image modality, image I is processed by a predetermined visual encoder, which segments image I into image blocks and converts them into embedded representations, as shown in Equation 10. It is the embedding of the image patch; Official 10; For the text modality, the text T is processed by a predetermined text encoder and encoded as shown in Formula 11, where ti′∈Rd is the embedding of the word in the text T; Official 11; For the audio modality, audio A is processed by a predetermined audio encoder, as shown in Formula 12, to extract the audio features of audio A and convert them into time series embeddings; Official 12; For the video modality, the video V is processed by a predetermined video encoder, as shown in Equation 13, treating each frame of the video V as an image and adding an embedding in the time dimension, where vit is the visual block embedding in each frame t. Official 13; A multimodal rotational position embedding method is adopted. In the 1D position embedding of the text T, as shown in Formula 14, 1D rotational position embedding is used to capture word order information, where i is the position of the word in the text sequence, j is the embedding dimension index, and d is the embedding dimension. Official 14; In the 2D position embedding of image I, the two-dimensional spatial position (h, w) of the image is rotated and encoded as shown in Formula 15, where h and w represent the positions of words in the image, and d h and d w For the corresponding embedding dimension; Official 15; In the 3D position embedding of the video V, as shown in Formula 16, corresponding position information is generated for each video word in both time and space; Official 16; The audio A is mainly time series data, and 1D rotation embedding is used to represent time information, as shown in Formula 17, where t represents the time step of each feature in the audio signal and j is the embedding dimension. Official 17.
5. The method for agricultural autonomous decision-making according to claim 1, characterized in that, The brain decision-making layer adopts a multi-agent collaborative architecture, including a management agent and at least one executive agent; the model combination is selected from a hybrid multi-model architecture, including a general large model, an intent recognition model, and multiple agricultural-specific models; The task decomposition and optimization steps further include: The task decomposition sub-step involves the management agent decomposing the agricultural task into at least one sub-task using the intent recognition model, with different sub-tasks being executed by different execution agents. The model matching sub-step employs the dynamic optimal matrix algorithm to perform reasoning for each model combination dynamically matched and optimized by the executing agent based on the real-time characteristics of each sub-task, and completes decision reasoning through multi-agent game collaboration.
6. The method for agricultural autonomous decision-making according to claim 5, characterized in that, The task decomposition sub-step further includes: The management agent performs task segmentation analysis on the agricultural task I using an intent recognition model. The output intent and task segmentation are shown in Formula 1. Simultaneously, an optimal allocation matrix is designed within the hybrid multi-model architecture for each subtask S. j For each ∈P, an optimal allocation matrix M is established, as shown in Formula 2; (1) Official 2; Where P represents the task segmentation and type, M i Let S represent the i-th model. j This represents the j-th task; The model matching sub-step further includes: For each of the subtasks S j Use the function Select Model(S) j The optimal model is dynamically selected using K models M1, M2, ..., Mn. K Each model M i For the subtask S j An adaptation score is calculated by evaluating the model's historical performance, current environment, and / or task characteristics; therefore, for each model M... i and each of the subtasks S j Define an fitness score A ij As shown in Formula 3, f is an evaluation function used to measure the model M. i For the subtask S j Adaptability; Official 3; The model M i The selection objective is to select the fitness score A. ij The highest model is shown in Equation 4; Official 4; The design constraints are shown in Equation 5, and the model M... i The resource consumption is R i R max This is the maximum allowed resource consumption, ensuring that the model M... i Choose to proceed within the acceptable resource limits; Official 5; The dynamic model selection process of the hybrid multi-model architecture is shown in Equation 6, where θ is a predetermined threshold used to determine the model M. i The lowest fitness score if the selected model M i If the score is below the threshold, the default model is returned, and M is in this case. selected (S) j () is the optimal model selected based on the allocation matrix; Official 6; During the reasoning process, for each selected model M selected (S) j As shown in Formula 7, perform reasoning and output the reasoning result; Official 7; The reasoning results of each of the sub-tasks are aggregated to generate autonomous decision-making instructions, as shown in Formula 8; Official 8.
7. The method for agricultural autonomous decision-making according to claim 5, characterized in that, Following the task decomposition and optimization step, the following also includes: In the knowledge enhancement step, based on the forced response mechanism of the general large model, the brain decision layer compares the reasoning results and / or the autonomous decision instructions with the professional knowledge patterns in the predetermined agricultural knowledge graph, and the brain decision layer forcibly corrects the autonomous decision instructions according to the feedback.
8. An agricultural autonomous decision-making device constructed based on the method described in any one of claims 1 to 7, characterized in that, The device includes: The data acquisition module is used to collect multi-source heterogeneous data related to agricultural production; A semantic alignment module is used to perform semantic alignment of the multi-source heterogeneous data in a shared embedding space. The task decomposition and optimization module is used to decompose agricultural tasks into at least one sub-task based on intent recognition by the brain decision layer, and to use a predetermined dynamic optimal matrix algorithm to match and schedule optimized model combinations for reasoning for each sub-task, and to aggregate the reasoning results of each sub-task to generate autonomous decision instructions. The equipment management module is used to output the autonomous decision-making instructions to the intelligent agricultural machinery equipment for closed-loop management.
9. A storage medium, characterized in that, Used to store a computer program for performing the method according to any one of claims 1 to 7.
10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.