Intelligent decision-making method based on multi-modal large model and related equipment
Through the intelligent decision-making method of multimodal large models, using visual-text alignment models and irrigation decision-making models, combined with historical multi-source data and multimodal data, the high cost and low efficiency problems of pest and disease diagnosis and irrigation plans in existing technologies are solved, and efficient and accurate pest and disease diagnosis and irrigation strategy generation are achieved.
Patent Information
- Application Number
- CN202510933810.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-26
AI Technical Summary
In existing technologies, pest and disease diagnosis and irrigation plan generation rely on image recognition model training and manual experience for specific pest and disease types, resulting in high migration costs, low efficiency and insufficient accuracy.
An intelligent decision-making method based on a multimodal large model is adopted. Through the visual-text alignment model and the irrigation decision-making model, historical multi-source data and multimodal data are used to diagnose pests and diseases and make irrigation decisions, generate irrigation strategies, and avoid reliance on individual training and manual experience for specific pest and disease types.
It reduces the cost of pest and disease diagnosis and irrigation strategy generation, improves efficiency and accuracy, and can generate precise irrigation strategies based on multimodal data.
Smart Images

Figure CN120707325A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to an intelligent decision-making method based on a multimodal large model and related equipment. Background Art
[0002] With the rapid development of science and technology, more and more industries are improving their work efficiency and quality by combining technology to complete corresponding tasks. For example, in agriculture, by combining technology to predict corresponding pest and disease diagnosis plans and irrigation plans, and generating corresponding irrigation strategies based on the predicted pest and disease diagnosis plans and irrigation plans, so that corresponding irrigation work can be carried out according to the generated irrigation strategies, thereby improving farmers' irrigation efficiency and increasing the survival probability of crops.
[0003] Specifically, existing technologies use sensors to collect corresponding field data, and then use human experience and image recognition models to specify corresponding pest and disease diagnosis plans based on this field data. This also uses human experience to specify corresponding irrigation plans based on this field data. However, this approach requires image recognition models to be individually trained for specific pest and disease types, resulting in high migration costs. Furthermore, this approach relies on human experience, which is not only inefficient, but also varies from user to user, making the generated pest and disease diagnosis and irrigation plans prone to inaccurate due to subjective human judgment. Summary of the Invention
[0004] In view of this, the present invention provides an intelligent decision-making method and related equipment based on a multimodal large model, with the purpose of reducing costs and improving the generation efficiency and accuracy of irrigation strategies.
[0005] In a first aspect, the present application provides an intelligent decision-making method based on a multimodal large model, the method comprising:
[0006] Collecting a current field image of the area to be decided, and preprocessing the current field image to obtain a target field image;
[0007] Inputting the target field image into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to make a prediction and obtain a pest and disease diagnosis plan corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data;
[0008] Dividing the area to be decided into a plurality of partitions, and obtaining current multimodal data of each of the partitions;
[0009] For each partition, processing current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality;
[0010] Inputting the spatiotemporal sequences of each modality of the partition into a pre-trained irrigation decision model, causing the irrigation decision model to make predictions based on the spatiotemporal sequences of each modality to obtain irrigation data for the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data;
[0011] An irrigation strategy for each of the subareas is generated according to the pest and disease diagnosis plan and the irrigation data of each of the subareas.
[0012] Optionally, the step of training the visual-text alignment model to be trained using historical multi-source data to obtain the visual-text alignment model includes:
[0013] Collecting historical multi-source data and preprocessing the historical multi-source data to obtain target multi-source data, wherein the target multi-source data includes historical field image data, historical ground data, and text data related to crops in the field;
[0014] Annotating the pest and disease type corresponding to each historical field image in the historical field image data according to the historical ground data and the text data, and obtaining a target historical field image corresponding to each historical field image;
[0015] Access to multiple agricultural literature;
[0016] Generate a plurality of image-text pairs based on each of the target historical field images and each of the agricultural documents through a visual-text alignment model to be trained; wherein the image-text pairs include the target historical field images and their matching agricultural documents;
[0017] Performing alignment processing on each of the image-text pairs using a visual-text alignment model to be trained to obtain a target image-text pair corresponding to each of the image-text pairs;
[0018] The visual-text alignment model to be trained is trained using each of the target image-text pairs and preset agronomic knowledge, and the trained visual-text alignment model is fine-tuned to obtain the visual-text alignment model.
[0019] Optionally, the step of labeling the pest and disease type corresponding to each image and text in the historical field image data according to the historical ground data and the text data to obtain a target historical field image corresponding to each historical field image includes:
[0020] For each historical field image in the historical field image data, perform image recognition on the historical field image to determine pest and disease information and pest and disease areas within the historical field image;
[0021] Determining the type of pests and diseases corresponding to the historical field image by analyzing the pest and disease information, the pest and disease atlas, the pest and disease identification guide, the historical pest and disease database, and the historical ground data; wherein the text data includes the pest and disease atlas, the pest and disease identification guide, and the historical pest and disease database;
[0022] The pest and disease type is marked on the pest and disease area of the historical field image using a preset marking method to obtain a target historical field image.
[0023] Optionally, performing alignment processing on each of the image-text pairs using a visual-text alignment model to be trained to obtain a target image-text pair corresponding to each of the image-text pairs includes:
[0024] For each of the image-text pairs, randomly masking the agricultural documents in the image-text pair using the visual-text alignment model to be trained, so as to randomly mask part of the text words in the agricultural documents;
[0025] A local area in the target historical field image in the image-text pair is randomly selected through the visual-text alignment model to be trained, and the local area is aligned with the keywords in the agricultural literature after masking to obtain the corresponding target image-text pair.
[0026] Optionally, for each partition, processing current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality includes:
[0027] For each partition, preprocessing the current multimodal data of the partition to obtain target current multimodal data; wherein the target current multimodal data includes target current data of multiple modalities;
[0028] For each of the modalities, performing intra-modal encoding on the target current data of the modality to obtain an embedding vector corresponding to the modality;
[0029] The embedding vectors of each of the modalities are cross-modally fused and encoded to obtain corresponding spatiotemporal sequences.
[0030] Optionally, the method further includes:
[0031] Determining the irrigation priority and irrigation status of each partition according to the irrigation recommendation value in the irrigation data of each partition; wherein the irrigation status is whether irrigation is required or not required;
[0032] Acquiring map data of each of the sub-areas and agricultural machinery information of the agricultural machinery used for irrigation, and determining terrain complexity and obstacle information of each of the sub-areas based on the map data of each of the sub-areas;
[0033] generating a path search map based on each of the partitions and their terrain complexity and obstacle information;
[0034] A preset path planning algorithm is used to generate a path planning solution for the area to be decided based on the path search graph, the agricultural machinery information, the irrigation priority and irrigation status of each partition.
[0035] Optionally, the method further includes:
[0036] Obtaining a path optimization target, and optimizing the path planning scheme according to the path optimization target to obtain a target path planning scheme;
[0037] Generate corresponding instructions according to the target path planning scheme and the irrigation strategies of each of the partitions, and send the instructions to the agricultural machine, so that the agricultural machine irrigates the area to be decided according to the target path planning scheme.
[0038] A second aspect of the present application provides an intelligent decision-making system based on a multimodal large model, the system comprising:
[0039] The current field image acquisition module is used to acquire the current field image of the area to be decided and pre-process the current field image to obtain the target field image;
[0040] A pest and disease diagnosis solution prediction module is configured to input the target field image into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to perform predictions and obtain a pest and disease diagnosis solution corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data by the training module;
[0041] A multimodal data acquisition module, configured to divide the area to be decided into a plurality of partitions and acquire current multimodal data of each of the partitions;
[0042] a data processing module, configured to process, for each partition, current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality;
[0043] an irrigation data prediction module, configured to input the spatiotemporal sequences of each modality of the partition into a pre-trained irrigation decision model, causing the irrigation decision model to perform predictions based on the spatiotemporal sequences of each modality to obtain irrigation data for the partition; wherein the irrigation decision model is trained using historical multimodal data to train the irrigation decision model to be trained;
[0044] The irrigation strategy generation module is used to generate an irrigation strategy for each of the partitions according to the pest and disease diagnosis plan and the irrigation data of each of the partitions.
[0045] The third aspect of the present application provides an electronic device, comprising: a processor and a memory, wherein the processor and the memory are connected via a bus; wherein the processor is used to call and execute a program stored in the memory; and the memory is used to store a program, and the program is used to implement an intelligent decision-making method based on a multimodal large model as provided in the first aspect of the present application.
[0046] The fourth aspect of the present application provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to execute an intelligent decision-making method based on a multimodal large model as provided in the first aspect of the present application.
[0047] The present application provides an intelligent decision-making method and related equipment based on a multimodal large model, which obtains a target field image by collecting a current field image of the area to be decided and preprocessing the current field image; inputs the target field image into a pre-trained visual-text alignment model, and enables the visual-text alignment model to use the target field image for prediction to obtain a disease and insect pest diagnosis scheme corresponding to the area to be decided; wherein, the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data; the area to be decided is divided into multiple partitions, and the image of each partition is obtained. The method comprises the following steps: obtaining current multimodal data of each partition; processing the current data of each mode in the current multimodal data of the partition to obtain a spatiotemporal sequence of each mode; inputting the spatiotemporal sequence of each mode of the partition into a pre-trained irrigation decision model, so that the irrigation decision model makes a prediction based on the spatiotemporal sequence of each mode to obtain the irrigation data of the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data; and generating an irrigation strategy for each partition according to the pest and disease diagnosis scheme and the irrigation data of each partition. It can be seen that the technical solution provided by the present application can pre-train the visual-text alignment model to be trained using historical multi-source data to obtain a visual-text alignment model, so that the visual-text alignment model can be used to predict the pest and disease diagnosis plan for the area to be decided based on the target field images of the area to be decided. The present application does not need to train a separate image recognition model for a specific type of pest and disease, nor does it need to be combined with manual experience. It is not only low-cost but also highly accurate. The present application uses historical multimodal data to pre-train the corresponding irrigation decision model, so that the irrigation data of each partition can be predicted using the irrigation decision model and the multimodal data of each partition. It does not need to be combined with manual experience and can determine accurate irrigation data. Finally, accurate irrigation strategies are formulated based on accurate pest and disease diagnosis plans and irrigation data. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0049] Figure 1 A flowchart of an intelligent decision-making method based on a multimodal large model provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of the structure of an intelligent decision-making system based on a multimodal large model provided by an embodiment of the present invention;
[0051] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0054] See also Figure 1 , shows a flow chart of an intelligent decision-making method based on a multimodal large model provided in an embodiment of the present application, the intelligent decision-making method based on a multimodal large model specifically includes the following steps:
[0055] S101: collecting a current field image of a region to be decided, and preprocessing the current field image to obtain a target field image.
[0056] In the specific process of executing step S101, after determining the area to be decided where a decision needs to be made, a drone equipped with a hyperspectral camera can be used to collect the current field image of the area to be decided; in order to eliminate the problem of blurred images collected due to the instability of the drone aircraft, the collected current field image can be deblurred; the current field image after deblurring is subjected to illumination correction to enhance the image with uneven illumination, improve the image contrast, and obtain the corresponding target field image; wherein, deblurring and illumination correction are both processes of preprocessing the current field image.
[0057] S102: Input the target field image into a pre-trained vision-text alignment model, so that the vision-text alignment model uses the target field image to perform prediction and obtain a pest and disease diagnosis plan corresponding to the area to be decided.
[0058] In the specific process of executing step S102, the visual-text alignment model to be trained can be pre-trained using historical multi-source data to obtain a visual-text alignment model, so that after obtaining the target field image of the area to be decided, the target field image is input into the pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to perform prediction and obtains the pest and disease diagnosis plan corresponding to the area to be decided.
[0059] In an embodiment of the present application, historical multi-source data can be acquired in advance and processed to obtain a plurality of target historical field images marked with pest and disease types; a plurality of image-text pairs are generated based on each target historical field image and agricultural literature, so that the visual-text alignment model to be trained can be trained using each generated image-text pair to obtain a visual-text alignment model.
[0060] It should be noted that the visual-text alignment model to be trained can be a Transformer model, which is not limited in the embodiments of this application.
[0061] Optionally, the process of using historical multi-source data to train the visual-text alignment model to be trained to obtain the visual-text alignment model can be: collecting historical multi-source data and preprocessing the historical multi-source data to obtain target multi-source data, wherein the target multi-source data includes historical field image data, historical ground data and text data related to crops in the field; annotating the pest and disease type corresponding to each historical field image in the historical field image data according to the historical ground data and the text data to obtain a target historical field image corresponding to each historical field image; obtaining multiple agricultural documents; generating multiple image-text pairs according to each target historical field image and each agricultural document through the visual-text alignment model to be trained; wherein the image-text pair includes the target historical field image and its matching agricultural document; aligning each image-text pair through the visual-text alignment model to be trained to obtain a target image-text pair corresponding to each image-text pair; training the visual-text alignment model to be trained using each target image-text pair and preset agronomic knowledge, and fine-tuning the trained visual-text alignment model to obtain a visual-text alignment model.
[0062] In an embodiment of the present application, historical field image data of the field can be collected through satellite images or drones equipped with hyperspectral cameras, and historical ground data of the field can be collected through a ground sensor network to obtain text data related to crops in the field; wherein, the historical ground data includes historical soil moisture, historical soil temperature and other information of the field; the text data may include pest and disease atlases, pest and disease identification guides, historical pest and disease databases and other data.
[0063] In some embodiments, historical multi-source data can be collected, wherein the historical multi-source data includes historical field image data before preprocessing, historical ground data and text data related to crops in the field, and the historical field image data includes multiple historical field images; in order to eliminate the problem that the collected field images are relatively blurred due to the instability of the drone aircraft, the collected historical field images can be deblurred; the historical field images after deblurring are subjected to illumination correction to enhance the image with uneven illumination and improve the image contrast; finally, in order to ensure that each historical field image has the same spatial resolution, the historical field images after illumination correction can be subjected to a resolution consistency check to obtain the preprocessed historical field image; wherein, deblurring, illumination correction and resolution consistency check are the process of preprocessing the historical field images.
[0064] As an implementation method of an embodiment of the present application, the process of labeling the pest and disease type corresponding to each historical field image in the historical field image data according to the historical ground data and text data, and obtaining the target historical field image corresponding to each historical field image can be: for each historical field image in the historical field image data, performing image recognition on the historical field image to determine the pest and disease information and pest and disease area in the historical field image; analyzing according to the pest and disease information, pest and disease atlas, pest and disease identification guide, historical pest and disease database and historical ground data to determine the pest and disease type corresponding to the historical field image; using a preset labeling method to label the pest and disease type on the pest and disease area of the historical field image to obtain the target historical field image.
[0065] In some embodiments, at least one initial pest and disease type that matches the pest and disease information can be determined from the pest and disease map based on the pest and disease identification guide and the historical pest and disease database, and the pest and disease type corresponding to the historical field image can be determined from each initial pest and disease type based on the historical ground data.
[0066] It should be noted that the preset annotation methods can be bounding box annotation, polygon annotation and pixel-level annotation.
[0067] In some embodiments, if the default annotation method is bounding box annotation, a rectangular box can be drawn to identify the corresponding pest and disease area in the historical field image, and the corresponding pest and disease type can be labeled on the corresponding pest and disease area using the corresponding annotation tool. For example, a red box can be used to mark the area of rice blast-infected leaves in the historical field image.
[0068] It should be noted that the labeling tool can be LabelImg, Label Studio, image annotation tool (VGG Image Annotato, VIA), visual data annotation tool (Computer Vision Annotation Tool, CVAT), which is not limited in this embodiment of the present application.
[0069] In other embodiments, if polygon annotation is the preset annotation method, the contours of diseased spots or insect bodies in historical field images can be precisely delineated, the corresponding pest areas can be determined, and the corresponding pest types can be labeled on the corresponding pest areas using corresponding annotation tools. For example, aphid clusters in historical field images can be precisely delineated using polygons.
[0070] In other embodiments, if the preset annotation method is pixel-level annotation, each pixel in the historical field image can be classified to distinguish healthy tissue from diseased tissue, and the portion corresponding to the diseased tissue can be identified as a pest area. The corresponding pest area can then be labeled with the corresponding pest type using a corresponding annotation tool. For example, pixels on leaves infected with rice blast in the historical field image can be labeled as category 1, and pixels on healthy leaves as category 0. The area corresponding to category 1 can then be identified as a pest area.
[0071] It should be noted that after the pest and disease type labels are completed for the historical field images, the target historical field images labeled with the pest and disease type can be stored in a structured file format. The structured file format can be JSON (such as the COCO dataset format), XML (such as the PASCAL VOC format), etc., which is not limited in this embodiment of the present application.
[0072] In an embodiment of the present application, a corresponding image-text matching loss function (CrossEntropy Loss) can be pre-configured in the visual-text alignment model to be trained, so that after obtaining multiple target historical field images, agricultural literature related to crop diseases and pests can be further obtained, and each target historical field image and each agricultural literature can be input into the visual-text alignment model to be trained, so that the visual-text alignment model to be trained uses the pre-configured image-text matching loss function to combine each target historical field image and each agricultural literature in pairs to obtain multiple initial image-text pairs; for each initial image-text pair, the correlation between the target historical field image and the agricultural literature in the initial image-text pair is calculated, and it is determined whether the correlation is greater than the correlation threshold; if it is greater, the initial image-text pair is determined to be an image-text pair; if it is not greater, the initial image-text pair is eliminated.
[0073] It should be noted that this application can also provide a corresponding dynamic knowledge graph so that various agricultural documents can be updated in real time using the dynamic knowledge graph, so as to achieve the purpose of grasping the latest information in real time and providing users with the most accurate suggestions.
[0074] It should also be noted that the dynamic knowledge graph not only contains a rich agricultural knowledge system, but also supports real-time content updates through an incremental knowledge injection mechanism, such as the characteristics of newly discovered crop varieties or the latest pest and disease control methods.
[0075] In some embodiments, after obtaining multiple image-text pairs, masked language modeling can be performed on each image-text pair. The target historical field image and agricultural literature in the masked language modeling image-text pairs can then be aligned to obtain a target image-text pair corresponding to each image-text pair. This allows the obtained target image-text pairs to be used in subsequent training of a visual-text alignment model to improve the model's ability to understand context.
[0076] Optionally, the process of aligning each of the image-text pairs through the visual-text alignment model to be trained to obtain the target image-text pair corresponding to each of the image-text pairs can be: for each image-text pair, randomly masking the agricultural documents in the image-text pair through the visual-text alignment model to be trained to randomly mask part of the text words in the agricultural documents; randomly selecting a local area in the target historical field image in the image-text pair through the visual-text alignment model to be trained, and aligning the local area with the keywords in the agricultural documents after masking to obtain the corresponding target image-text pair.
[0077] In actual application, the masked language modeling function (MLM Loss) and the image-text contrast learning function (InfoNCE Loss) can also be pre-configured in the visual-text alignment model to be trained; so that after obtaining multiple image pairs, the visual-text alignment model to be trained uses the masked language modeling function to randomly mask the agricultural documents in the image-text pairs, so as to randomly mask some text words in the agricultural documents, and randomly select local areas in the target historical field images in the image-text pairs; use the image-text contrast learning function to align the local areas with the keywords in the masked agricultural documents to obtain the corresponding target image-text pairs.
[0078] It should be noted that the image-text comparison learning function can effectively explore the potential correlations between different modal data. At the same time, the contrastive learning mechanism can also narrow the semantic gap, thereby realizing the effective integration and understanding of image, text and sensor data, and further improving the accuracy of the model's agricultural data analysis, thereby providing strong support for the development of precision agriculture.
[0079] In an embodiment of the present application, after obtaining each target image-text pair, the visual-text alignment model to be trained can be trained using each target image-text pair, and the trained visual-text alignment model can be adapted to the agricultural scene of a specific crop or region. At the same time, the parameter fine-tuning method (Low-Rank Adaptation, LoRA) is used to inject corresponding agronomic knowledge into the visual-text alignment model in the agricultural scene, and a low-rank matrix is inserted into the Transformer layer of the visual-text alignment model to update the agronomic knowledge in real time, thereby achieving the first stage fine-tuning of the trained visual-text alignment model, reducing the computational overhead of the visual-text alignment model, and retaining the generalization ability of the original model. Finally, a small number of target image-text pairs in each target image-text pair and customized image-text pairs constructed by expert knowledge can be used to continue to perform second-stage fine-tuning on the visual-text image after the first stage fine-tuning to obtain the final visual-alignment model.
[0080] It should be noted that using LoRA technology to inject agronomic knowledge of a specific region is essentially a process of introducing local agricultural field knowledge into the visual-text alignment model. It can be adapted to actual application scenarios of specific regions, crops or pest and disease types with low parameter cost and high efficiency without significantly changing the original model structure and parameters.
[0081] It should also be noted that this application uses a large-scale model with a two-stage pre-training and fine-tuning training strategy as the core of the inference engine. First, the model is pre-trained on a wide range of agricultural-related datasets (respectively historical field images of various targets and various agricultural literature) to acquire basic semantic understanding capabilities. Then, it is fine-tuned using task-specific datasets (a small number of target image-text pairs from each target image-text pair and customized image-text pairs constructed based on expert knowledge). In particular, combined with the efficient parameter fine-tuning technology of low-rank adaptation (LoRA), the performance of the visual-text alignment model for specific agricultural applications is significantly improved while reducing computing resource consumption.
[0082] In this embodiment of the present application, after obtaining the final visual-text alignment model, it can be further deployed to corresponding edge devices (such as agricultural machinery terminals) to support offline diagnosis. This allows the visual-text alignment model to input the target field image of the decision-making area into the visual-text alignment model, which then uses the target field image to make predictions and obtain a pest and disease diagnosis solution for the decision-making area.
[0083] It should be noted that the pest and disease diagnosis plan may include the pest and disease type, prevention plan, pesticide dosage, etc. corresponding to the target field image.
[0084] It should also be noted that in order to overcome the high latency and large bandwidth consumption in the traditional cloud computing model, and to meet the needs of agricultural production sites for real-time decision support, the trained visual-text alignment model can be deployed to the corresponding edge devices, while retaining the core processing unit to run in the cloud. In this way, the edge-cloud collaborative architecture can not only reduce the data transmission cost, but also ensure the corresponding speed of the corresponding system, reduce low-latency response, and provide possibilities for application scenarios such as smart agricultural monitoring and instant pest and disease warning, greatly enhancing the intelligence level and efficiency of agricultural production.
[0085] S103: Divide the area to be decided into multiple partitions, and obtain current multimodal data of each partition.
[0086] In the specific process of executing step S103, the area to be decided can be divided into multiple regular or irregular grid units, wherein one grid unit corresponds to a partition that needs to be irrigated; for each partition, the current multimodal data corresponding to the partition can be obtained; wherein the current multimodal data includes current data of multiple modalities; the multiple modalities may include sensor data modalities, meteorological data modalities, and biological or agricultural science data modalities.
[0087] In some embodiments, the area to be decided may be divided into multiple partitions using a hexagonal grid division method.
[0088] It should be noted that for each zone, soil conductivity sensor data (current data corresponding to the sensor data modality) can be collected. This soil conductivity sensor data includes at least the zone's soil conductivity. Soil conductivity is a measure of the soluble salt content in the soil, indirectly reflecting soil fertility and moisture content. The soil conductivity sensor can collect zone-specific soil conductivity data every 30 minutes to monitor soil moisture trends in real time.
[0089] The current data in the meteorological data modality can include weather forecasts for the next period (e.g., the next 72 hours). Specifically, this can include key information such as rainfall probability and evaporation. This data is crucial for predicting future water supply and demand. For example, if the forecast indicates a high probability of rainfall, irrigation may need to be reduced. High evaporation, on the other hand, means crops may need additional water, so irrigation may need to be increased.
[0090] The current data of the biological or agricultural science data modality may include crop growth stage information of crops within the partition.
[0091] It's important to note that different crops have different water requirements at different stages of their growth cycle. For example, corn has a specific water requirement threshold at the heading stage. Understanding this can help precisely adjust irrigation strategies to meet the crop's water needs at that specific growth stage, thereby optimizing water use efficiency and promoting healthy crop growth.
[0092] S104: For each partition, process the current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality.
[0093] In the specific execution step S104, after obtaining the partitioned current multimodal data, for each modality, the current data of the modality can be preprocessed to obtain the corresponding target current data, and the target current data can be encoded to obtain the corresponding spatiotemporal sequence.
[0094] Optionally, for each partition, the current data of each modality in the current multimodal data of the partition is processed to obtain the spatiotemporal sequence of each modality. The process can be: for each partition, the current multimodal data of the partition is preprocessed to obtain the target current multimodal data; wherein the target current multimodal data includes target current data of multiple modalities; for each modality, the target current data of the modality is intra-modally encoded to obtain the embedding vector corresponding to the modality; the embedding vectors of each modality are cross-modally fused and encoded to obtain the corresponding spatiotemporal sequence.
[0095] In actual application, for each partition, the current data of each mode of the partition can be cleaned and normalized to complete the preprocessing of the current data and obtain the target current data; feature extraction and embedding coding are performed on the target current data to map the target current data to a vector space of unified dimension to obtain the corresponding embedding vector; the embedded vector is expanded modal fusion coded in the time and feature dimensions to obtain the corresponding spatiotemporal sequence.
[0096] S105: Input the spatiotemporal sequences of each mode of the partition into a pre-trained irrigation decision model, so that the irrigation decision model makes predictions based on the spatiotemporal sequences of each mode to obtain the irrigation data of the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data.
[0097] In an embodiment of the present application, corresponding historical multimodal data and corresponding irrigation decision data can be obtained in advance, and the historical data of each mode in the historical multimodal data can be preprocessed to obtain target historical data; feature extraction and embedding coding are performed on the target historical data to map the target historical data into a vector space of unified dimension to obtain a corresponding historical embedding vector; the historical embedding vector is subjected to modal fusion coding in the time and feature dimensions to obtain a corresponding historical spatiotemporal sequence; the historical spatiotemporal sequence and the corresponding irrigation decision data are input into the irrigation decision model to be trained, so that the irrigation decision model to be trained uses the input historical spatiotemporal sequence for prediction, and the predicted irrigation decision is close to the corresponding irrigation decision data as the training goal, and the parameters (dynamic weight coefficients) of the irrigation decision model to be trained are adjusted until the irrigation decision model to be trained reaches convergence to obtain the final irrigation decision model.
[0098] In the specific process of executing step S105, after obtaining the irrigation decision model, the spatiotemporal sequence corresponding to the partition can be input into the irrigation decision model, so that the irrigation decision model uses the input spatiotemporal sequence to perform prediction and obtain the irrigation data corresponding to the partition.
[0099] It should be noted that the spatiotemporal series can include the soil moisture status index, future rainfall probability and evaporation and transpiration of the partition; the irrigation decision model uses the input spatiotemporal series to make predictions, and the process of obtaining the irrigation data corresponding to the partition is shown in formula (1):
[0100] Q=α·Wsoil+β·Prain−γ·Et (1)
[0101] It should also be noted that α, β, and γ are dynamic weight coefficients; Wsoil represents the current soil moisture status of the zone, typically derived from soil conductivity, volumetric water content, or other sensor data. Higher values indicate moister soil; however, if it falls below the crop water requirement threshold, it indicates a water shortage. Prain represents the probability of future rainfall; and Et represents evapotranspiration, which represents the total amount of water evaporated from the soil and crop evaporation per unit time, reflecting the intensity of crop water demand. Higher Et values indicate higher crop water requirements.
[0102] S106: Generate an irrigation strategy for each sub-area based on the pest and disease diagnosis plan and the irrigation data of each sub-area.
[0103] In the specific process of executing step S106, after obtaining the irrigation data of each partition, for each partition, the pest and disease type, prevention and control plan and pesticide dosage of the partition can be determined according to the pest and disease diagnosis plan of the area to be decided. Finally, the irrigation strategy of the partition is generated according to the pest and disease type, prevention and control plan, pesticide dosage and irrigation recommendation value in the irrigation data of the partition.
[0104] The present application provides an intelligent decision-making method based on a multimodal large model, which collects current field images of the area to be decided and preprocesses the current field images to obtain target field images; inputs the target field images into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field images to make predictions and obtains a pest and disease diagnosis plan corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data; the area to be decided is divided into multiple partitions, and the current multimodal data of each partition is obtained; for each partition, the current data of each mode in the current multimodal data of the partition is processed to obtain a spatiotemporal sequence of each mode; the spatiotemporal sequences of each mode of the partition are input into a pre-trained irrigation decision model, so that the irrigation decision model makes predictions based on the spatiotemporal sequences of each mode to obtain irrigation data for the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data; and an irrigation strategy for each partition is generated based on the pest and disease diagnosis plan and the irrigation data of each partition. It can be seen that the technical solution provided by the present application can pre-train the visual-text alignment model to be trained using historical multi-source data to obtain a visual-text alignment model, so that the visual-text alignment model can be used to predict the pest and disease diagnosis plan for the area to be decided based on the target field images of the area to be decided. The present application does not need to train a separate image recognition model for a specific type of pest and disease, nor does it need to be combined with manual experience. It is not only low-cost but also highly accurate. The present application uses historical multimodal data to pre-train the corresponding irrigation decision model, so that the irrigation data of each partition can be predicted using the irrigation decision model and the multimodal data of each partition. It does not need to be combined with manual experience and can determine accurate irrigation data. Finally, accurate irrigation strategies are formulated based on accurate pest and disease diagnosis plans and irrigation data.
[0105] Furthermore, in an embodiment of the present application, after obtaining the irrigation strategy for each partition of the area to be decided, the irrigation priority and irrigation status of each partition can be further determined based on the irrigation recommendation value in the irrigation data of each partition; wherein the irrigation status is whether irrigation is required or not; the map data of each partition and the agricultural machinery information of the agricultural machinery used for irrigation are obtained, and the terrain complexity and obstacle information of each partition are determined based on the map data of each partition; a path search map is generated based on each partition and its terrain complexity and obstacle information; a preset path planning algorithm is used to generate a path planning scheme for the area to be decided based on the path search map, agricultural machinery information, the irrigation priority and irrigation status of each partition; the path optimization target is obtained, and the path planning scheme is optimized according to the path optimization target to obtain a target path planning scheme; corresponding instructions are generated according to the target path planning scheme and the irrigation strategy of each partition, and the instructions are sent to the agricultural machinery, so that the agricultural machinery irrigates the area to be decided according to the target path planning scheme.
[0106] In actual application, since path planning depends on a variety of spatial and task-related data, it is possible to further obtain map data of each partition and agricultural machinery information of agricultural machinery used for irrigation, where the agricultural machinery information includes the mechanical parameters and real-time status of the agricultural machinery; use the large model to perform spatiotemporal reasoning based on the irrigation recommendation value of each partition to obtain the irrigation priority of each partition; judge whether the irrigation recommendation value of the partition is 0; if it is not zero, determine that the irrigation status of the partition is that irrigation is required; if it is zero, determine that the irrigation status of the partition is that irrigation is not required; determine the terrain complexity and obstacle information of the partition based on the field boundary, partition number, terrain elevation, obstacle location and other data in the partition map data, where the obstacle information includes the existence of obstacles in the partition and the location of the obstacle or the absence of obstacles in the partition.
[0107] It should be noted that the mechanical status of agricultural machinery may include parameters such as the model size, turning radius, maximum speed, unit energy consumption, etc.; the real-time status of agricultural machinery may include the current location of the agricultural machinery, power / oil level, whether it is in operation, etc.
[0108] It's also worth noting that map data for each zone can be collected through methods such as Graphic Information Systems (GIS), drones, and Real-time Kinematic (RTK) positioning. The real-time status of agricultural machinery can be obtained through Global Positioning Systems (GPS) or IoT sensors.
[0109] In some embodiments, the area to be decided can be abstracted into a graph structure, wherein the nodes in the graph structure are the center points, starting points, or focus points of each partition, and the lines between the nodes (edges in the graph structure) represent feasible paths from one partition to another; based on comprehensive indicators such as the edges in the graph structure, edge length (path length), energy consumption cost, terrain complexity of each partition, and obstacle information, a corresponding path search graph is generated.
[0110] It should be noted that the generated graph structure can be represented using an adjacency matrix or an adjacency list to facilitate the processing of the subsequent path planning algorithm.
[0111] In some embodiments, the preset path planning algorithm can be a heuristic search algorithm, a meta-heuristic algorithm, etc.; among them, the heuristic search algorithm can be an A* algorithm or Dijkstra, the A* algorithm can combine the corresponding heuristic function with the actual cost to find the shortest path, and is suitable for small farmlands without obstacles; Dijkstra can find the global shortest path, and is suitable for farmlands with high requirements on path accuracy; the meta-heuristic algorithm can be a genetic algorithm, a particle swarm optimization algorithm, an ant colony algorithm, etc.; among them, the genetic algorithm has a strong global search capability and is suitable for multi-objective optimization; the particle swarm optimization algorithm has a fast convergence speed and is suitable for continuous space optimization; the ant colony algorithm can simulate the foraging behavior of ants and is suitable for path exploration.
[0112] Furthermore, after determining the preset path planning algorithm, the corresponding path planning model can be trained using reinforcement learning (RL) in combination with the preset path planning algorithm, so that the path search map, agricultural machinery information, irrigation priority of each partition, and irrigation status can be input into the path planning model to obtain the corresponding path planning scheme; wherein, the path planning scheme includes at least the moving direction and speed of each path.
[0113] It should be noted that the use of reinforcement learning algorithms can automatically find the best path and obtain the optimal path planning solution.
[0114] In some embodiments, path optimization objectives may include minimum road length, minimum energy consumption, irrigation priority scheduling, and multi-machine collaboration; optimizing the path planning scheme through path optimization objectives can effectively reduce redundant detour paths in the path planning scheme, give priority to platform sections, reduce sharp turns, give priority to accessing high-priority partitions, and use task allocation algorithms to assign corresponding paths to multiple agricultural machines, so that multiple agricultural machines can cooperate with each other.
[0115] It should be noted that the higher the irrigation recommendation value, the higher the priority of the corresponding partition.
[0116] In some embodiments, after obtaining the corresponding target path planning scheme, it can be visually displayed in the corresponding interface for manual review, and the target path planning scheme and the irrigation strategy of each partition can be converted into instructions in an instruction format (such as a Waypoint sequence) recognized by the agricultural machinery controller. Finally, the instructions are sent to the agricultural machinery controller so that the agricultural machinery controller can irrigate the decision area according to the target path planning scheme and the irrigation strategy of each partition.
[0117] It should be noted that during the irrigation process of agricultural machinery, the status of the agricultural machinery can be monitored in real time, and the path can be dynamically adjusted based on feedback (if there are sudden obstacles or task changes).
[0118] It should also be noted that before the instructions generated according to the target path planning scheme and the irrigation strategy of each partition are sent to the agricultural machinery controller, irrigation simulation can be carried out according to the target path planning scheme and the irrigation strategy of each partition through digital twins, and the changes in soil moisture in each partition after simulated irrigation can be detected. This can not only avoid the risk of over-irrigation, but also ensure that the actual implementation effect achieves the expected goal.
[0119] Based on the intelligent decision-making method based on the multimodal large model provided in the above embodiment of the present application, correspondingly, the embodiment of the present application also provides an intelligent decision-making system based on the multimodal large model, such as Figure 2 As shown in Figure 1, the intelligent decision-making system based on the multimodal large model includes:
[0120] The current field image acquisition module 21 is used to acquire the current field image of the area to be decided and pre-process the current field image to obtain the target field image;
[0121] The pest and disease diagnosis solution prediction module 22 is used to input the target field image into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to make a prediction and obtain a pest and disease diagnosis solution corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data in the training module;
[0122] The multimodal data acquisition module 23 is used to divide the area to be decided into multiple partitions and obtain the current multimodal data of each partition;
[0123] A data processing module 24 is configured to process, for each partition, the current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality;
[0124] The irrigation data prediction module 25 is used to input the spatiotemporal sequences of each modality of the partition into a pre-trained irrigation decision model, so that the irrigation decision model makes predictions based on the spatiotemporal sequences of each modality to obtain irrigation data for the partition; wherein the irrigation decision model is trained using historical multimodal data to train the irrigation decision model to be trained;
[0125] The irrigation strategy generating module 26 is used to generate an irrigation strategy for each sub-area based on the pest and disease diagnosis plan and the irrigation data of each sub-area.
[0126] The specific principles and execution processes of each unit in the intelligent decision-making system based on the multimodal large model disclosed in the above-mentioned embodiment of the present application are the same as the intelligent decision-making method based on the multimodal large model disclosed in the above-mentioned embodiment of the present application. Please refer to the corresponding parts of the intelligent decision-making method based on the multimodal large model disclosed in the above-mentioned embodiment of the present application, and no further details will be given here.
[0127] The present application provides an intelligent decision-making system based on a multimodal large model, which collects current field images of the area to be decided and preprocesses the current field images to obtain target field images; inputs the target field images into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field images to make predictions and obtains a pest and disease diagnosis plan corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data; the area to be decided is divided into multiple partitions, and the current multimodal data of each partition is obtained; for each partition, the current data of each mode in the current multimodal data of the partition is processed to obtain a spatiotemporal sequence of each mode; the spatiotemporal sequences of each mode of the partition are input into a pre-trained irrigation decision model, so that the irrigation decision model makes predictions based on the spatiotemporal sequences of each mode to obtain irrigation data for the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data; and an irrigation strategy for each partition is generated based on the pest and disease diagnosis plan and the irrigation data of each partition. It can be seen that the technical solution provided by the present application can pre-train the visual-text alignment model to be trained using historical multi-source data to obtain a visual-text alignment model, so that the visual-text alignment model can be used to predict the pest and disease diagnosis plan for the area to be decided based on the target field images of the area to be decided. The present application does not need to train a separate image recognition model for a specific type of pest and disease, nor does it need to be combined with manual experience. It is not only low-cost but also highly accurate. The present application uses historical multimodal data to pre-train the corresponding irrigation decision model, so that the irrigation data of each partition can be predicted using the irrigation decision model and the multimodal data of each partition. It does not need to be combined with manual experience and can determine accurate irrigation data. Finally, accurate irrigation strategies are formulated based on accurate pest and disease diagnosis plans and irrigation data.
[0128] Optionally, the visual-text alignment model to be trained is trained using historical multi-source data to obtain a training module for the visual-text alignment model, specifically for:
[0129] Collecting historical multi-source data and preprocessing the historical multi-source data to obtain target multi-source data, wherein the target multi-source data includes historical field image data, historical ground data, and text data related to crops in the field;
[0130] Label the pest and disease type corresponding to each historical field image in the historical field image data according to the ground data and the text data, and obtain the target historical field image corresponding to each historical field image;
[0131] Access to multiple agricultural literature;
[0132] Generate multiple image-text pairs based on each target historical field image and each agricultural document using a visual-text alignment model to be trained; wherein the image-text pairs include the target historical field image and its matching agricultural document;
[0133] Align each image-text pair using the visual-text alignment model to be trained to obtain the target image-text pair corresponding to each image-text pair;
[0134] The visual-text alignment model to be trained is trained using each target image-text pair and preset agronomic knowledge, and the trained visual-text alignment model is fine-tuned to obtain a visual-text alignment model.
[0135] Optionally, the pest and disease type corresponding to each historical field image in the historical field image data is annotated according to the ground data and the text data, and a training module for a target historical field image corresponding to each historical field image is obtained, specifically for:
[0136] For each historical field image in the historical field image data, image recognition is performed on the historical field image to determine pest and disease information and pest and disease areas within the historical field image;
[0137] Analyze pest and disease information, pest and disease maps, pest and disease identification guides, historical pest and disease databases, and historical ground data to determine the pest and disease types corresponding to historical field images; the text data includes pest and disease maps, pest and disease identification guides, and historical pest and disease databases;
[0138] The pest and disease types are marked on the pest and disease areas of the historical field images using a preset labeling method to obtain the target historical field images.
[0139] Optionally, a training module for generating multiple image-text pairs based on each target historical field image and each agricultural document is used by the visual-text alignment model to be trained, specifically for:
[0140] For each image-text pair, the agricultural literature in the image-text pair is randomly masked using the trained visual-text alignment model to randomly mask some of the text words in the agricultural literature.
[0141] The visual-text alignment model to be trained randomly selects local areas in the target historical field image in the image-text pair, and aligns the local areas with the keywords in the agricultural literature after masking to obtain the corresponding target image-text pair.
[0142] Optionally, for each partition, the current data of each modality in the current multimodal data of the partition is processed to obtain a data processing module of the spatiotemporal sequence of each modality, specifically for:
[0143] For each partition, preprocessing the current multimodal data of the partition to obtain target current multimodal data; wherein the target current multimodal data includes target current data of multiple modalities;
[0144] For each modality, the target current data of the modality is encoded intramodally to obtain the embedding vector corresponding to the modality;
[0145] The embedding vectors of each modality are cross-modally fused and encoded to obtain the corresponding spatiotemporal sequence.
[0146] Optionally, the intelligent decision-making system based on the multimodal large model provided in the embodiment of the present application further includes: a path planning solution module, which is specifically used to:
[0147] Determine the irrigation priority and irrigation status of each zone according to the irrigation recommendation value in the irrigation data of each zone; wherein the irrigation status is irrigation required or irrigation not required;
[0148] Obtaining map data of each sub-area and agricultural machinery information of agricultural machinery used for irrigation, and determining terrain complexity and obstacle information of each sub-area based on the map data of each sub-area;
[0149] Generate a path search map based on each partition and its terrain complexity and obstacle information;
[0150] The preset path planning algorithm is used to generate a path planning plan for the area to be decided based on the path search map, agricultural machinery information, irrigation priority and irrigation status of each partition.
[0151] Optionally, the intelligent decision-making system based on the multimodal large model provided in the embodiment of the present application further includes an optimization module, which is specifically used to:
[0152] Obtaining the path optimization target, and optimizing the path planning scheme according to the path optimization target to obtain the target path planning scheme;
[0153] According to the target path planning scheme and the irrigation strategy of each partition, corresponding instructions are generated and sent to the agricultural machinery, so that the agricultural machinery can irrigate the decision area according to the target path planning scheme.
[0154] The present application also provides a storage medium, which stores program instructions. When the program instructions are loaded and executed by a processor, any of the above-mentioned embodiments of the intelligent decision-making method based on a multimodal large model is implemented.
[0155] The present application also provides an electronic device, such as Figure 3 As shown, the device includes a processor 301 and a memory 302, and the processor and the memory are connected via a bus; program instructions are stored in the memory; the processor calls the program instructions in the memory to execute any of the above-mentioned embodiments of the intelligent decision-making method based on the multimodal large model.
[0156] The processor in this article can be the CPU of the terminal, or the MCU integrated in the terminal, or a combination of the CPU and MCU. In addition, the processor includes a core, which calls the corresponding program from the memory. There can be one or more cores.
[0157] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0158] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.
[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0160] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
[0161] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An intelligent decision-making method based on a multimodal large model, characterized in that: The method comprises: Collecting a current field image of the area to be decided, and preprocessing the current field image to obtain a target field image; Inputting the target field image into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to make a prediction and obtain a pest and disease diagnosis plan corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data; Dividing the area to be decided into a plurality of partitions, and obtaining current multimodal data of each of the partitions; For each partition, processing current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality; Inputting the spatiotemporal sequences of each modality of the partition into a pre-trained irrigation decision model, causing the irrigation decision model to make predictions based on the spatiotemporal sequences of each modality to obtain irrigation data for the partition; wherein the irrigation decision model is obtained by training the irrigation decision model to be trained using historical multimodal data; An irrigation strategy for each of the subareas is generated according to the pest and disease diagnosis plan and the irrigation data of each of the subareas.
2. The method according to claim 1, characterized in that The method of training the visual-text alignment model to be trained using historical multi-source data to obtain the visual-text alignment model includes: Collecting historical multi-source data and preprocessing the historical multi-source data to obtain target multi-source data, wherein the target multi-source data includes historical field image data, historical ground data, and text data related to crops in the field; Annotating the pest and disease type corresponding to each historical field image in the historical field image data according to the historical ground data and the text data, and obtaining a target historical field image corresponding to each historical field image; Access to multiple agricultural literature; Generate a plurality of image-text pairs based on each of the target historical field images and each of the agricultural documents through a visual-text alignment model to be trained; wherein the image-text pairs include the target historical field images and their matching agricultural documents; Performing alignment processing on each of the image-text pairs using a visual-text alignment model to be trained to obtain a target image-text pair corresponding to each of the image-text pairs; The visual-text alignment model to be trained is trained using each of the target image-text pairs and preset agronomic knowledge, and the trained visual-text alignment model is fine-tuned to obtain the visual-text alignment model.
3. The method according to claim 2, characterized in that The step of labeling the pest and disease type corresponding to each historical field image in the historical field image data according to the historical ground data and the text data, and obtaining a target historical field image corresponding to each historical field image, includes: For each historical field image in the historical field image data, perform image recognition on the historical field image to determine pest and disease information and pest and disease areas within the historical field image; Determining the type of pests and diseases corresponding to the historical field image by analyzing the pest and disease information, the pest and disease atlas, the pest and disease identification guide, the historical pest and disease database, and the historical ground data; wherein the text data includes the pest and disease atlas, the pest and disease identification guide, and the historical pest and disease database; The pest and disease type is marked on the pest and disease area of the historical field image using a preset marking method to obtain a target historical field image.
4. The method according to claim 2, characterized in that The step of aligning each of the image-text pairs using the visual-text alignment model to be trained to obtain a target image-text pair corresponding to each of the image-text pairs includes: For each of the image-text pairs, randomly masking the agricultural documents in the image-text pair using the visual-text alignment model to be trained, so as to randomly mask part of the text words in the agricultural documents; A local area in the target historical field image in the image-text pair is randomly selected through the visual-text alignment model to be trained, and the local area is aligned with the keywords in the agricultural literature after masking to obtain the corresponding target image-text pair.
5. The method according to claim 1, wherein For each partition, processing the current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality includes: For each partition, preprocessing the current multimodal data of the partition to obtain target current multimodal data; wherein the target current multimodal data includes target current data of multiple modalities; For each of the modalities, performing intra-modal encoding on the target current data of the modality to obtain an embedding vector corresponding to the modality; The embedding vectors of each of the modalities are cross-modally fused and encoded to obtain corresponding spatiotemporal sequences.
6. The method according to claim 1, characterized in that The method further comprises: Determining the irrigation priority and irrigation status of each partition according to the irrigation recommendation value in the irrigation data of each partition; wherein the irrigation status is whether irrigation is required or not required; Acquiring map data of each of the sub-areas and agricultural machinery information of the agricultural machinery used for irrigation, and determining terrain complexity and obstacle information of each of the sub-areas based on the map data of each of the sub-areas; generating a path search map based on each of the partitions and their terrain complexity and obstacle information; A preset path planning algorithm is used to generate a path planning solution for the area to be decided based on the path search graph, the agricultural machinery information, the irrigation priority and irrigation status of each partition.
7. The method according to claim 6, characterized in that The method further comprises: Obtaining a path optimization target, and optimizing the path planning scheme according to the path optimization target to obtain a target path planning scheme; Generate corresponding instructions according to the target path planning scheme and the irrigation strategies of each of the partitions, and send the instructions to the agricultural machine, so that the agricultural machine irrigates the area to be decided according to the target path planning scheme.
8. An intelligent decision-making system based on a multimodal large model, characterized in that: The system comprises: The current field image acquisition module is used to acquire the current field image of the area to be decided and pre-process the current field image to obtain the target field image; A pest and disease diagnosis solution prediction module is configured to input the target field image into a pre-trained visual-text alignment model, so that the visual-text alignment model uses the target field image to perform predictions and obtain a pest and disease diagnosis solution corresponding to the area to be decided; wherein the visual-text alignment model is obtained by training the visual-text alignment model to be trained using historical multi-source data by the training module; A multimodal data acquisition module, configured to divide the area to be decided into a plurality of partitions and acquire current multimodal data of each of the partitions; a data processing module, configured to process, for each partition, current data of each modality in the current multimodal data of the partition to obtain a spatiotemporal sequence of each modality; an irrigation data prediction module, configured to input the spatiotemporal sequences of each modality of the partition into a pre-trained irrigation decision model, causing the irrigation decision model to perform predictions based on the spatiotemporal sequences of each modality to obtain irrigation data for the partition; wherein the irrigation decision model is trained using historical multimodal data to train the irrigation decision model to be trained; The irrigation strategy generation module is used to generate an irrigation strategy for each of the partitions according to the pest and disease diagnosis plan and the irrigation data of each of the partitions.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor and the memory are connected via a bus; wherein the processor is configured to call and execute a program stored in the memory; The memory is used to store a program, and the program is used to implement an intelligent decision-making method based on a multimodal large model as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to execute an intelligent decision-making method based on a multimodal large model as described in any one of claims 1-7.
Citation Information
Cited By
Intelligent cherry irrigation decision-making method based on multi-mode and space-time prediction
CN121525882A
A Smart Irrigation Decision-Making Method for Cherries Based on Multimodal and Spatiotemporal Prediction
CN121525882B
Sweet potato disease and insect pest diagnosis, prevention and control system and method based on multi-agent collaboration
CN122265139A