Large-model multi-modal and multi-dimensional data augmentation method and device based on man-machine interaction
By employing a human-computer interactive large-scale model multimodal and multidimensional data augmentation method, this approach utilizes large-scale language models and CLIP models in the agricultural field, combined with agricultural knowledge graphs and multiple modality generators, to address the problem of insufficient multimodal data in smart agriculture. This generates high-quality, semantically consistent augmented data, thereby improving the model's generalization ability and decision-making accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING RES CENT FOR INFORMATION TECH & AGRI
- Filing Date
- 2025-12-05
- Publication Date
- 2026-05-08
AI Technical Summary
In the field of smart agriculture, there is a lack of multimodal data collection, poor data quality, and insufficient diversity. Traditional data augmentation methods cannot effectively link multidimensional data, and the generated data is out of touch with the actual scenario, affecting the generalization ability of the model and the accuracy of decision-making.
We employ a large-scale, multimodal, and multidimensional data augmentation method based on human-computer interaction. By combining large-scale language models, CLIP models, and agricultural knowledge graphs in the agricultural field with data generators of various modalities, we perform data preprocessing, feature extraction, cross-modal consistency verification, and fusion to generate high-quality augmented data that meets agronomic requirements.
The generated multimodal data is semantically consistent and agronomically sound, which improves the generalization ability and decision-making accuracy of downstream agricultural large models and solves the problems of insufficient data quantity, poor quality and lack of diversity.
Smart Images

Figure CN121996338A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for augmenting large-scale multimodal and multidimensional data based on human-computer interaction. Background Technology
[0002] In the field of smart agriculture, there are significant bottlenecks in the collection and application of multimodal data. Regarding data scale, data on the entire crop growth cycle is severely insufficient. Taking cabbage as an example, effective samples at key stages such as the early heading and expansion stages are scarce, with single experimental fields often containing fewer than a thousand samples, making it difficult to support large-scale deep learning models. In terms of data quality, field sensors are easily interfered with; for example, soil moisture sensors may exhibit abrupt changes due to mud and water adhesion, and drone images may be overexposed due to backlighting. This type of noise reduces the model's recognition accuracy. Data diversity is also significantly insufficient. Crop phenotypic data under extreme climates are scarce, and the growth characteristics of different varieties are unevenly covered, resulting in weak generalization ability of models in complex field scenarios and difficulty in accurately identifying abnormal states caused by rare pests and diseases.
[0003] Traditional data augmentation methods suffer from significant limitations in agricultural scenarios. Their single-modal processing logic struggles to handle the multi-dimensional relationships inherent in agricultural data. For instance, while random cropping and rotation of crop images can increase the number of image samples, they fail to correlate with key environmental factors such as soil moisture and meteorological data, leading to a disconnect between the augmented data and actual conditions. In cross-modal processing, spatiotemporal alignment accuracy is low; for example, the time error in matching meteorological text with remote sensing images often exceeds one hour, and the spatial deviation reaches several meters, failing to meet the real-time linkage requirements for disaster early warning. These limitations make it difficult for augmented data to support large models learning the complex relationships between crop growth, environment, and agricultural activities, impacting the performance of tasks such as precision irrigation and pest and disease prediction.
[0004] Furthermore, traditional augmentation processes lack the effective integration of agronomic expert knowledge, leading to significant deviations between generated data and actual field conditions. For example, it may generate false data such as "cabbage grows rapidly under high winter temperatures," ignoring its optimal growth temperature (15–20℃); or generate incorrect samples such as "aphids gather on cabbage heads," contradicting its biological characteristic of primarily parasitizing the underside of leaves. Such "pseudo-data" can mislead the model into forming incorrect perceptions, such as misjudging slow growth under normal low temperatures as a pathological condition. Agronomic experience (such as "shallow cultivation during the seedling stage and water control during the heading stage") is also difficult for machines to capture through autonomous learning, resulting in conflicts between generated agricultural operation data and crop growth stages, ultimately reducing the practical application value of smart agriculture technologies. Summary of the Invention
[0005] This invention provides a method and apparatus for augmenting large-scale multimodal and multidimensional data based on human-computer interaction, in order to solve the above-mentioned problems and achieve the effective generation of high-quality multimodal augmented data that is semantically consistent and agronomically reasonable.
[0006] This invention provides a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction, comprising the following steps: Multimodal and multidimensional agricultural data from different data sources are collected, preprocessed and formatted to obtain processed agricultural data, and multimodal and multidimensional agricultural data are obtained through a human-computer interaction interface to obtain multimodal demand commands input by users. Based on the multimodal demand instructions and the processed agricultural data, the data augmentation model is invoked to obtain augmented agricultural data; The augmented agricultural data is subjected to cross-modal consistency verification, and the verified augmented agricultural data is fused with the processed agricultural data to obtain the augmented agricultural dataset. The data augmentation model includes a large-scale language model for agriculture, a contrastive language-image pre-trained CLIP model for agriculture, an agricultural knowledge graph, and data generators corresponding to multiple modalities; the data generators corresponding to multiple modalities are associated through an attention mechanism that shares a semantic space.
[0007] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. Based on the multimodal demand instruction and the processed agricultural data, a data augmentation model is invoked to obtain augmented agricultural data, including: By using a large-scale language model, a CLIP model, and an agricultural knowledge graph in the agricultural field, features are extracted and fused from the multimodal demand instructions to obtain structured constraints, and the target modality type in the multimodal demand instructions is determined. The processed agricultural data is encoded into multimodal feature vectors, and the structured constraints are encoded into target constraint boundaries. Based on the multimodal feature vectors and the target constraint boundaries, augmented agricultural data is generated by the data generator corresponding to the target modality type.
[0008] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, wherein the multiple modal types include one or more of image data types, time-series data types, and text data types; Among them, the data generator corresponding to the image data type consists of the Generative Adversarial Network (GAN) and the CLIP model; The data generator corresponding to the time series data type consists of a Transformer time series model and a physical constraint rule library; The data generator corresponding to the text data type consists of a large language model that has been fine-tuned with knowledge from the agricultural field.
[0009] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. This method involves extracting and fusing features from the multimodal demand instructions using a large-scale language model, a CLIP model, and an agricultural knowledge graph to obtain structured constraints. The method includes: By using a large-scale language model in the agricultural field, features are extracted from the text instructions in the multimodal demand instructions to obtain a structured entity list and entity pair relationship types; The image instructions in the multimodal demand instructions are extracted using the CLIP model in the agricultural field to obtain a quantitative phenotypic feature set and a feature credibility report; Based on the agricultural knowledge graph, the structured entity list, the entity pair relationship type, the quantitative phenotypic feature set, and the feature credibility report are fused to obtain the fusion result; The fusion result is logically verified by calling the agricultural expert rule base to obtain the verification result; Based on the constraint generation algorithm, the fusion result and the logical verification result are transformed to obtain structured constraints.
[0010] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. This method involves extracting features from textual instructions in multimodal demand instructions using a large-scale language model in the agricultural field to obtain a structured entity list and entity pair relationship types, including: Based on agricultural corpus, the agricultural-specific large language model AgriLLM is fine-tuned to obtain the large language model for the agricultural field. The semantic features of the text instructions are obtained through a large-scale language model in the agricultural field. The semantic features of the text instructions are input into an entity recognition model based on a Conditional Random Field (CRF) to obtain a structured entity list. Input the structured entity list into the dual Transformer model to obtain the initial relationship type and entity association probability of each entity pair in the structured entity list; The initial relationship type is filtered by the entity association probability to obtain the final entity pair relationship type.
[0011] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. The method involves extracting features from image instructions in the multimodal demand instructions using a CLIP model in the agricultural field to obtain a quantified phenotypic feature set and a feature credibility report, including: The region of interest is extracted from the original image in the image instruction, and the extracted image is standardized by bicubic interpolation algorithm to obtain a standardized image. The standardized image is input into the CLIP model of the agricultural field, which has been fine-tuned for agricultural scenarios, and matched with the preset agricultural text labels to output key phenotypic data and form an initial quantitative phenotypic feature set. The original image is scored using a ResNet-50-based image quality assessment model, and the feature confidence level is obtained based on the scoring results. The initial quantified phenotypic feature set is verified and filtered according to the preset decision rules and the feature credibility to obtain the final quantified phenotypic feature set and feature credibility report.
[0012] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, which extracts the region of interest from the original image in the image instruction, including: For any candidate threshold within the candidate grayscale search range, calculate the inter-class variance between the foreground and background corresponding to the arbitrary candidate threshold; Compare the inter-class variances corresponding to all candidate thresholds within the search interval, select the candidate threshold that maximizes the inter-class variance as the optimal segmentation threshold, and use the optimal segmentation threshold to perform binarization segmentation on the original image to obtain the region of interest. The candidate grayscale search interval is set based on prior knowledge of the grayscale distribution of the target and background in agricultural images.
[0013] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, wherein the method for cross-modal consistency verification of the augmented agricultural data includes: The semantic similarity between image data and text data in the augmented agricultural data is calculated using the CLIP model and compared with a preset threshold; and / or, based on an agricultural knowledge graph, the logical correlation between time-series data and image data in the augmented agricultural data is verified.
[0014] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, the method further comprising: The system validation metrics for the augmented agricultural dataset are obtained, and user feedback ratings for the augmented agricultural dataset are obtained through a human-computer interaction interface. The user feedback rating and the system verification index are merged into a reward signal; and optimized training samples are generated based on the augmented agricultural dataset with a preset number of interactions as the iteration cycle. The parameters of the data augmentation model are optimized based on the reward signal and the optimized training samples using a proximal policy optimization algorithm.
[0015] The present invention also provides a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction, comprising the following modules: The data acquisition and processing module is used to acquire multimodal and multidimensional agricultural data from different data sources, preprocess and format the multimodal and multidimensional agricultural data to obtain processed agricultural data, and obtain multimodal demand commands input by users through a human-computer interaction interface. The data augmentation module is used to call the data augmentation model based on the multimodal demand command and the processed agricultural data to obtain augmented agricultural data; The verification and fusion module is used to perform cross-modal consistency verification on the augmented agricultural data, and to fuse the verified augmented agricultural data with the processed agricultural data to obtain the augmented agricultural dataset. The data augmentation model includes a large-scale language model for agriculture, a contrastive language-image pre-trained CLIP model for agriculture, an agricultural knowledge graph, and data generators corresponding to multiple modalities; the data generators corresponding to multiple modalities are associated through an attention mechanism that shares a semantic space.
[0016] The present invention provides a method and apparatus for augmenting large-scale multimodal and multidimensional data based on human-computer interaction. By introducing a human-computer interaction interface to obtain multimodal demand instructions input by users, a data augmentation model driven by a large agricultural model and in which multiple modality generators work collaboratively through an attention mechanism sharing a semantic space is constructed. This effectively generates high-quality multimodal augmented data that is semantically consistent and agronomically reasonable. It solves the problems of insufficient quantity, poor quality, lack of diversity, and disconnect from real-world scenarios in agricultural scenarios, and can improve the generalization ability and decision-making accuracy of downstream agricultural large-scale models. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the large-scale multimodal and multidimensional data augmentation method based on human-computer interaction provided by the present invention.
[0019] Figure 2 This invention provides a large-scale multimodal and multidimensional data augmentation system based on human-computer interaction.
[0020] Figure 3This is a schematic diagram of the structure of the large model dynamic generation and processing module provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the structure of the large-scale multimodal and multidimensional data augmentation device based on human-computer interaction provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Figure 1 This is a flowchart illustrating the large-scale multimodal and multidimensional data augmentation method based on human-computer interaction provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 100: Collect multimodal and multidimensional agricultural data from different data sources, preprocess and format the multimodal and multidimensional agricultural data to obtain processed agricultural data; and obtain multimodal demand instructions input by users through the human-computer interaction interface.
[0024] Step 101: Based on the multimodal demand instructions and the processed agricultural data, call the data augmentation model to obtain augmented agricultural data.
[0025] Step 102: Perform cross-modal consistency verification on the augmented agricultural data, and merge the verified augmented agricultural data with the processed agricultural data to obtain the augmented agricultural dataset.
[0026] Among them, the data augmentation models include large-scale language models in the agricultural field, contrastive language-image pre-training (CLIP) models in the agricultural field, agricultural knowledge graphs, and data generators corresponding to multiple modalities; the data generators corresponding to multiple modalities are associated through an attention mechanism that shares a semantic space.
[0027] Specifically, the execution entity of the large-scale multimodal and multidimensional data augmentation method based on human-computer interaction provided by the present invention can be a system composed of multiple sensors and computing processing devices with human-computer interaction interfaces.
[0028] First, the system initiates a multimodal and multidimensional data acquisition and preprocessing process. It can comprehensively collect multimodal and multidimensional agricultural data related to environmental monitoring, crop phenotyping, and farmland management through various methods such as IoT sensor grids deployed in the field, handheld phenotyping devices, multispectral drones, and manual mobile recording.
[0029] In some implementations, each sample in the multimodal, multidimensional agricultural data includes image data of the same plant, soil temperature and humidity data, environmental data, and agricultural log records.
[0030] In some implementations, multimodal and multidimensional agricultural data includes multimodal and multidimensional data collected at different times and frequencies during different stages of the crop growth cycle.
[0031] In some implementations, IoT sensors can be deployed in cabbage fields in a 50m×50m grid to collect real-time data on soil moisture, air temperature and humidity (e.g., every 15 minutes). Field equipment such as handheld phenotyping instruments can be used to measure plant height and stem diameter weekly, and multispectral drones can be used to acquire canopy images to extract leaf area index and vegetation cover. Simultaneously, manual recording can be combined, such as growers entering agricultural operations (e.g., fertilization time and amount) via mobile apps, and technicians photographing and labeling pest and disease samples (e.g., "2024-07-10, Plot A, Downy mildew level 3"). Furthermore, historical data can be integrated, including cabbage planting records for the target area over the past three years, monitoring data from agricultural technology extension stations, and historical climate data from meteorological stations. These data collectively constitute multimodal and multidimensional agricultural data, including environmental monitoring data (such as air temperature, humidity, light intensity, precipitation, wind speed, soil temperature, soil moisture, pH value, and available nutrients), crop phenotypic data (such as plant height, stem diameter, number of leaves, leaf area index, head diameter, SPAD value, pest and disease severity, and wilting degree), and farmland management data (such as sowing date, fertilizer type and amount, irrigation records, pesticide spraying, sensor deployment location, and drone inspection time).
[0032] Then, the collected multimodal and multidimensional agricultural data can be preprocessed and formatted to obtain processed agricultural data.
[0033] In some implementations, preprocessing and formatting can be done by cleaning data, such as removing outliers caused by sensor malfunctions (e.g., air temperature > 40°C) and filling in short-term missing values (e.g., data missing for 1-2 hours) using linear interpolation. Preprocessing and formatting can also be done by standardization, such as normalizing data in different units (e.g., mapping temperature to the 0-1 range) and unifying the time granularity (summarizing daily averages by "day").
[0034] In some implementations, the structured transformation can include converting unstructured data such as images and text into structured indicators. For example, drone images can be converted into structured indicators such as leaf area index, ultimately forming a three-dimensional data table of "time-plot-indicator" (example: plot B, 2024-06-15, air temperature 26℃, plant height 22cm, fertilizer application rate 3kg / mu). Regarding formatting, image data can be converted to JPEG format and named according to "plot ID-collection time-camera number"; point cloud data can be compressed into LAS format, retaining XYZ coordinates and reflection intensity fields; and text logs can be parsed into JSON format, extracting structured information such as "agricultural operation type-time-executor".
[0035] In some implementations, multimodal user input instructions can be obtained through a human-computer interaction interface. These instructions can include text and image commands. The interface can be divided into a visual interactive interface and a natural language understanding unit, supporting users to input constraints in various ways. For example, users can perform semantic annotations on images, selecting key areas (such as "cabbage head") and associating them with text descriptions ("a diameter of 5-8cm indicates maturity"); they can also set rules through drop-down menus, selecting spatiotemporal constraints (such as "augmented data must meet the requirement that when soil moisture > 60%, the leaf spread of the crop image > 80%").
[0036] Based on the multimodal demand instructions and the processed agricultural data, the data augmentation model can be invoked to obtain augmented agricultural data.
[0037] This data augmentation model can include large language models (LLMs) in the agricultural field, CLIP models in the agricultural field, agricultural knowledge graphs, and data generators corresponding to various modalities.
[0038] Among them, large-scale language models can be used to understand textual instructions input by users and parse them into entities and entity relationships; CLIP models can be used to understand image instructions input by users and parse them into quantified phenotypic features; agricultural knowledge graphs can be used to perform bidirectional mapping between entities and entity relationships obtained from textual instructions and quantified phenotypic features obtained from image instructions, and determine whether there are conflicts.
[0039] Data generators for various modalities can be responsible for generating different forms of data, such as images, time series (e.g., temperature and humidity changes), text logs (e.g., agricultural records), or point clouds. For example, an image data generator can perform style transfer or content addition on existing cabbage canopy images (e.g., simulating different disease spots), while a time series data generator can interpolate or extrapolate data such as air temperature and soil moisture to generate more dimensional or finer-grained time series samples.
[0040] Data generators corresponding to multiple modalities can be linked through an attention mechanism that shares a semantic space. The generation processes of different modalities are not carried out in isolation, but rather achieve information interaction and consistency constraints through a unified semantic representation space. For example, when an image generator generates the feature "leaf slightly curled", it passes a signal to the time series generator "soil moisture needs to be associated" through an attention mechanism, thereby ensuring that the generated data is logically consistent.
[0041] In some implementations, diversity control and constraint verification can be performed during the data generator's generation process. For example, by inputting a scenario label such as "drought year," the data generator can be guided to generate data showing that the average soil moisture is 20% lower than in a normal year; simultaneously, by matching a rule base in real time, samples that violate constraints can be filtered out, such as removing unreasonable data like "heading temperature <10℃." In this way, high-quality augmentation data that meets both user needs and agricultural principles can be generated.
[0042] Then, cross-modal consistency verification can be performed on the augmented agricultural data generated by the data augmentation model to ensure semantic and logical consistency between different modalities. For example, if the generated image shows that cabbage is in the heading stage and the leaves are dark green, the corresponding text log should contain reasonable agricultural records (such as "topdressing with nitrogen fertilizer during the heading stage"), the SPAD value in the time series data should be no less than 30, and the available nitrogen content in the soil should also be at a high level. If the image shows severe wilting, but the soil moisture data shows that the field water holding capacity is 75%, there may be a modal conflict. Such inconsistent samples can be identified and removed through rule base matching, expert rules, or model-driven verification mechanisms, thereby ensuring the rationality and credibility of the augmented agricultural data.
[0043] Finally, the validated augmented agricultural data can be merged with the processed agricultural data to obtain the augmented agricultural dataset. This fusion process can be carried out according to a preset ratio, such as using an integration strategy of "real data: generated data = 1:2" (e.g., combining 10,000 real data records with 20,000 generated data records) to significantly expand the sample size while maintaining data authenticity.
[0044] In some implementations, the augmented agricultural dataset can be divided according to a predetermined ratio, such as dividing it into a training set, validation set, and test set in an 8:1:1 ratio. Simultaneously, all fields can be standardized, unifying the format of fields such as time and temperature, ensuring that the dataset contains comprehensive information across all dimensions, including "time-plot-growth stage-environmental indicators-crop indicators-management measures," thus forming a structured data table.
[0045] In some implementations, the augmented agricultural dataset can be provided in various formats, including common data formats such as CSV, JSON, and Parquet, along with detailed metadata descriptions, such as the meaning of indicators and constraint rules. Simultaneously, quality reports containing validation results, such as distribution comparison charts and expert evaluation opinions, can be generated to demonstrate the reliability and usability of the dataset.
[0046] In some implementations, dedicated API interfaces can be developed to support flexible data access. For example, users can filter data by criteria such as "growth stage" or "indicator type," and perform specific queries such as "get soil nitrogen content data during the rosette stage," making it convenient for downstream applications such as precision irrigation systems to directly access and use the data.
[0047] The present invention provides a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction. By introducing a human-computer interaction interface to obtain multimodal demand instructions input by users, a data augmentation model driven by a large agricultural model and in which multiple modality generators work collaboratively through an attention mechanism sharing a semantic space is constructed. This effectively generates high-quality multimodal augmented data that is semantically consistent and agronomically reasonable. It solves the problems of insufficient quantity, poor quality, lack of diversity, and disconnect from real-world scenarios in agricultural scenarios, and can improve the generalization ability and decision-making accuracy of downstream agricultural large models.
[0048] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. Based on multimodal demand instructions and processed agricultural data, a data augmentation model is invoked to obtain augmented agricultural data, including: By using a large-scale language model, a CLIP model, and an agricultural knowledge graph, features are extracted and fused from multimodal demand instructions to obtain structured constraints; and the target modality type in the multimodal demand instructions is determined. The processed agricultural data is encoded into multimodal feature vectors, and the structured constraints are encoded into target constraint boundaries. Augmented agricultural data is generated based on multimodal feature vectors and target constraint boundaries through a data generator corresponding to the target modality type.
[0049] Specifically, in this embodiment, the text instructions in the multimodal demand instructions input by the user can first be parsed using a large-scale language model in the agricultural field, the image instructions in the multimodal demand instructions input by the user can be parsed using a CLIP model in the agricultural field, and the text instructions and image instructions can be fused using an agricultural knowledge graph to obtain the structured constraints of the multimodal demand instructions.
[0050] Meanwhile, during the process of parsing the text instructions in the multimodal requirement instructions input by the user, the target modality type indicated by the multimodal requirement instructions can be determined.
[0051] In some implementations, large-scale language models in the agricultural field can be used to break down user instructions into "modal types" (such as images, time series, text), "core elements" (such as crop = cabbage, growth stage = heading stage) and "attribute constraints" (such as leaf yellow spot ratio <10%, soil moisture 50-70%).
[0052] In some implementations, the constraints indicated by the multimodal demand directive may include one or more of the following: (1) Phenological period constraints are used to limit the time boundaries and phenotypic index ranges of each growth stage of crops.
[0053] Specifically, the constraints indicated by multimodal demand instructions can include phenological constraints, which are used to limit the time boundaries and phenotypic index ranges of each growth stage of the crop. For example, the time boundaries of each growth stage of cabbage can be clearly defined, such as 30-40 days for the seedling stage, 25-30 days for the rosette stage, and 40-50 days for the heading stage, while limiting the index ranges for each stage, such as the plant height during the heading stage not being less than 40cm.
[0054] (2) Physiological constraints are used to define the correlation between different agricultural data indicators and to exclude combinations of indicators that violate crop physiological laws.
[0055] Specifically, the constraints indicated by multimodal demand instructions can include physiological constraints, used to define the correlation between different agricultural data indicators and to exclude indicator combinations that violate crop physiological laws. For example, a positive correlation rule can be set for "when soil nitrogen content > 50 mg / kg, SPAD value ≥ 35". At the same time, unreasonable combinations can be excluded. For example, if an image or record of "no disease spots on leaves" is generated under the condition of "air humidity > 90% and in the heading stage", it may violate common agronomical knowledge.
[0056] (3) Scenario constraints are used to limit the scope of scenarios for expanding the coverage of agricultural data.
[0057] Specifically, the constraints indicated by the multimodal demand instructions can include scenario constraints to limit the scope of planting scenarios covered by augmented agricultural data, so that the generated content is regionally and climatically representative. For example, for "spring cabbage in North China", the augmented data should include the scenario of late spring cold snap (temperature 0–5℃); while for "autumn cabbage in South China", it needs to cover high-temperature environments (30–35℃).
[0058] By structuring one or more of the above constraints, we obtain the structured constraints of the multimodal demand instructions. In some implementations, a "constraint rule base" can be established to store rules in the form of "IF-THEN", such as: "IF growth stage = seedling stage THEN plant height ∈ [5cm, 15cm]". At the same time, the system can develop a visual interactive interface, allowing experts to intuitively adjust parameters through sliders and other controls (such as dragging the upper limit of the heading stage temperature to 25℃). The system will automatically convert such operations into mathematical expressions that the model can recognize (such as temperature constraint: T≤25℃).
[0059] In some implementations, large-scale language models in the agricultural field can generate a unified task descriptor for the semantic parsing process. For example, a JSON format can be used to include the generation objectives of each modality (such as image resolution of 1920×1080 and time-series sampling frequency of 1 time / hour) and association rules (such as yellow spots on leaves in the image → the text must contain "may be nitrogen deficient"), thereby achieving a standardized representation of the instructions.
[0060] Then, the processed agricultural data can be encoded into multimodal feature vectors, while the obtained structured constraints can be transformed into target constraint boundaries.
[0061] In some implementations, the process of encoding processed agricultural data into multimodal feature vectors can be aided by an encoder that extracts features from different types of data. For example, cabbage canopy images can be encoded into visual feature vectors to extract phenotypic information such as leaf morphology and color; time-series data such as soil moisture can be encoded into temporal feature vectors to capture their diurnal fluctuation patterns; and agricultural texts can be encoded into semantic vectors to interpret the agronomic meanings of operations such as "irrigation" and "fertilization." These multimodal feature vectors collectively constitute a structured representation of real agricultural scenarios, providing contextual support for the data generator.
[0062] In some implementations, by transforming structured constraints into target constraint boundaries, attribute constraints contained in the task descriptor (such as "soil moisture 50–70%)" can be mapped to constraint boundaries in the feature space as limitations in the generation process. For example, when generating time-series soil moisture data, the model output value can be constrained to the range of 50% to 70% field capacity.
[0063] Finally, augmented agricultural data can be generated using data generators corresponding to the target modality type, based on multimodal feature vectors and target constraint boundaries. For example, if the target modality is an image, an image data generator is invoked to generate new samples that meet the requirements, guided by real image features and constraint boundaries (such as light intensity and leaf health). If the target modality is a time series, a time series data generator is invoked to perform interpolation or extrapolation within specified constraints, based on the fluctuation patterns of the original time series. If the target modality is text, a text data generator is invoked to generate log records that conform to agricultural logic. All generators work collaboratively under an attention mechanism sharing a semantic space, ensuring semantic consistency between different modal outputs, ultimately forming high-quality, multi-dimensional, and multimodal augmented agricultural data.
[0064] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, wherein the multiple modal types include one or more of image data types, time-series data types, and text data types; Among them, the data generator corresponding to the image data type consists of a Generative Adversarial Network (GAN) and a CLIP model; The data generator corresponding to the time series data type consists of a Transformer time series model and a physical constraint rule library; The data generator corresponding to the text data type consists of a large language model that has been fine-tuned with knowledge from the agricultural field.
[0065] Specifically, the various modal types in the embodiments of this application may include one or more of image data types, time-series data types, and text data types.
[0066] In some implementations, for image data types, the corresponding data generator can consist of GAN and CLIP models. For example, a Constrained Generative Adversarial Network (ConstrainedGAN) can be used, which includes a generator to output "pseudo-data" and a discriminator to distinguish between "real data" and "pseudo-data." Simultaneously, the CLIP model can be used to align textual semantics with image content, guiding the data generator to produce semantically consistent visual samples. Furthermore, to ensure agricultural rationality, a "constraint penalty term" can be added to the generator's loss function. For example, if the generated image corresponds to the seedling stage but shows head formation, it is considered a violation of the cabbage's physiological laws, and the system will increase the loss value, forcing the generator to adjust its output to conform to the characteristics of the actual growth stage.
[0067] In some implementations, for time-series data types, the corresponding data generator can consist of a Transformer time-series model and a physical constraint rule base. This generator can leverage the sequence modeling capabilities of the Transformer to capture long-term dependencies and can also incorporate Long Short-Term Memory (LSTM) layers to better capture short-term temporal correlations, such as analyzing humidity trends over seven consecutive days. Simultaneously, the physical constraint rule base can embed agronomical knowledge about crop growth, such as "daily average plant height growth of cabbage ≤ 2cm." The physical constraint rule base allows for real-time validation of the generated data, filtering out samples that violate agricultural principles.
[0068] In some implementations, for text-based data types, the corresponding data generator can consist of a large-scale language model fine-tuned with agricultural knowledge. Based on a general-purpose language model, this model can be fine-tuned using extensive agricultural corpora such as crop planting logs, agricultural manuals, and expert Q&As, enabling it to understand and generate text content consistent with agronomic contexts. This fine-tuning allows the large-scale language model not only to generate grammatically correct sentences but also to accurately reflect the causal relationships between operations such as fertilization, irrigation, and disease control and crop status.
[0069] The aforementioned data generators of different modalities can be dynamically invoked through a policy routing mechanism: when the system identifies the target modality as an image data type, it invokes the "GAN+CLIP large model" combination; when the target modality is a time-series data type, it activates the "Transformer time-series model + physical constraint rule library"; and when the target modality is a text data type, it triggers the "instruction fine-tuning + LLM" process. These data generators perform cross-modal adaptation through an attention mechanism that shares a semantic space. For example, when the image data generator outputs an image containing the feature of "yellow spots on leaves," it can automatically send a signal to the time-series data generator that "the soil nitrogen content needs to be correlated," while simultaneously triggering the text data generator to output descriptions such as "potential nitrogen deficiency, nitrogen fertilizer supplementation is recommended," thereby ensuring that the image, time-series, and text are highly consistent in semantics and logic, forming structurally complete and multi-dimensional collaborative augmented agricultural data.
[0070] According to the present invention, a large-scale multimodal and multidimensional data augmentation method based on human-computer interaction is provided. This method extracts and fuses features from multimodal demand instructions using a large-scale language model, a CLIP model, and an agricultural knowledge graph, to obtain structured constraints, including: By using a large-scale language model in the agricultural field, features are extracted from the text instructions in multimodal demand instructions to obtain a structured entity list and entity pair relationship types; The CLIP model in the agricultural field is used to extract features from image instructions in multimodal demand instructions, resulting in a quantitative phenotypic feature set and a feature credibility report. Based on agricultural knowledge graphs, the structured entity list, entity pair relationship types, quantitative phenotypic feature sets, and feature credibility reports are fused to obtain the fusion result; The agricultural expert rule base is invoked to perform logical verification on the fusion results, and the verification results are obtained. Based on the constraint generation algorithm, the fusion result and the logical verification result are transformed to obtain structured constraints.
[0071] Specifically, in this embodiment of the invention, features are first extracted from the text instructions in the multimodal demand instructions using a large-scale language model in the agricultural field to obtain a structured entity list and entity pair relationship types.
[0072] In some implementations, features are extracted from textual instructions in multimodal demand instructions using a large-scale language model in the agricultural field to obtain a structured entity list and entity pair relationship types, including: Based on agricultural corpus, the agricultural-specific large language model AgriLLM is fine-tuned to obtain a large language model for the agricultural field. Semantic features of text instructions are obtained through large-scale language models in the agricultural field; The semantic features of the text instructions are input into an entity recognition model based on a Conditional Random Field (CRF) to obtain a structured list of entities. Input the structured entity list into the dual Transformer model to obtain the initial relationship type and entity association probability of each entity pair in the structured entity list; The initial relation types are filtered by the entity association probability to obtain the final entity pair relation types.
[0073] Specifically, the agricultural-specific large language model (AgriLLM) can be fine-tuned based on agricultural corpus to obtain a large-scale language model for the agricultural field. The agricultural corpus is derived from knowledge in fields such as crop cultivation, plant protection, and soil and fertilizer science.
[0074] In some implementations, the fine-tuning of the AgriLLM agricultural language model can employ LoRA (Low-Rank Adaptation) technology. This involves freezing 95% of the parameters of the LLaMA-2 base model and training only the adapter layer to ensure the model can accurately recognize agricultural language. The specific execution flow for fine-tuning using LoRA is as follows: 1. Initialization: Freeze all parameters of the pre-trained model ( Randomly initialize low-rank matrices (Usually using normal distribution) , Need to be based on Adjustments, such as , low-rank matrix Initialize as an all-zero matrix (ensuring initialization) The initial output of the model is consistent with the pre-trained model, avoiding fluctuations in the early stages of training. 2. Forward Propagation: Input Features After pre-training matrix Get the basic output: The input features are simultaneously processed and Get incremental output: The final output of this layer is the superposition of the basic output and the incremental output (a scaling factor can be optionally added). (Balancing incremental weights) in It is a hyperparameter (usually set to ) ,like =16 =16), its function is to let The magnitude and Matching is used to avoid incremental outputs being too small or too large; 3. Backpropagation: Only calculate and gradient ( The gradient is frozen and not updated; updates are performed via gradient descent. and The parameters are adjusted until the model converges on the task; 4. Reasoning stage: This can be... The increment is directly added to superior( No additional calculations are required during reasoning. and The reasoning process is completely consistent with the original model.
[0075] The semantic features of text instructions can be obtained through large-scale language models in the agricultural field.
[0076] Subsequently, a CRF-based entity recognition model can be used to identify modal type entities, core element entities, attribute constraint entities, etc. in the text, and obtain a structured entity list.
[0077] The structured entity list is then input into the dual Transformer model. The dual Transformer model extracts relations to obtain the initial relation type and entity association probability of each entity pair in the structured entity list. The initial relation type is then filtered by the entity association probability to obtain the final entity pair relation type.
[0078] The following are examples of extraction steps in specific application scenarios: 1. First layer: Accurately identify key entities in the instructions. The original command was segmented using the Jieba extended agricultural terminology dictionary. Based on the preset "BIO + agricultural entity category" tagging system, each segmentation result was labeled with agricultural domain knowledge. The segmentation and labeling of "cabbage in the heading stage has aphids" is shown in Table 1. According to the segmentation order, the complete labeling sequence is [B-STAGE, I-STAGE, B-CROP, B-ORG, O, B-STRESS].
[0079] Table 1 Examples of word segmentation annotation
[0080] According to the BIO annotation rules, the annotation sequences are merged into entities, as shown in Table 2: Table 2 Example of Label Merging
[0081] The merged complete entities are stored in a structured format of "entity ID + entity content + entity category" to form an "entity set" that can be called by subsequent relation modeling branches. The final output is shown in Table 3. Table 3 Entity Sets
[0082] 2. Second layer: Modeling the potential relationships in entities First, the relation modeling branch receives the initial entity set output by the entity recognition branch. For the agricultural instruction "cabbage leaves have aphids during the heading stage", it will obtain entities with category information such as "heading stage (growth stage)", "cabbage (crop)", "leaf (organ)" and "aphid (stress)". Then, with the help of the Transformer's self-attention mechanism, each entity is embedded into a high-dimensional semantic space.
[0083] The system then proceeds to entity pair relationship modeling and classification. First, for the received entity set, the system generates all possible entity pairs, such as <cabbage, heading stage>, <cabbage leaves, aphids>, and <cabbage, soil nitrogen content>. Next, for each entity pair, the model models a "relationship type" based on agricultural characteristics. By learning from a large amount of agricultural instruction data, the model constructs semantic representations and recognition patterns for these relationships. Finally, the system performs a relationship classification task. Predefined relationship categories, such as "growth stage-crop," "stress-organ," and "attribute-entity," conform to agricultural knowledge logic, are used. Based on the embedding features of entity pairs in the high-dimensional semantic space, the model uses algorithms such as the Softmax classifier to predict and classify the relationship type. For example, <cabbage, heading stage> will be accurately classified into the "growth stage-crop" category, clearly defining the relationship between the two.
[0084] 3. Third layer: Calculate the association probability Instruction decomposition needs to avoid "false positive associations," therefore, the entity relationship probabilities of the output dual Transformers are calculated. The association scores of entity pairs are converted into probability values between 0 and 1 using the "softmax function," as shown in the following formula: in, For entities and Existence Relationship The probability, Entity pairs computed for the model < > Correspondence The original score, R For a predefined set of agricultural relationships; set the probability threshold to 0.8, if If the entity relationship is deemed "reliable," it is retained until the structured decomposition result is reached; if If so, it is marked as "suspicious association" and requires further verification using an agricultural knowledge graph.
[0085] Then, the image instructions in the multimodal demand instructions can be feature extracted using the CLIP model in the agricultural field to obtain a quantitative phenotypic feature set and a feature credibility report.
[0086] In some implementations, the image instructions in multimodal demand instructions are feature extracted using the CLIP model in the agricultural field to obtain a quantified phenotypic feature set and a feature confidence report, including: The region of interest is extracted from the original image in the image instruction, and the extracted image is standardized by bicubic interpolation algorithm to obtain a standardized image; The standardized image is input into the CLIP model of the agricultural field, which has been fine-tuned for the agricultural scene, and matched with the preset agricultural text labels to output key phenotypic data and form an initial quantified phenotypic feature set. The original images are scored using a ResNet-50-based image quality assessment model, and the feature credibility is obtained based on the scoring results. The initial quantified phenotypic feature set is verified and filtered according to the preset decision rules and feature credibility to obtain the final quantified phenotypic feature set and feature credibility report.
[0087] Specifically, the image instruction can be preprocessed first, the region of interest can be extracted from the original image in the image instruction, and the extracted image can be standardized by bicubic interpolation algorithm to obtain a standardized image.
[0088] In some implementations, region of interest extraction is performed on the original image in the image instruction, including: For any candidate threshold within the candidate grayscale search range, calculate the inter-class variance between the foreground and background corresponding to any candidate threshold. Compare the inter-class variances corresponding to all candidate thresholds within the search interval, select the candidate threshold that maximizes the inter-class variance as the optimal segmentation threshold, and use the optimal segmentation threshold to perform binarization segmentation on the original image to obtain the region of interest. The candidate grayscale search interval is set based on prior knowledge of the grayscale distribution of the target and background in agricultural images.
[0089] Specifically, in this embodiment of the invention, an improved Otsu thresholding algorithm can be used for ROI extraction. The optimal segmentation threshold is found by calculating the inter-class variance, and then a bicubic interpolation algorithm is used to uniformly scale the image to a standard resolution, resulting in a standardized image. The following are examples of image standardization steps in specific application scenarios: Image quality was optimized using OpenCV image preprocessing tools. First, threshold segmentation was used to remove irrelevant backgrounds such as soil and weeds, retaining only the cabbage leaf as the ROI (region of interest). Then, the image resolution was uniformly adjusted to 1920×1080 to ensure consistency and accuracy in subsequent feature extraction, resulting in a denoised reference image.
[0090] Traditional Otsu's algorithm calculates inter-class variance solely based on the global grayscale histogram of the image, neglecting the scene-specific nature of agricultural images. This leads to incomplete ROI extraction or missegmentation. Addressing the core challenges of agricultural images, traditional Otsu's algorithm requires traversing all grayscale values from 0 to 255 to find the optimal threshold, and it fails to distinguish between the grayscale characteristics of the "agricultural target region" and the "irrelevant background." Specific improvements are as follows: Based on prior knowledge of agricultural images, namely that the gray values of cabbage leaves are usually concentrated in the range of 50-180, and the gray values of soil background are mostly distributed in the range of 20-60 or 200-255, the "target gray value candidate range" for leaves is preset to be 40-200. The threshold search is performed only within a preset range, rather than traversing all 256 grayscale values; at the same time, it excludes the interference of extreme grayscale values (such as dark noise <20, overexposed areas >230) on the threshold calculation.
[0091] The specific implementation process is as follows: 1. An improved Otsu thresholding algorithm is used for ROI extraction to find the threshold that maximizes the variance between foreground and background classes. That is, the foreground and background have the highest distinguishability, and the formula for the variance between classes is: in, The grayscale threshold to be selected typically ranges from 40 to 200 for image grayscale values. Take an integer between 40 and 200; Threshold The proportion of foreground pixels in the segmentation. , Foreground pixel count This represents the total number of pixels in the image. Threshold Segmenting the background pixel proportions , This represents the number of background pixels. The average grayscale value of the foreground pixels. , The grayscale value is The number of pixels; The average grayscale value of the background pixels. ; Then, iterate through all possible grayscale thresholds. (40-200), calculate the corresponding , take The largest As the optimal segmentation threshold : The traditional Otsu algorithm requires traversing 256 grayscale values. The improved algorithm reduces the search range to 40-200 (only 161 grayscale values) through "grayscale range constraints," reducing the computational load by approximately 37%. Furthermore, when combined with batch processing scenarios in agricultural image preprocessing, the improved algorithm can shorten the ROI extraction time for a single image from the traditional 0.8-1.2 seconds to 0.4-0.6 seconds, significantly improving the overall efficiency of multimodal data acquisition and preprocessing, and meeting the needs of "real-time data augmentation" in smart agriculture.
[0092] 2. Perform image standardization processing to ensure consistency in subsequent feature extraction. The ROI needs to be uniformly resized to 1920×1080 resolution. Bicubic interpolation can be calculated by weighting the gray values of the 16 neighboring pixels to preserve leaf details (such as aphid attachment points and leaf texture).
[0093] For the scaled image pixels Its grayscale value is derived from the original image. China and Israel The 16 neighboring pixels corresponding to the original coordinates. The weighted sum of the gray values is obtained as follows: in: Original image Medium pixel grayscale value; The bicubic interpolation weighting function uses the classic B-spline function, and the formula is as follows: In the formula, The difference is the pixel coordinates. These are commonly used parameters to ensure that the interpolated image is smooth and free of jagged edges.
[0094] Then, the standardized image can be input into the CLIP model of the agricultural field, which has been fine-tuned for the agricultural scene, and matched with the preset agricultural text labels to output key phenotypic data and form an initial quantified phenotypic feature set.
[0095] The following are examples of steps for obtaining phenotypic features in specific application scenarios: For example, in terms of phenotypic feature quantification, the CLIP model, which has been fine-tuned for agricultural scenarios, is used to accurately match the standardized image of the user's input image command with the agricultural text label "cabbage leaf yellowing". Key phenotypic data are obtained by calculating feature vectors.
[0096] The phenotypic feature quantification algorithm for the yellow cabbage leaf scene is as follows: First, the yellowed area is segmented using a fine-tuned CLIP model. Phenotypic label matching is performed on the input cabbage leaf image to determine the confidence level of the "yellowing" label. To accurately locate the yellowed area in the leaf, the HSV (hue-saturation-brightness) color space threshold segmentation method is used, and segmentation parameters are set based on the color characteristics of cabbage leaf yellowing. The formula for converting RGB format leaf images to HSV format is as follows. : Then, HSV color space threshold segmentation is used to assist in locating the yellowed area; color threshold range: H∈[20,60] (yellow hue range), S∈[40,255] (saturation threshold, excluding light background), V∈[50,255] (brightness threshold, excluding overly dark areas).
[0097] Then, a ResNet-50-based image quality assessment model can be used to score the original images, and feature credibility can be obtained based on the scoring results. The initial quantized phenotypic feature set is then validated and filtered according to preset decision rules and feature credibility to obtain the final quantized phenotypic feature set and feature credibility report.
[0098] The following are examples of steps for obtaining feature credibility in specific application scenarios: For example, to avoid image quality issues affecting feature accuracy, a ResNet-50-based image quality assessment model is used to score the reference image. If the score is ≥80 (out of 100), the feature reliability is considered ≥90%, and the extracted phenotypic features are retained. If the score is below 80, the user should be prompted to re-upload the reference image. This step ultimately outputs a quantified phenotypic feature set (including specific numerical ranges) and a feature reliability report, providing accurate feature data support for text-image feature fusion.
[0099] Credibility mapping and decision rules are used to assign image quality scores. With phenotypic feature reliability Establish a mapping relationship and set the following decision rules: 1. If Score: Image quality is judged to be excellent, feature quantization error is low. 10%, corresponding to credibility Directly retain the quantified phenotypic feature set (such as...) (pixels) 2. If Score: Image quality is deemed moderate, feature quantization error is 10%-20%, corresponding to a certain level of confidence. A "manual review recommended" prompt should be added; 3. If Score: Image quality is deemed unacceptable, feature quantization error >20%, corresponding confidence level. This triggers a user interaction prompt: "Please re-upload a clear image of the cabbage leaf," and provides shooting instructions (such as "Keep the lens perpendicular to the leaf and avoid shooting against the light").
[0100] After obtaining the structured entity list, entity pair relation types, quantitative phenotypic feature sets, and feature credibility reports, the structured entity list, entity pair relation types, quantitative phenotypic feature sets, and feature credibility reports can be fused based on the agricultural knowledge graph to obtain the fusion result.
[0101] The following are examples of fusion steps in specific application scenarios: By relying on agricultural knowledge graphs, we can verify the consistency between image extraction features and text requirements, thus avoiding conflicts between "text requirements and image features".
[0102] Taking the scenario of yellowing cabbage leaves as an example, a knowledge graph specifically for the cabbage cultivation field is introduced. , where the node set Includes entity types such as "crop phenotype", "environmental factors", and "growth stage"; edge set This represents the agronomic relationships between entities, with each edge carrying a weight indicating the strength of that relationship. Based on agronomic literature and expert experience annotation, the scope , The closer to 1, the more significant the association.
[0103] A bidirectional feature mapping is performed to map the core requirement "yellow cabbage leaves" after parsing the text instructions to the target phenotypic node in the knowledge graph. (like =“Yellow leaves”); The quantified phenotypic features extracted from the image, “35% of the leaves are yellow and slightly curled”, are mapped to image feature nodes in the knowledge graph. .
[0104] For the mapped nodes, the logical matching degree between the text requirements and image features is calculated through "direct association verification + indirect association deduction". The process is as follows: 1. Direct association matching degree: Calculate the target text node Image feature nodes Direct correlation score If the two are the same node, such as both being "yellow leaves", then If there are directly related edges, such as "Yellow leaves" and "Nitrogen deficiency" (Edge weight); if there is no direct association, then .
[0105] 2. Indirect association derivation degree: If the direct association score is low, it is calculated using Dijkstra's algorithm for the shortest path in the knowledge graph. and Indirect correlation score Suppose the shortest path includes Edges, with weights respectively ,but: The text requirement is "yellow cabbage leaves". "Yellow leaves" and the image feature "slightly curled leaves" The shortest path for "curling" is "yellow leaves → nitrogen deficiency → insufficient soil fertility → insufficient soil moisture → curling". If the weights of the edges are as follows: ,but .
[0106] 3. Comprehensive matching degree of text-image features The weights are the weighted sum of the scores for direct and indirect associations. Used to highlight the priority of direct relationships: Set matching threshold ,like If the feature association is determined to be unconflicted; Marked as "logical conflict," this needs to be reported to the human-computer interaction interface to prompt the user to correct the requirement; when the text "cabbage yellow leaves" corresponds to the image "leaves not yellowing," This triggers a conflict warning.
[0107] When determining whether cabbage leaves are yellow, the text requirement is "cabbage leaves are yellow". The association matching calculation with the image features "yellow leaves, slight curling" is as follows: Regarding the characteristics of "yellow leaves": , , Match successful; Regarding the "mildly curled" characteristic: , , Further environmental constraints need to be considered (such as "soil moisture < 50%" in the text parsing). After adding environmental constraints, the direct association edge weights between "slight curling" and "soil moisture < 50%" are determined. , , The determination was found to be conflict-free.
[0108] Finally, the agricultural expert rule base can be called to perform logical verification on the fusion results, obtain the verification results, and based on the constraint generation algorithm, the fusion results and logical verification results can be transformed to obtain structured constraints.
[0109] The following are examples of the constraint verification and output steps after fusion in specific application scenarios: First, agronomic logic verification is performed by calling an agricultural expert rule base, which contains a large number of practically verified agronomic rules. Matching is then performed using the Rete algorithm, with the following rules: Network construction and single-condition matching+ The network is constructed and multiple conditions are combined to perform conflict detection. Then, the rules are sorted in descending order of confidence and the rule with the highest confidence is selected as the final matching rule.
[0110] The quantized feature model is transformed into structured constraint terms in standardized JSON format using a constraint generation algorithm. The transformation rules are as follows: 1. For numerical features, generate key-value pairs of “feature name - type - range”.
[0111] { "feature_name": "Percentage of yellowed area", "feature_type": "numerical", "constraint": { "lower_bound": 30.0, "upper_bound": 40.0, "unit": "%" } } 2. For location-based features, generate key-value pairs of “feature name - type - category set”.
[0112] { "feature_name": "Location of yellowing area", "feature_type": "categorical", "constraint": { "category_set": ["leaf edge"], "description": "The yellowing area is concentrated at the leaf margins, with no interveinal yellowing." } } 3. Integrate the structured constraints of image conversion with the environmental constraints of text parsing ("soil moisture < 50%", "growth stage = heading stage") to form a unified set of constraints. The integrated JSON is as follows: { "constraint_pair": [ { "feature_constraint": "Yellowed area accounts for 30%-40%", "environment_constraint": "Soil moisture < 50%" } ], "association_rule": "Soil moisture < 50% → Increased percentage of leaves with yellowing area (association strength 0.85)", "rule_confidence": 0.85} After feature association matching and constraint supplementation and integration are completed, the "conflict-free fusion result" and "structured constraint set" are output, and standardized according to the preset format. All constraint items are stored in the form of a JSON array, thus obtaining the structured constraints.
[0113] According to the present invention, a method for cross-modal consistency verification of augmented agricultural data based on human-computer interaction for large-scale multimodal and multidimensional data augmentation includes: The CLIP model is used to calculate the semantic similarity between image data and text data in augmented agricultural data and compare it with a preset threshold; and / or, based on agricultural knowledge graphs, the logical correlation between time-series data and image data in augmented agricultural data is verified.
[0114] Specifically, in some implementations, cross-modal consistency verification of augmented agriculture data may include using a CLIP model to calculate the semantic similarity between image data and text data in the augmented agriculture data, and comparing this similarity with a preset threshold. For example, after generating an image of a cabbage plant and its corresponding text description (such as "green leaves, no disease spots, in the heading stage"), the system can use the CLIP model to encode the image and text into embedding vectors respectively, and calculate the cosine similarity between them; if the similarity is lower than a preset threshold (such as 0.8), it may indicate a semantic mismatch, for example, the text description "green leaves" corresponds to an image that is obviously yellowed, such samples can be judged as failing verification.
[0115] In some implementations, cross-modal consistency verification may also include verifying the logical correlation between time-series data and image data in augmented agricultural data based on agricultural knowledge graphs. Agricultural knowledge graphs encode causal or correlational relationships between crop physiology, environmental stress, and phenotypic responses; for example, "soil moisture < 50% (field holding capacity)" typically leads to "leaf wilting or curling." During the verification phase, the system can query this knowledge graph to determine whether the generated time-series data (e.g., soil moisture of 45%) is logically consistent with the image content at the corresponding time point (e.g., whether the leaves exhibit curling characteristics). If the image shows unfolded leaves while soil moisture is significantly low, and there is no irrigation record, it may violate the physiological characteristics of cabbage, and the sample can be deemed to have failed verification.
[0116] It should be noted that the two verification methods mentioned above can be used individually or in combination (i.e., "AND / OR") to form a multi-layered verification mechanism.
[0117] In some implementations, data that fails validation can be fed back into the augmentation model for regeneration. This feedback mechanism ensures that the final augmented agriculture data meets cross-modal consistency requirements.
[0118] According to the present invention, a method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction is provided, the method further includes: The system validation metrics for the augmented agricultural dataset were obtained, as well as user feedback ratings for the augmented agricultural dataset were obtained through the human-computer interaction interface. User feedback ratings and system validation metrics are integrated into a reward signal; and optimized training samples are generated based on the augmented agricultural dataset, with a preset number of interactions as the iteration cycle. The parameters of the data augmentation model are optimized based on the reward signal and optimized training samples using a proximal policy optimization algorithm.
[0119] Specifically, the method provided in the embodiments of this application may further include a step of feedback optimization of the data augmentation model.
[0120] First, system validation metrics can be obtained by performing quality verification and cross-modal consistency verification on augmented agricultural data; and user feedback scores for the augmented agricultural dataset can be obtained through the human-computer interaction interface. These two types of feedback together constitute a multi-dimensional evaluation of the quality of augmented data.
[0121] In some implementations, system validation metrics can automatically record single-modal quality (such as the structural similarity index and peak signal-to-noise ratio of an image) and cross-modal consistency scores (such as CLIP semantic similarity between an image and text).
[0122] In some implementations, user feedback ratings can be collected via an interactive interface in the form of 1–5 stars, and users can highlight issues such as “does not conform to agronomic principles” or “distortion of details”.
[0123] Then, user feedback ratings and system validation metrics can be combined into a reward signal to guide model optimization. This reward signal reflects both user subjective preferences and retains the constraints of objective technical metrics, thus balancing practicality and generation quality during the optimization process.
[0124] In some implementations, user ratings can be normalized to the [0,1] interval, and system verification metrics (such as PSNR, semantic similarity, etc.) can also be mapped to the [0,1] interval through weighted fusion. The two are then summed according to a preset weight (e.g., 6:4) to form a comprehensive reward signal.
[0125] After receiving the reward signal, the system can generate optimized training samples based on the augmented agricultural dataset, with a preset number of interactions as the iteration cycle. For example, the system can set every 100 human-computer interactions as an iteration cycle, during which user ratings, system indicators, and corresponding augmented data generation records are continuously collected. After the cycle ends, the data generated in these 100 interactions and their context (including original instructions, input features, generation results, reward signals, etc.) are integrated into a batch of training samples for reinforcement learning, serving as the basis for subsequent model updates.
[0126] Finally, the parameters of the data-enhanced model can be optimized using the Proximal Policy Optimization (PPO) algorithm, based on the aforementioned reward signal and optimized training samples. Specifically, the generation policy of the large model can be treated as a learnable policy function, with the model parameters as the optimization object. Within the PPO framework, the probability ratio of actions generated by the new and old policies is calculated, and a clipping mechanism (e.g., clipping range) is introduced. To limit the policy update magnitude, a threshold of 0.2 is set, thereby minimizing the clipping loss function that includes reward weighting. Simultaneously, this process can jointly optimize the value function error and policy entropy to improve training stability and generalization.
[0127] In some implementations, the Adam optimizer can be used to iteratively update model parameters, and the reward improvement rate can be verified after each iteration. If the reward signal is significantly improved compared to the previous cycle (e.g., the increase exceeds a preset threshold), the current model parameters can be fixed and the next interactive optimization cycle can begin; otherwise, the learning rate can be rolled back or adjusted. This mechanism helps ensure that the model steadily converges to an optimal state that meets both user preferences and agricultural technology constraints during long-term human-computer interaction, thereby continuously improving the practicality and reliability of augmented data.
[0128] The method provided by the present invention will be further illustrated below through specific application scenarios.
[0129] Figure 2 This invention provides a large-scale multimodal and multidimensional data augmentation system based on human-computer interaction, such as... Figure 2 As shown, the system includes: Multimodal and multidimensional data acquisition and preprocessing module: Used to acquire multimodal and multidimensional data from different data sources. Based on the characteristics of different modalities, preprocessing techniques are used to process the multimodal and multidimensional data, and the data is formatted and standardized.
[0130] The human-computer interaction decision interface module provides a user interface for interaction with the system. Users input their data augmentation needs and constraints through this interface, such as commands like "generate high-resolution images of leaves affected by pests and diseases" or "generate instructions for pesticide application in response to rainy weather." Simultaneously, the system displays the data augmentation results to the user through the interface, allowing the user to evaluate and correct the results, thus forming a closed-loop human-computer interaction mechanism.
[0131] Large model dynamic generation processing module: Figure 3 This is a schematic diagram of the structure of the large model dynamic generation and processing module provided by the present invention, as shown below. Figure 3 As shown, this is used to receive preprocessed multimodal and multidimensional data. First, the large model receives the user's multimodal and multidimensional requirement instructions and transforms them into structured constraints through semantic parsing and format standardization. Second, the preprocessed raw data and structured constraints are mapped to a shared semantic space, and cross-modal feature alignment is achieved through raw data encoding and constraint encoding. Finally, based on the "modal type" in the task descriptor, the corresponding generation strategy is invoked, and an adapter ensures strategy coordination, thereby achieving the augmentation of multimodal and multidimensional data.
[0132] Cross-modal consistency verification and fusion: After generating multimodal and multidimensional data, the interface ensures correlation through multi-layered verification, returning unqualified data for regeneration. First, the CLIP model is used to calculate the semantic similarity (≥0.8) between the "generated image" and the "generated text description" to avoid misattributing the text "green leaves" to an "image of yellow leaves." An agricultural knowledge graph is used to verify the association between "time-series data" and "images" (e.g., when soil moisture is <50%, the leaves in the image must show curled features). Finally, the verified multimodal and multidimensional data are bound by "spatiotemporal tags" (such as images, moisture levels, and text with the same timestamp) to form a complete sample package. If the data does not conform, it is fed back to the large model processing module for regeneration.
[0133] The feedback optimization module is used to return the fused multimodal augmented data to the user via the interface. It also collects two types of feedback: user ratings (1-5 stars, highlighting issues such as "inconsistent with agronomical patterns" and "distorted details") and system metrics: automatically recording single-modal quality (e.g., the Structural Similarity Index, SSIM) and cross-modal consistency scores as the basis for subsequent optimization. The model is optimized once every 100 interactions, using "user ratings + validation metrics" as the reward signal.
[0134] In some preferred embodiments, each multimodal multidimensional data sample in the multimodal multidimensional dataset includes image data of the same plant, soil temperature and humidity data, environmental data, and agricultural log records.
[0135] In some preferred embodiments, the multimodal and multidimensional data includes multimodal and multidimensional data collected at different times and frequencies during different stages of the crop growth cycle.
[0136] In some preferred embodiments, the preprocessing techniques include denoising and normalizing the image data, and filtering, downsampling, and feature extraction of the point cloud data.
[0137] In some preferred embodiments, the data formatting in the preprocessing is uniform, which means converting the image to JPEG format, naming it according to "plot ID-collection time-camera number", compressing the point cloud data to LAS format, retaining the XYZ coordinates and reflection intensity fields, parsing the text log to JSON format, and extracting the structured information of "agricultural operation type-time-executor".
[0138] In some preferred embodiments, the user-system interaction interface in the human-computer interaction decision interface module is divided into two parts: a visual interaction interface and a natural language understanding unit, and supports users to input constraints in the following ways: Semantic annotation: Select key areas (such as "cabbage head") on the image and associate them with text descriptions ("5-8cm in diameter indicates maturity"). Rule settings: Select spatiotemporal constraints via the drop-down menu (e.g., "When the augmentation data meets the requirement that the soil moisture is >60%, the leaf spread of the crop image is >80%"). Feedback rating: Rate the machine-generated augmented samples from 1 to 5 stars, highlighting issues such as "semantic errors" and "modal conflicts".
[0139] In some preferred embodiments, the semantic parsing in the large model dynamic generation processing module refers to using agricultural LLM to decompose instructions into "modal types" (images, time series, text), "core elements" (crop = cabbage, growth stage = heading stage), and "attribute constraints" (leaf yellow spot ratio <10%, soil moisture 50-70%). In some preferred embodiments, the format standardization in the large model dynamic generation processing module refers to generating a unified task descriptor (JSON format) that includes the generation objectives of each modality (such as "image resolution 1920×1080" and "time-series sampling frequency 1 time / hour") and association rules (such as "yellow spots on leaves in the image → text must contain 'may be nitrogen deficient'").
[0140] In some preferred embodiments, the raw data encoding in the large model dynamic generation processing module refers to processing the input data with a unified encoder. For example, images are encoded into visual feature vectors (extracting leaf morphology, color, etc.); soil moisture is encoded into temporal feature vectors (capturing diurnal fluctuation patterns); and agricultural text is encoded into semantic vectors (interpreting the agronomic meaning of operations such as "irrigation" and "fertilization"). In some preferred embodiments, the constraint encoding in the large model dynamic generation processing module refers to converting the attribute constraints in the task descriptor (such as "soil moisture 50-70%)" into spatial constraint boundaries to ensure that the generated data does not deviate from the range.
[0141] In some preferred embodiments, the strategy routing in the large model dynamic generation processing module refers to assigning generators based on modality type. For image data, the "GAN+CLIP large model" combination is invoked; for time-series data, the "Transformer time-series model + physical constraints" is enabled; and for text data, "instruction fine-tuning + LLM" is triggered. Cross-modal adaptation is performed by associating different generators through an attention mechanism that shares a semantic space. For example, when the image generator outputs the "yellow spots on leaves" feature, it automatically sends a signal to the time-series generator that "low soil nitrogen content needs to be associated" to ensure data logic consistency.
[0142] In some preferred embodiments, the multimodal and multidimensional data augmentation in the large model dynamic generation and processing module refers to analyzing data requirements and constraints based on user-input text descriptions and dynamically invoking corresponding generation strategies. For example, performing style conversion and content addition on existing images; performing interpolation and extrapolation on time-series data to generate more dimensional data.
[0143] In some preferred embodiments, the specific optimization strategy in the feedback optimization module is as follows: Using large model parameters as the optimization object, each iteration cycle consists of 100 human-computer interactions. The user's 1-5 star rating of the augmented data (normalized to [0,1]) and system validation metrics (image PSNR, text semantic similarity, etc., weighted and fused to [0,1]) are summed at a weight ratio of 6:4 to serve as the reward signal. Next, augmented data generation records from 100 interactions are collected to construct PPO training samples. Using model parameters as the strategy, the old and new strategies are edited by ratio (…). =0.2) Limit the update magnitude, minimize the reward-weighted clipping loss function, and simultaneously optimize the value function error and policy entropy. Finally, use the Adam optimizer to iteratively update the model parameters. After each iteration, verify the reward improvement rate. If the target is met, solidify the parameters and enter the next cycle to ensure that the model stably converges to the optimal state that meets user preferences and technical constraints.
[0144] The following uses the augmentation of monitoring data for cabbage planting in smart agriculture as an example to explain in detail the implementation process of this invention: 1. Data Collection 1.1 Determine data dimensions and metrics Based on the physiological characteristics of cabbage (such as its preference for cool temperatures, poor tolerance to waterlogging, and sensitivity to light during the heading stage) and the planting scenario, the following collection indicators were defined: (1) Environmental monitoring data Meteorological data: air temperature (key range 5-30℃, optimal 15-20℃ during the heading stage), relative humidity (60%-80%, excessive humidity can easily cause downy mildew), light intensity (daily average ≥4 hours of direct sunlight, affecting heading compactness), precipitation, and wind speed (affecting transpiration). Soil data: soil temperature (10-25℃), soil moisture (60%-80% of field capacity, irrigation is required if it is below 50%), soil pH (5.5-6.5, slightly acidic is suitable), and soil available nitrogen / phosphorus / potassium content (affecting leaf growth and head quality).
[0145] (2) Crop phenotypic data Growth indicators: plant height (5-15cm in seedling stage, 20-40cm in rosette stage, 40-60cm in heading stage), stem diameter (≥0.8cm for healthy seedlings), number of leaves (4-6 leaves in seedling stage, 15-20 leaves in rosette stage), leaf area index (LAI, ≥3.5 in heading stage), head diameter (15-25cm at maturity), leaf chlorophyll value (SPAD value, reflecting nitrogen level, ≥30 in seedling stage); Stress data: pest and disease severity (e.g., downy mildew level 1-5, aphid density (heads / plant)), leaf wilting degree (0-3, reflecting drought level).
[0146] (3) Farmland management data Agricultural operations: sowing date, fertilizer type (organic fertilizer / compound fertilizer) and dosage (e.g., 5 kg / mu of nitrogen fertilizer during the seedling stage), irrigation time and water volume (high water requirement during the heading stage, 30 m³ of water per irrigation). 3 / acre), pesticide spraying records (such as the time of application for controlling cabbage caterpillars); equipment records: sensor deployment location (such as monitoring air temperature and humidity 1.2m above the ground), drone inspection time (once a week to acquire canopy images).
[0147] 1.2 Data Sources and Collection Methods IoT sensors: Sensors are deployed in the cabbage field in a 50m×50m grid to collect soil moisture and air temperature and humidity data in real time (every 15 minutes). The data is automatically uploaded to a cloud platform (such as an agricultural IoT management system). Field equipment data collection: Plant height and stem diameter were measured weekly using a handheld phenotyping instrument; canopy images were acquired using a multispectral drone to extract leaf area index and vegetation coverage. Manual recording: Farmers record agricultural operations (such as fertilization time and amount) through a mobile APP; technicians take photos of pest and disease samples and label them with their severity level (e.g., "2024-07-10, Plot A, Downy mildew level 3"). Historical data integration: Collect cabbage planting records for the target area over the past three years (such as monitoring data from agricultural technology extension stations and historical climate data from meteorological stations).
[0148] 1.3 Data Preprocessing Cleaning: Remove outliers (such as air temperature >40℃ caused by sensor failure) and fill in short-term missing values (such as missing data for 1-2 hours) with linear interpolation. Standardization: Normalize data from different units (e.g., map temperature to the 0-1 range) and unify the time granularity (summarize daily averages by "day"); Structured data conversion: Convert unstructured data such as images and text into structured indicators (e.g., drone image → leaf area index) to form a three-dimensional data table of "time-plot-indicator" (Example: Plot B, 2024-06-15, air temperature 26℃, plant height 22cm, fertilizer application rate 3kg / mu).
[0149] 2. Human-computer interaction constraints 2.1 Expert Knowledge Input Format Phenological constraints: Clearly define the time boundaries of each growth stage (seedling stage 30-40 days, rosette stage 25-30 days, heading stage 40-50 days) and the range of indicators (e.g., plant height during heading stage must not be less than 40cm). Physiological constraints: Define the correlation of key indicators (e.g., when soil nitrogen content is >50mg / kg, SPAD value ≥35); exclude unreasonable combinations (e.g., "air humidity >90% + heading stage" will not result in "no leaf disease spots"); Scenario constraints: Limit the scope of scenarios covered by the data (e.g., "Spring cabbage in North China" must include scenarios of late spring cold snap (0-5℃), and "Autumn cabbage in South China" must include scenarios of high temperature (30-35℃).
[0150] 2.2 Structuring of Constraint Rules Establish a "constraint rule base" and store it in the form of "IF-THEN" (e.g., IF growth stage = seedling stage THEN plant height ∈ [5cm, 15cm]). A visual interface is developed, allowing experts to adjust parameters via sliders (such as dragging to set the upper limit of the heading stage temperature to 25℃). The system automatically converts the rules into mathematical expressions that the model can recognize (such as temperature constraint: T≤25℃).
[0151] 3. Large model generation 3.1 Model Selection and Construction The basic model uses a "Constrained Generative Adversarial Network (GAN)" which includes a generator (outputting "pseudo-data") and a discriminator (distinguishing between "real data" and "pseudo-data"). Constraint fusion: Add a "constraint penalty term" to the generator loss function—if the generated data violates the rules (such as seedling head formation), the loss value is increased, forcing the generator to adjust its output; Time series module: Add an LSTM layer to capture time correlations (such as the humidity change trend over 7 consecutive days) to ensure that the generated data conforms to the continuity of cabbage growth (such as the average daily plant height growth ≤ 2cm).
[0152] 3.2 Model Training and Generation Training setup: Train the model using preprocessed real data (70% training set, 30% validation set), iterate 5000 times (batch size=32), and output intermediate results every 100 iterations for expert evaluation; Diversity control: By inputting "scenario labels" (such as "drought year" and "high fertilizer management"), the generator is guided to cover different scenarios (such as the average soil moisture in drought year data being 20% lower than that in normal years). Constraint verification: Generate a real-time data matching rule base to filter samples that violate constraints (such as removing data with "heading temperature < 10℃").
[0153] 4. Verification and Optimization 4.1 Authenticity Verification Statistical validation: Compare the distribution of generated data with that of real data (such as mean and variance), requiring that the difference in key indicators be <5% (such as the difference in the mean air temperature <1℃). Expert evaluation: Three cabbage cultivation experts were invited to blind evaluate 100 data points (a mixture of real and unreal data), and the percentage of "reasonable data" must be ≥90%. Application Validation: Train the "Cabbage Yield Prediction Model" using generated data. If the accuracy (R²) is... 2 If the improvement is ≥5% compared to using only real data, it indicates that the data is effective.
[0154] 4.2 Optimization and Iteration If the generated data shows that "the proportion of data during the heading stage is too low", adjust the model sampling strategy and increase the training weight for this stage; if experts report that "the drought scenario data is unreasonable", supplement the physiological rules under drought stress (such as the correlation between leaf wilting degree and soil moisture) and retrain the model.
[0155] 5. Output dataset 5.1 Dataset Integration and Structuring Ratio and division: The data is integrated according to the ratio of "real data: generated data = 1:2" (e.g., 10,000 real data points + 20,000 generated data points), and then divided into training set, validation set, and test set in a ratio of 8:1:1. Field standardization: Unify field formats (such as time "YYYY-MM-DD", temperature "℃"), and include full-dimensional information such as "time-plot-growth stage-environmental indicators-crop indicators-management measures".
[0156] 5.2 Output Content and Format Data format: CSV, JSON, and Parquet formats are provided, along with metadata descriptions (such as indicator meanings and constraint rules). Quality report: Includes validation results (such as distribution comparison charts and expert evaluation opinions), demonstrating the reliability of the dataset; API Interface: Develop an API interface that supports filtering data by "growth stage" and "indicator type" (such as "get soil nitrogen content data during the rosette stage"), making it convenient for precision irrigation and other systems to directly call the API.
[0157] The following describes the large-scale multimodal and multidimensional data augmentation device based on human-computer interaction provided by the present invention. The large-scale multimodal and multidimensional data augmentation device based on human-computer interaction described below can be referred to in correspondence with the large-scale multimodal and multidimensional data augmentation method based on human-computer interaction described above.
[0158] Figure 4 This is a schematic diagram of the structure of the large-scale multimodal and multidimensional data augmentation device based on human-computer interaction provided by the present invention, as shown below. Figure 4 As shown, the device includes the following modules: The data acquisition and processing module 400 is used to acquire multimodal and multidimensional agricultural data from different data sources, preprocess and format the multimodal and multidimensional agricultural data to obtain processed agricultural data, and obtain multimodal demand commands input by users through a human-computer interaction interface. The data augmentation module 410 is used to call the data augmentation model based on the multimodal demand instructions and the processed agricultural data to obtain augmented agricultural data; The verification fusion module 420 is used to perform cross-modal consistency verification on augmented agricultural data and fuse the verified augmented agricultural data with the processed agricultural data to obtain the augmented agricultural dataset. Among them, the data augmentation model includes a large-scale language model in the agricultural field, a CLIP model in the agricultural field, an agricultural knowledge graph, and data generators corresponding to multiple modal types; the data generators corresponding to multiple modal types are associated through an attention mechanism that shares a semantic space.
[0159] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided. Based on multimodal demand commands and processed agricultural data, it invokes a data augmentation model to obtain augmented agricultural data, including: By using a large-scale language model, a CLIP model, and an agricultural knowledge graph in the agricultural field, features are extracted and fused from multimodal demand instructions to obtain structured constraints and determine the target modality type in the multimodal demand instructions. The processed agricultural data is encoded into multimodal feature vectors, and the structured constraints are encoded into target constraint boundaries. Augmented agricultural data is generated based on multimodal feature vectors and target constraint boundaries through a data generator corresponding to the target modality type.
[0160] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided, wherein the multiple modal types include one or more of image data types, time-series data types and text data types; Among them, the data generator corresponding to the image data type consists of GAN and CLIP models; The data generator corresponding to the time series data type consists of a Transformer time series model and a physical constraint rule library; The data generator corresponding to the text data type consists of a large language model that has been fine-tuned with knowledge from the agricultural field.
[0161] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided. This device extracts and fuses features from multimodal demand instructions using a large-scale language model, a CLIP model, and an agricultural knowledge graph, to obtain structured constraints, including: By using a large-scale language model in the agricultural field, features are extracted from the text instructions in multimodal demand instructions to obtain a structured entity list and entity pair relationship types; The CLIP model in the agricultural field is used to extract features from image instructions in multimodal demand instructions, resulting in a quantitative phenotypic feature set and a feature credibility report. Based on agricultural knowledge graphs, the structured entity list, entity pair relationship types, quantitative phenotypic feature sets, and feature credibility reports are fused to obtain the fusion result; The agricultural expert rule base is invoked to perform logical verification on the fusion results, and the verification results are obtained. Based on the constraint generation algorithm, the fusion result and the logical verification result are transformed to obtain structured constraints.
[0162] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided. This device extracts features from textual instructions in multimodal demand instructions using a large-scale language model in the agricultural field, obtaining a structured entity list and entity pair relationship types, including: Based on agricultural corpus, the agricultural-specific large language model AgriLLM is fine-tuned to obtain a large language model for the agricultural field. Semantic features of text instructions are obtained through large-scale language models in the agricultural field; The semantic features of the text instructions are input into a CRF-based entity recognition model to obtain a structured entity list; Input the structured entity list into the dual Transformer model to obtain the initial relationship type and entity association probability of each entity pair in the structured entity list; The initial relation types are filtered by the entity association probability to obtain the final entity pair relation types.
[0163] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided. This device extracts features from image instructions in multimodal demand instructions using a CLIP model in the agricultural field, obtaining a quantified phenotypic feature set and a feature credibility report, including: The region of interest is extracted from the original image in the image instruction, and the extracted image is standardized by bicubic interpolation algorithm to obtain a standardized image; The standardized image is input into the CLIP model of the agricultural field, which has been fine-tuned for the agricultural scene, and matched with the preset agricultural text labels to output key phenotypic data and form an initial quantified phenotypic feature set. The original images are scored using a ResNet-50-based image quality assessment model, and the feature credibility is obtained based on the scoring results. The initial quantified phenotypic feature set is verified and filtered according to the preset decision rules and feature credibility to obtain the final quantified phenotypic feature set and feature credibility report.
[0164] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided, which extracts the region of interest from the original image in the image command, including: For any candidate threshold within the candidate grayscale search range, calculate the inter-class variance between the foreground and background corresponding to the arbitrary candidate threshold; Compare the inter-class variances corresponding to all candidate thresholds within the search interval, select the candidate threshold that maximizes the inter-class variance as the optimal segmentation threshold, and use the optimal segmentation threshold to perform binarization segmentation on the original image to obtain the region of interest. The candidate grayscale search interval is set based on prior knowledge of the grayscale distribution of the target and background in agricultural images.
[0165] According to the present invention, a method for cross-modal consistency verification of augmented agricultural data based on a human-computer interactive large-model multimodal multidimensional data augmentation device includes: The CLIP model is used to calculate the semantic similarity between image data and text data in augmented agricultural data and compare it with a preset threshold; and / or, based on agricultural knowledge graphs, the logical correlation between time-series data and image data in augmented agricultural data is verified.
[0166] According to the present invention, a large-scale multimodal and multidimensional data augmentation device based on human-computer interaction is provided, the device further comprising a model optimization module for: The system validation metrics for the augmented agricultural dataset were obtained, as well as user feedback ratings for the augmented agricultural dataset were obtained through the human-computer interaction interface. User feedback ratings and system validation metrics are integrated into a reward signal; and optimized training samples are generated based on the augmented agricultural dataset, with a preset number of interactions as the iteration cycle. The parameters of the data augmentation model are optimized based on the reward signal and optimized training samples using a proximal policy optimization algorithm.
[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction, characterized in that, include: Multimodal and multidimensional agricultural data from different data sources are collected, and the multimodal and multidimensional agricultural data are preprocessed and formatted to obtain processed agricultural data. And obtain multimodal demand instructions from users through the human-computer interaction interface; Based on the multimodal demand instructions and the processed agricultural data, the data augmentation model is invoked to obtain augmented agricultural data; The augmented agricultural data is subjected to cross-modal consistency verification, and the verified augmented agricultural data is fused with the processed agricultural data to obtain the augmented agricultural dataset. The data augmentation model includes a large-scale language model for agriculture, a contrastive language-image pre-trained CLIP model for agriculture, an agricultural knowledge graph, and data generators corresponding to various modalities. The data generators corresponding to the various modal types are associated through an attention mechanism that shares a semantic space.
2. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 1, characterized in that, Based on the multimodal demand instructions and the processed agricultural data, a data augmentation model is invoked to obtain augmented agricultural data, including: By using a large-scale language model, a CLIP model, and an agricultural knowledge graph in the agricultural field, features are extracted and fused from the multimodal demand instructions to obtain structured constraints; and the target modality type in the multimodal demand instructions is determined. The processed agricultural data is encoded into multimodal feature vectors, and the structured constraints are encoded into target constraint boundaries. Based on the multimodal feature vectors and the target constraint boundaries, augmented agricultural data is generated by the data generator corresponding to the target modality type.
3. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 1 or 2, characterized in that, The multiple modal types include one or more of image data types, time-series data types, and text data types; Among them, the data generator corresponding to the image data type consists of the Generative Adversarial Network (GAN) and the CLIP model; The data generator corresponding to the time series data type consists of a Transformer time series model and a physical constraint rule library; The data generator corresponding to the text data type consists of a large language model that has been fine-tuned with knowledge from the agricultural field.
4. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 2, characterized in that, By using a large-scale language model, a CLIP model, and an agricultural knowledge graph, features are extracted and fused from the multimodal demand instructions to obtain structured constraints, including: By using a large-scale language model in the agricultural field, features are extracted from the text instructions in the multimodal demand instructions to obtain a structured entity list and entity pair relationship types; The image instructions in the multimodal demand instructions are extracted using the CLIP model in the agricultural field to obtain a quantitative phenotypic feature set and a feature credibility report; Based on the agricultural knowledge graph, the structured entity list, the entity pair relationship type, the quantitative phenotypic feature set, and the feature credibility report are fused to obtain the fusion result; The fusion result is logically verified by calling the agricultural expert rule base to obtain the verification result; Based on the constraint generation algorithm, the fusion result and the logical verification result are transformed to obtain structured constraints.
5. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 4, characterized in that, The process involves extracting features from the textual instructions in the multimodal demand instructions using a large-scale language model in the agricultural field, resulting in a structured entity list and entity pair relationship types, including: Based on agricultural corpus, the agricultural-specific large language model AgriLLM is fine-tuned to obtain the large language model for the agricultural field. The semantic features of the text instructions are obtained through a large-scale language model in the agricultural field. The semantic features of the text instructions are input into an entity recognition model based on a Conditional Random Field (CRF) to obtain a structured entity list. Input the structured entity list into the dual Transformer model to obtain the initial relationship type and entity association probability of each entity pair in the structured entity list; The initial relationship type is filtered by the entity association probability to obtain the final entity pair relationship type.
6. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 4, characterized in that, The process involves extracting features from image instructions within the multimodal demand instructions using the CLIP model in the agricultural field, resulting in a quantified phenotypic feature set and a feature credibility report, including: The region of interest is extracted from the original image in the image instruction, and the extracted image is standardized by bicubic interpolation algorithm to obtain a standardized image. The standardized image is input into the CLIP model of the agricultural field, which has been fine-tuned for agricultural scenarios, and matched with the preset agricultural text labels to output key phenotypic data and form an initial quantitative phenotypic feature set. The original image is scored using a ResNet-50-based image quality assessment model, and the feature confidence level is obtained based on the scoring results. The initial quantified phenotypic feature set is verified and filtered according to the preset decision rules and the feature credibility to obtain the final quantified phenotypic feature set and feature credibility report.
7. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 6, characterized in that, Extracting the region of interest from the original image in the image instruction includes: For any candidate threshold within the candidate grayscale search range, calculate the inter-class variance between the foreground and background corresponding to the arbitrary candidate threshold; Compare the inter-class variances corresponding to all candidate thresholds within the search interval, select the candidate threshold that maximizes the inter-class variance as the optimal segmentation threshold, and use the optimal segmentation threshold to perform binarization segmentation on the original image to obtain the region of interest. The candidate grayscale search interval is set based on prior knowledge of the grayscale distribution of the target and background in agricultural images.
8. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 1, characterized in that, Methods for cross-modal consistency verification of the augmented agriculture data include: The semantic similarity between image and text data in the augmented agriculture data is calculated using the CLIP model and compared with a preset threshold; and / or, Based on agricultural knowledge graphs, the logical correlation between time-series data and image data in the augmented agricultural data is verified.
9. The method for augmenting large-scale multimodal and multidimensional data based on human-computer interaction according to claim 1 or 8, characterized in that, The method further includes: The system validation metrics for the augmented agricultural dataset are obtained, and user feedback ratings for the augmented agricultural dataset are obtained through a human-computer interaction interface. The user feedback rating and the system verification index are merged into a reward signal; and optimized training samples are generated based on the augmented agricultural dataset with a preset number of interactions as the iteration cycle. The parameters of the data augmentation model are optimized based on the reward signal and the optimized training samples using a proximal policy optimization algorithm.
10. A large-scale multimodal and multidimensional data augmentation device based on human-computer interaction, characterized in that, include: The data acquisition and processing module is used to acquire multimodal and multidimensional agricultural data from different data sources, preprocess and format the multimodal and multidimensional agricultural data, and obtain processed agricultural data. And obtain multimodal demand instructions from users through the human-computer interaction interface; The data augmentation module is used to call the data augmentation model based on the multimodal demand command and the processed agricultural data to obtain augmented agricultural data; The verification and fusion module is used to perform cross-modal consistency verification on the augmented agricultural data, and to fuse the verified augmented agricultural data with the processed agricultural data to obtain the augmented agricultural dataset. The data augmentation model includes a large-scale language model for agriculture, a contrastive language-image pre-trained CLIP model for agriculture, an agricultural knowledge graph, and data generators corresponding to various modalities. The data generators corresponding to the various modal types are associated through an attention mechanism that shares a semantic space.
Citation Information
Patent Citations
Brassica oleracea knowledge dynamic expression and interaction method and system based on multi-modal fusion
CN119862954A
Intelligent data augmentation and cleaning system and method based on multi-modal consistency detection
CN120179996A
Knowledge injection-based multi-modal agricultural disease and insect pest large model prediction method
CN120256812A
Smart farm knowledge warehouse establishment method and system
CN120910104A
Server, display device, and digital human processing method
WO2025066217A1