Field large model construction and question and answer service method oriented to whole course of grass production

By constructing multimodal data and knowledge graphs, and combining human-computer collaboration and multimodal retrieval, the problems of insufficient professionalism and system fragmentation of large language models in grassland production are solved, and high-precision, personalized decision support for grassland production is achieved.

CN122065880APending Publication Date: 2026-05-19INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
Filing Date
2026-03-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing large language models lack professional depth in grassland production, are prone to factual errors, lack high-quality domain instruction corpora, exhibit fragmentation, lack multimodal fusion decision-making capabilities, and cannot meet the precise guidance needs of the entire grassland production process.

Method used

We construct multimodal data covering the entire grassland industry process, generate high-quality instruction fine-tuning corpus through human-machine collaboration, establish a structured knowledge graph and unstructured vector library, design a workflow architecture with central routing and expert sub-agents, achieve enhanced multimodal retrieval, and generate highly accurate answers.

Benefits of technology

It ensures the scientific nature and safety of grassland production decisions, improves the efficiency of corpus construction, provides personalized decision support covering the entire process, and solves the professional problems of general models in grassland production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065880A_ABST
    Figure CN122065880A_ABST
Patent Text Reader

Abstract

The invention discloses a field large model construction and question-and-answer service method oriented to the whole course of grass production. The method comprises the following steps: acquiring whole-process original data of the prairie industry, and preprocessing to form a basic training corpus and a referenceable evidence library; secondly, constructing a grassland full-process knowledge graph and a multi-modal embedded vector library, and forming a graph and vector double-track knowledge base; and then generating a high-quality instruction fine-tuning corpus by using a pre-generated question and answer draft and performing adversarial error correction, and performing fine tuning to obtain a large model in the field of the grassland industry. In the question and answer service stage, a workflow of a central route and a field expert sub-agent is constructed, and intention splitting and scheduling are carried out on user query; and synchronously executing accurate logic routing and retrieval in a double-track base, inputting a retrieval enhanced context into a large model in the field of the prairie industry for reasoning, and outputting a final professional solution. According to the method, the knowledge illusion of a general large model in the vertical field of the grassland industry is effectively overcome, and high-precision and multi-modal intelligent decision support of the whole grassland industry production process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and smart agricultural information, and in particular to a domain-wide model construction technology based on multimodal retrieval enhancement and multi-agent collaboration, specifically to a domain-wide model construction and question-answering service method for the entire process of grass production. Background Technology

[0002] The forage industry is a crucial pillar of modern agriculture and ecological civilization construction, encompassing multiple complex stages including seed selection and breeding, planting management, harvesting and processing, storage, transportation, and sales. With the industry's large-scale development, practitioners are increasingly demanding precise and scientific agricultural technology guidance and decision support. In recent years, artificial intelligence technologies, represented by large language models, have made groundbreaking progress, and their powerful natural language understanding and generation capabilities have provided new solutions for smart agriculture.

[0003] However, directly applying existing large-scale modeling techniques to real-world grassland production scenarios still faces the following prominent technical bottlenecks: First, general-purpose models lack specialized depth, easily leading to factual "illusions." While the training corpus of general-purpose large-scale language models is broad, it lacks depth across specific domains. When faced with specialized grassland technical issues (such as distinguishing the subtle differences between "alfalfa" and "oat grass" in water requirements and pest and disease characteristics), the model often generates seemingly reasonable but seriously factually flawed content based on probability, or can only provide vague general suggestions, failing to meet the stringent accuracy requirements of agricultural production. Second, high-quality domain instruction corpora are scarce, and traditional fine-tuning is too costly. Training a specialized domain-specific large-scale model requires a massive amount of high-quality instruction-question-answer pairs. Traditional fine-tuning corpus construction relies entirely on manual writing by domain experts, which is not only time-consuming, labor-intensive, and costly, but experts often struggle to fully cover the complex, colloquial question variations of end-users (such as farmers), resulting in insufficient generalization ability of the model in real-world interaction scenarios. Third, existing systems exhibit "fragmentation," lacking full lifecycle collaborative decision-making capabilities. Current information technology solutions for the grass industry are mostly "information silos," such as image recognition systems that only target pests and diseases, or rule engines that only provide irrigation decisions. These fragmented systems cannot understand the logical connections between upstream and downstream stages of grass production and struggle to handle complex cross-stage planning tasks (e.g., recommending grass species based on regional soil conditions and generating matching overwintering pest control plans). Finally, existing knowledge bases are rigid in their interaction and lack multimodal fusion reasoning capabilities. Existing agricultural expert systems are mostly based on fixed rule trees or keyword matching, unable to understand the complex semantics of natural language; at the same time, grass production heavily relies on visual information (such as pest and disease symptom images and growth maps), and existing large-scale text models struggle to perform deep mapping and multimodal joint reasoning between external image features and underlying professional knowledge graphs.

[0004] Therefore, there is an urgent need for an intelligent question-answering and decision-making service method that can deeply integrate multimodal expertise in the grassland industry, overcome the illusion of factuality in models, and cover the entire grassland production chain, in order to meet the precise guidance needs of farmers, herdsmen, and related enterprises in complex production environments. Summary of the Invention

[0005] This invention aims to address the shortcomings of general-purpose large-scale models in the grass industry, such as lack of specialization and susceptibility to factual errors, as well as the limitations of existing intelligent solutions in covering only a single stage and lacking multimodal fusion decision-making capabilities. It provides a method for constructing a domain-specific large-scale model and providing question-answering services for the entire grass industry production process, thereby achieving high-precision, multimodal intelligent decision support for the entire grass industry production process. The core idea of ​​this invention is as follows: First, comprehensively collect multimodal data covering the entire grass industry production process. After cleaning and decoupling text and graph, construct a basic training corpus and a referable evidence library. Second, adopt a "human-machine collaboration" mechanism. Utilize the general-purpose large-scale model to pre-generate question-answer drafts, which are then subjected to factual adversarial error correction and multimodal annotation by domain experts to generate high-quality instruction fine-tuning corpus. This corpus is then used to further pre-train and supervisedly fine-tune the base large-scale model across the entire grass industry production process. Simultaneously, extract entity relationships from the evidence library to construct a hierarchical knowledge graph of the entire grass industry process. Combine this with a multimodal embedding model to establish a vector library, forming a dual-track knowledge base composed of a structured graph and unstructured vectors. Finally, a workflow architecture is constructed that includes a central routing agent and multiple domain expert sub-agents. When receiving multimodal queries from users, the system performs precise logical pathfinding and cross-modal similarity semantic retrieval simultaneously within a dual-track knowledge base through intent decomposition and routing scheduling, combined with the user's historical memory archive. The fused multimodal retrieval enhanced context is then input into a pre-trained large-scale grassland industry model for reasoning, ultimately generating a professional solution with high accuracy, traceability, and full-process coverage.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0007] This paper provides a method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process. The specific steps include: S1. Collect multimodal raw data covering the entire process of grassland production. The dataset was then cleaned and desensitized to obtain a standardized dataset. According to the preset division criteria, the Classified into basic training corpus and citationable evidence library and the above The long texts in the document are semantically sliced ​​and metadata-tagged, while the multimodal documents within them are decoupled from the text and image and associated with bidirectional mapping. S2, Utilizing a general large language model In conjunction with the above Pre-generate a draft of instruction question and answer that includes various question variations and multimodal interaction intents. After review and error correction by domain experts and multimodal annotation, a high-quality instruction fine-tuning corpus was formed. and with the stated Further pre-training of the base multimodal large model in the relevant domain was performed, using the aforementioned Supervised fine-tuning was performed to obtain a large-scale model for the grassland industry and its final parameters. ; S3, from the above Extracting professional knowledge elements from the grassland industry to construct a domain ontology and a hierarchical knowledge graph of the entire grassland industry process and the Text slices and image content are encoded into shared semantic space vectors and stored in a multimodal vector library. Thus constructing the structure described and Together they form a dual-track knowledge foundation ; S4. Construct a central routing agent. and multiple domain expert sub-intelligent agents Workflow architecture Receive user multimodal query input , by the Perform intent recognition and task decomposition and schedule to the corresponding task. The scheduled Combine user memory profile In the Perform logical retrieval in the and in the Cross-modal similarity retrieval is performed, and the results of the two-way retrieval are fused to form a multimodal retrieval enhanced context. Then the above With the Common input to the mounting parameters A large-scale model in the grassland industry was developed to generate the final solution output. .

[0008] Furthermore, the specific method of step S1 includes the following sub-steps: S1-1. Collect multimodal raw datasets covering the entire process of grassland production. ,in It is a collection of unstructured text. For a collection of images, and The initial number of text and images are defined; a text cleaning function based on regular expressions and a domain dictionary is defined. Image cleaning functions Desensitization function with privacy information mask Regarding the above After deduplication, error correction, format standardization, and data anonymization, a standard dataset is obtained. ,in , According to the preset information entropy threshold With the structured degree scoring function Define the partitioning function:

[0009] The Classified into basic training corpus and citationable evidence library ,in Indicates basis and A parameterized data flow routing function that determines the trainability and citationability of samples and completes the aggregation. S1-2, Regarding the aforementioned citation evidence library Long text collections in For any long text A slicing algorithm based on sliding window and semantic boundary markers is adopted. Semantic slicing is performed to obtain evidence slice sequences. And summarize them to form a collection of evidence slices. ,in, The number of slices generated from this long text. Define the total number of slices; define the multidimensional metadata feature space. ,in , and These represent the stages of grass production, suitable regions, and grass species, respectively; metadata extraction functions are used. Slice each piece of evidence Generate the corresponding metadata vector And assign a globally unique identifier to each piece of evidence. Based on this, construct a set of evidence text slices with metadata tags. ; S1-3, Regarding the aforementioned citation evidence library Multimodal literature For any multimodal literature Using a layout analysis model Parse it into a collection of page text blocks With image collection ,in and These represent the number of text blocks and images parsed from the document, respectively; for any image object Based on their spatial adjacency relationship from Selecting a set of adjacent text blocks And describe the function through a visual language model. Generate structured image descriptions To satisfy the dimensionality reduction constraints of matrix operations, the local image sets and their image descriptions corresponding to all multimodal literature are flattened sequentially to construct a globally citationable image sequence. With global image description sequence ,in Total number of images; Define a global image object. With evidence slices Association discriminant function If and only if there is a contextual referential relationship between the two. Otherwise, it is 0; a bidirectional mapping correlation matrix is ​​constructed accordingly. Its elements satisfy Construct a multimodal mapping set This enables the decoupling of text and graphics and the bidirectional mapping association of multimodal documents.

[0010] Furthermore, the specific method of step S2 includes the following sub-steps: S2-1, Set the general large language model as... With the aforementioned citation evidence library The text paragraphs and multimodal document content serve as contextual input for prompts, based on a preset set of prompt word templates. trigger Generate initial instruction question-and-answer pairs ,in , The initial total number of instructions; for each initial problem , using the Semantic rewriting and scenario expansion are performed to generate a set of question variants that include colloquial and scenario-based features. And define the multimodal interaction intent tag as ,in For the number of variants generated, The number of intent categories and For the corresponding multi-hot or one-hot encoding; construct a draft instruction question and answer set accordingly. , For the traversal index of the query variant; S2-2, The draft set of instructions and questions. We will import a customized questionnaire system to receive feedback evaluation matrices from domain experts for each draft. ;in, To represent the accuracy of factual audit scores, The standard answer after adversarial error correction. To align the corrected multimodal intent with references; set an acceptance threshold. Eliminate those that meet the requirements The draft entries, regarding the reserved entries replace ,use renew And convert the standardized fine-tuning serialization template into instruction fine-tuning sample pairs. To form a high-quality instruction fine-tuning corpus ,in For input prompt sequence, The target output sequence after expert verification. For sample size; S2-3. Initialize the network weight parameters of the multimodal large model of the base as follows: In the first stage, the aforementioned Continue pre-training in the domain, assuming a single training sequence is... , For sequence time step index, The pre-training sequence length is used, and an autoregressive language modeling loss is employed:

[0011] Pre-trained parameters are obtained through gradient descent. Loading in the second phase Using the Perform supervised fine-tuning, assuming the first... The target output sequence of the sample is , For the length of the target output sequence, a modified supervised fine-tuning loss function is used:

[0012] The model underwent parameter alignment and fine-tuning, ultimately converging to obtain the parameters of the large-scale grassland industry model with grassland industry professional interaction capabilities. .

[0013] Furthermore, the specific method of step S3 includes the following sub-steps: S3-1. Define the referable evidence database as follows: A joint extraction model based on a pre-defined grassland domain dictionary and a large language model. Extract entity sets from them Attribute set and relation sets ,in , , These represent the total number of entities, attributes, and relationships extracted, respectively. , , For the corresponding traversal index; construct a domain ontology for grassland production. , of which Concept set, A set of hierarchical relationships of concepts. Define a set of conceptual attributes; with the aforementioned For the pattern layer, the extracted , , Mapping generates a set of triples in the form of a resource description framework. ,in For the head entity, For tail entities, To establish the current relationship between the two, a hierarchical knowledge graph of the entire grassland industry process is constructed. ; S3-2, Extract the cited evidence library Text slice collection With image feature set The dimension of the shared semantic space is set to be Call the pre-trained multimodal embedding model To obtain the text vector Image vector representation A hierarchical navigable small-world indexing algorithm is used to construct a vector index. And store the vector and its identifier in the multimodal vector library:

[0014] in A globally unique identifier assigned to each type of modal object; S3-3, Define the cross-modal alignment mapping function Its input is a globally unique identifier. The output is an aligned pair of knowledge graph objects and vector library objects. ,in Indicates that The corresponding map nodes or map elements, Indicates that The corresponding vector library entries; based on this, a dual-track knowledge base is constructed. It is used to support simultaneous retrieval of structured and unstructured data.

[0015] Furthermore, the specific method of step S4 includes the following sub-steps: S4-1. Initialize the workflow architecture ,in As a central routing intelligent agent, To cover different production stages Domain expert sub-agents The total number of domain expert sub-agents; receiving multimodal query requests from users. ,in For text query, For image queries; by Calling the intent recognition and planning function complex query requests Parsed as a directed acyclic dependent sequence ,in The total number of subtasks after decomposition, and for each subtask Generate task feature vectors Define the scheduling mapping matrix. And calculate the allocation probability. Subtasks Scheduled until satisfied Domain expert sub-intelligent agents ; S4-2, For the activated domain expert sub-agent Extract user memory profile ,in This serves as a historical window for the recent rounds of dialogue. Define a retrieval rewriting function to include long-term personalized features such as region and planting scale. Generate standard search instructions Parallel-triggered dual-path retrieval: one path is based on the knowledge graph. Inference functions based on rule traversal and graph constraints Obtain the set of logical relationship paths ,in The total number of paths for recall; another path will Encoded as query vector and in the Normalized cosine similarity between China and Israel:

[0016] Compared with the entry vectors in the vector library Calculate the matching degree and recall the desired outcome. The former A set of cross-modal semantic entries ,in For similarity threshold, This refers to the number of items recalled. S4-3, Constructing a Knowledge Integration Module The precise logical relationship path set is processed through a cross-attention mechanism. With cross-modal semantic slice set Deduplication and weight reallocation are performed to generate a multimodal retrieval enhancement context sequence that integrates structured and unstructured data. Define the concatenation operator This indicates that the sequence is concatenated according to a preset template to construct a combined input prompt. And input the mounting parameters The large-scale model in the grassland industry is subjected to autoregressive decoding to output the final solution.

[0017] in For length is The candidate output sequence, The time step decoding index for the output sequence, the This serves as the final professional answer for the entire grass industry production process, returned to the user.

[0018] The beneficial effects of this invention are as follows: 1. This invention constructs a dual-track knowledge foundation integrating a structured knowledge graph and an unstructured multimodal vector library, and innovatively designs a dual-path retrieval mechanism that combines logical pathfinding and semantic similarity. The large model is strictly constrained by the strong logical evidence and multimodal literature retrieved when generating responses, solving the problem of general large models being prone to fabrication in specialized issues such as grass seed selection and disease control, thus ensuring the safety and scientific nature of agricultural decision-making.

[0019] 2. This invention changes the high-cost model that relies entirely on manually writing question-and-answer pairs. By utilizing a general large model to pre-generate drafts containing colloquial and scenario-based variations, domain experts only need to perform factual adversarial error correction and intent verification. This reverse feedback flow not only significantly improves the efficiency of corpus construction but also ensures that the data injected during the instruction fine-tuning stage has extremely high professional value, enabling the final grassland industry large model to perfectly adapt to the real interaction habits of farmers and herdsmen.

[0020] 3. This invention designs a system architecture of "central routing + domain expert sub-agent". Through precise intent decomposition, task routing, and the attachment of long and short-term memory archives (including personalized features such as region and planting scale), the system can cross multiple business silos in "breeding-planting-harvesting-sales" and provide coherent, personalized, and globally optimal decision support services for complex mixed text and image queries (such as uploading disease images and combining them with regional historical climate to inquire about countermeasures). Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the overall process of this method; Figure 2 This is an example diagram of a questionnaire survey conducted using this method. Detailed Implementation

[0022] The specific embodiments of the present invention are described below to facilitate understanding of the invention by those skilled in the art. However, it should be understood that the described embodiments are only a part of the embodiments of the present invention, and not all of them. The present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0023] like Figure 1 As shown in the overall process, in one embodiment of the present invention, taking the scenario of "diagnosis and fertilization of overwintering diseases in alfalfa" as an example, the overall process includes: collecting and cleaning multimodal data from the entire grass industry process; constructing a basic corpus for model pre-training and a referable evidence library after semantic slicing and bidirectional graph-text mapping; using a general large model to pre-generate question-answer drafts and introducing domain experts for adversarial error correction, generating high-quality instruction fine-tuning corpus to complete the pre-training and supervised fine-tuning of the large model in the grass industry domain; extracting entity relationships from the evidence library to construct a structured knowledge graph of the entire grass industry process, and combining multimodal embedding technology to establish an unstructured vector library, forming a dual-track knowledge base supporting synchronous retrieval; constructing a workflow that coordinates central routing and expert sub-agents, combining user memory archives to perform dual-track retrieval of graph logic and vector semantics in the dual-track base, fusing and generating context, and having the large model infer the final answer. Detailed steps include: S11. Collect multimodal raw datasets covering the entire process of grassland production. For example, in the alfalfa scenario, an unstructured text collection Include Documents, such as document 1 This is the second document in the "Alfalfa Overwintering Cultivation Technology Manual". Examples include pest reports issued by the Agricultural Department of xx Province; image collection. Include Images, such as the first image. The second image shows alfalfa root diseases photographed by farmers. Leaf diagrams showing different nutrient deficiency symptoms, etc. Text cleaning functions were used. Remove garbled characters and non-standard characters from documents using image cleaning functions. Images with too low resolution or blurriness are filtered out, and a privacy information masking function is used for desensitization. The farmer's name (e.g., replacing "Zhang San" with "Farmer A") and contact information in the technical guidance records are masked to obtain a standardized dataset. That is, a clean text collection after desensitization and cleaning. With image set Based on the preset information entropy threshold (set up ) and the scoring function of structured degree (set up (representing highly structured text), execute parameterized data stream routing functions. The text "Fundamentals of Forage Science," which has broad content and a low structured score, was included in the basic training corpus. The "Standard for Prevention and Control of Alfalfa Root Rot" and the "Guidelines for Soil Testing and Fertilizer Recommendation," which are highly structured, meet information entropy standards, and are extremely accurate, were deemed to have high citation value and were included in the citationable evidence database. .

[0024] S12. Regarding the aforementioned citation evidence library Long text collections in For any long text (For example, the "Overwintering Cultivation Technology Manual" uses a slicing algorithm based on sliding windows and semantic boundary symbols.) Semantic slicing is performed to obtain evidence slice sequences. For example, the manual was divided into... A semantically coherent slice, of which the first slice "Alfalfa should be given additional phosphorus and potassium fertilizer before winter, with 15-20 kg of superphosphate per acre..." Summarize all the long text segments to form a total. A collection of evidence slices. Define the feature space of multidimensional metadata. Slicing is done through metadata extraction functions. (i.e., the above) Generate the corresponding metadata vector Specifically, the assigned value is: grass production stage. "Overwintering Management", Suitable Regions Grass species in the North China Plain "Alfalfa". The system assigns a globally unique identifier to this slice. Based on this, a set of evidence text slices with metadata tags is constructed. If it contains tuples .

[0025] S13. Regarding the aforementioned citation evidence library Multimodal literature For any multimodal literature (For example, the "Atlas of Alfalfa Diseases," which includes numerous real-life photos of disease spots, uses a layout analysis model.) Parse it into a collection of page text blocks (For example, parsing out) (text paragraphs) and image set (If parsed out) (Illustrations of diseases). One of the images depicts "brown spots appearing on the root collar." Based on their spatial adjacency relationship from Selecting a set of adjacent text blocks (For example, extracting the caption "Figure 3: Early Root Rot Lesions" and the main text analysis, which are located directly below the image), and describing the function using a visual language model. Generate structured image descriptions "Typical symptoms of alfalfa root rot include browning of the vascular bundles at the root collar." To meet the matrix dimensionality reduction constraints for subsequent map and vector construction, all local images from the literature were sequentially flattened to construct a matrix containing... Zhang Tu's globally referenceable image sequence and corresponding global image description sequence By using the correlation discriminant function When the global image (i.e., the root spot diagram above) and evidence slices describing "root rot control measures" When there is a contextual referential relationship, determine Otherwise, it is 0; a bidirectional mapping correlation matrix is ​​constructed accordingly. Set the corresponding element values ​​in this matrix. Finally, a multimodal mapping set is constructed based on this matrix. For example, the set contains elements This will completely decouple the text and images and establish a two-way mapping relationship among the many modal documents in the "Atlas of Alfalfa Diseases".

[0026] S21. Utilizing a general large model (This example uses the Qwen-Max model) combined with an evidence base. The text paragraphs and multimodal document content serve as contextual input for prompts, based on a preset set of prompt word templates. (For example, set the template as: "Generate a question-and-answer pair for farmers based on the following alfalfa disease prevention and control text") Trigger Generate initial instruction question-and-answer pairs Set the initial total number of instructions For example, the first initial problem generated. "Please briefly describe the key management techniques for alfalfa before winter," corresponding to the initial... "Harvesting should cease before winter, and phosphorus and potassium fertilizers should be applied more frequently." This applies to each initial issue. , using the Semantic rewriting and scenario expansion are performed to generate a set of question variants that include colloquial and scenario-based features. The generated variants include How should alfalfa be managed before winter? How to ensure alfalfa survives the winter safely? "What preparations need to be made for alfalfa before winter?" Meanwhile, the multimodal interaction intent label is defined as... Set the number of intent categories (Note: This may refer to text-only interaction, disease images, or soil / weather tables.) If the overwintering management and fertilization issue is frequently accompanied by soil testing and disease consultation, then its multimodal interaction tag should be set to the corresponding multi-thermal coding. Based on this, a draft set of instruction questions and answers containing a total of 15,000 sets of data was generated. .

[0027] S22. The aforementioned instruction question and answer draft set A customized questionnaire system was introduced, and domain experts were brought in for "human-machine collaboration" verification. This questionnaire system required experts to perform the following core tasks: 1) List the most common decision-making scenarios encountered throughout the entire forage production process (e.g., "diagnosing the cause of yellow spots on alfalfa leaves" or "selecting suitable forage grasses for newly reclaimed saline-alkali land"); 2) Provide typical user questions and their expected professional answers; 3) Imagine colloquial expressions used by ordinary users (e.g., changing "developing a first-harvest time plan" to "when is the best time to harvest my alfalfa for the first time?"); 4) Identify key information that needs to be combined with images or tables (e.g., soil composition tables). The system receives feedback evaluation matrices from domain experts for each draft input. For example, regarding the draft generated above, if experts find that the machine-generated answer ignores regional soil differences, they will provide a factual review score representing accuracy. Set acceptance threshold The system automatically rejects the one that meets the requirements. The unqualified items. If the expert's score is If the entry is incorrect, it will be retained, and experts will conduct further adversarial error correction to determine the standard answer. The revised version reads, "In the saline-alkali lands of North China, potassium sulfate should be used preferentially over potassium chloride to prevent salt accumulation," and adds multimodal intent and citation association annotations. The instruction states, "Users must be prompted to upload soil testing reports." replace ,use renew The original instructions and answer texts provided by experts are standardized by correcting typos, unifying professional terminology, and standardizing unit formats. They are then converted into instruction fine-tuning samples according to a fine-tuning serialization template. To form a high-quality instruction fine-tuning corpus The sample size .

[0028] Initialize the network weight parameters of the base multi-modal large model as ; In the first stage, use for domain continuation pre-training. Let a single training sequence be , for example, extract a training sequence of length characters from "Fundamentals of Forage Science", which contains a local subsequence "root rot of alfalfa" describing the physiological characteristics of alfalfa (i.e., "purple", "flower", ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​("How to fertilize before winter?") and the generated target output prefix ("Additional measures") are used together as conditions to calculate the predicted target character. Find the conditional probability distribution of ("phosphorus") and calculate its negative log-likelihood loss. The outer loop of the lost function traverses all Each expert fine-tunes the sample; the inner layer adjusts each sample... The decoding path is summed step by step, and the model is fine-tuned to align the parameters. Finally, the parameters of the large model for the grassland industry with professional interaction capabilities are obtained by convergence. .

[0031] S31. Define the referable evidence library as... A joint extraction model based on a pre-defined grassland domain dictionary and a large language model. Extract entity sets from them Attribute set and relation sets For example, setting the total number of entities to be extracted. Total number of attributes Total number of relationships Taking the overwintering scenario of alfalfa as an example, the extracted specific entities include... "Alfalfa" "Alfalfa root rot" "Carbendazim" and "Superphosphate," etc.; the specific properties extracted include... "Lethal temperature: -20℃" and "Suitable soil pH: 7.0-8.0", etc.; the specific relationships extracted include... "Easily infected" "Prevention and treatment drugs" and "Suitable base fertilizer," etc. Constructing a domain ontology oriented towards grassland production. The concept set It includes high-level abstract concepts such as "pasture," "diseases," "pesticides," and "fertilizers," forming a hierarchical relationship. It stipulates that "alfalfa" is a subcategory of "forage grass". The definition of "disease" includes attributes such as "location of disease"; [The following is a separate, unrelated sentence:] For the pattern layer, the extracted , , Mapping generates a set of triples in the form of a resource description framework. For example, the generated specific triplets include those representing the disease-health relationship (alfalfa, susceptible to infection, alfalfa root rot), those representing the treatment relationship (alfalfa root rot, preventive medication, carbendazim), and those representing winter nutrient supplementation (alfalfa, suitable base fertilizer, superphosphate), thereby constructing a hierarchical knowledge graph of the entire grassland industry process with a rigorous logical system. .

[0032] S32. Extract the cited evidence library. Text slice collection (For example A slice, containing slices "Use tebuconazole or carbendazim for root drenching before overwintering to prevent root rot..." and image feature set (For example A picture, containing images (A photograph of browning lesions on the vascular bundles of alfalfa roots), with the shared semantic space dimension set to [value missing]. Call the pre-trained multimodal embedding model (In this example, the CLIP model is used), resulting in text vectors. (i.e., text slices) Mapped to a 1024-dimensional dense vector Image vector representation (i.e., lesion image) Mapped to a 1024-dimensional vector in a unified space A hierarchical navigable small-world indexing algorithm is used to construct a vector index. This supports extremely fast retrieval of near nearest neighbors for massive high-dimensional vectors, and stores the vectors and their identifiers in a multimodal vector library.

[0033] in A globally unique identifier assigned to various modal objects, such as a text slice. Assigning disease images .

[0034] S33. Define the cross-modal alignment mapping function. Its input is a globally unique identifier. The output is an aligned pair of knowledge graph objects and vector library objects. For example, setting a unified identifier. The function maps it to a tuple: where the map elements are... Pointing to entity nodes in the graph "Alfalfa root rot", and the vector library entry It then precisely points to an image containing the characteristics of the disease. Vectors and related prevention and control manual text paragraphs The set of vectors. Based on this, an anchor mapping is established between the nodes of the graph topology and the high-dimensional continuous vector space, constructing a dual-track knowledge base. It is used to support the synchronous dual-path retrieval of the large model between structured logical rules (what medicine to administer) and unstructured semantic features (what lesions look like) in subsequent question-answering services.

[0035] S41. Initialize the workflow architecture ,in For the central routing agent, set the total number of domain expert sub-agents. These are sub-agents corresponding to five different production stages: "breeding," "planting management," "plant protection (disease diagnosis)," "harvesting," and "sales." to Users initiated multimodal query requests via mobile devices before winter in November. Text query "The leaves of my alfalfa are turning yellow and there are black spots on the roots. It's almost winter, what pesticide should I use?" (Image search) This is a close-up image of the blackened vascular bundles at the roots of alfalfa, taken and uploaded in real-time by a user. Calling the intent recognition and planning function complex query requests Parsed as a directed acyclic dependent sequence For example, the total number of subtasks after decomposition The first subtask "Diagnosing root diseases in images," the second subtask "Develop a winter fertilization and pesticide application plan." And for each sub-task... Generating hidden layer representations of task feature vectors Define the scheduling mapping matrix. And calculate the allocation probability. ,by For example, through matrix multiplication and After calculation, the probability value of the third dimension (corresponding to the plant protection sub-agent) is the highest, that is... This will transform the subtasks Dispatch to domain expert sub-agent Execution; similarly, fertilization-related matters. Dispatch to planting management sub-agent implement.

[0036] S42, for the activated domain expert sub-agent (For example, those responsible for disease diagnosis) Extract user memory profile Upon learning of the recent multiple rounds of dialogue historical windows The record states "watered for overwintering just last week," indicating a long-term personalized characteristic. The record states that the farmer is located in Cangzhou City, Hebei Province (saline-alkali land), with a planting area of ​​500 mu. Define the retrieval rewrite function. Generate standard search instructions "Diagnose root black spot disease and search for appropriate pesticides on alfalfa grown in saline-alkali soil in Cangzhou, Hebei, that has just been watered for overwintering"; then, a dual-path search is triggered in parallel: one path is based on the knowledge graph. Inference functions based on rule traversal and graph constraints Obtain the set of logical relationship paths For example, the total number of recall paths ,in (Alfalfa) [Susceptible to infection] root rot [Medication] Tebuconazole), (Alfalfa) [Environmental Susceptibility] (High humidity in saline-alkali soil induces root rot); another approach will... Encoded as query vector (Shared semantic space dimension) ), and in the The normalized cosine similarity formula and the entry vectors in the vector library are compared. Calculate the matching degree:

[0037] Recall satisfied (Set similarity threshold) (before) A set of cross-modal semantic entries Set the number of recalls For example, recalls These are historical root rot images from the vector library that are highly similar to the current lesions. and It is a text slice from the "Atlas of Alfalfa Diseases" concerning the treatment plan for saline-alkali land rot.

[0038] S43. Constructing a knowledge fusion module The precise logical relationship path set is processed through a cross-attention mechanism. With cross-modal semantic slice set Deduplication and weight reallocation are performed to generate a multimodal retrieval enhancement context sequence that integrates structured and unstructured data. Define the concatenation operator This indicates that the sequence is concatenated according to a preset template to construct a combined input prompt. This involves combining the user's original question (text and images), the retrieved knowledge context about root rot, and the user profile of the Cangzhou saline-alkali land area into a coherent Prompt sequence, which is then input into the onboard parameters. Autoregressive decoding was performed on a large-scale model in the grassland industry to output the final solution:

[0039] Candidate output sequence The length is Chinese characters, This is the decoding index for the time step of the output sequence. During the decoding process, when... At that time, the model relies on prior prompts. With the first 9 generated characters Calculate and output the 10th Chinese character that maximizes the joint probability. Returning to the user with the final professional answer covering the entire process of forage production: "Based on the pictures you uploaded and the climate conditions of Cangzhou, Hebei, your alfalfa is suffering from root rot, a disease that commonly occurs before winter. Recommendations: 1. Immediately apply tebuconazole suspension for root irrigation (based on the 'Alfalfa Disease Atlas'); 2. Considering your saline-alkali soil, do not use chlorine-containing fertilizers before winter. It is recommended to apply 20 kg of potassium sulfate compound fertilizer per acre to improve the root system's frost resistance."

Claims

1. A method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process, characterized in that: Includes the following steps: S1. Collect multimodal raw data covering the entire process of grassland production. The dataset was then cleaned and desensitized to obtain a standardized dataset. According to the preset division criteria, the Classified into basic training corpus and citationable evidence library and the above The long texts in the document are semantically sliced ​​and metadata-tagged, while the multimodal documents within them are decoupled from the text and image and associated with bidirectional mapping. S2, Utilizing a general large language model In conjunction with the above Pre-generate a draft of instruction question and answer that includes various question variations and multimodal interaction intents. After review and error correction by domain experts and multimodal annotation, a high-quality instruction fine-tuning corpus was formed. and with the stated Further pre-training of the base multimodal large model in the relevant domain was performed, using the aforementioned Supervised fine-tuning was performed to obtain a large-scale model for the grassland industry and its final parameters. ; S3, from the above Extracting professional knowledge elements from the grassland industry to construct a domain ontology and a hierarchical knowledge graph of the entire grassland industry process and the Text slices and image content are encoded into shared semantic space vectors and stored in a multimodal vector library. Thus constructing the structure described and Together they form a dual-track knowledge foundation ; S4. Construct a central routing agent. and multiple domain expert sub-intelligent agents Workflow architecture Receive user multimodal query input , by the Perform intent recognition and task decomposition and schedule to the corresponding task. The scheduled Combine user memory profile In the Perform logical retrieval in the and in the Cross-modal similarity retrieval is performed, and the results of the two-way retrieval are fused to form a multimodal retrieval enhanced context. Then the above With the Common input to the mounting parameters A large-scale model in the grassland industry was developed to generate the final solution output. .

2. The method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process as described in claim 1, is characterized in that, Step S1 specifically includes the following steps: S1-1. Collect multimodal raw datasets covering the entire process of grassland production. ,in It is a collection of unstructured text. For a collection of images, and The initial number of text and images are defined; a text cleaning function based on regular expressions and a domain dictionary is defined. Image cleaning functions Desensitization function with privacy information mask Regarding the above After deduplication, error correction, format standardization, and data anonymization, a standard dataset is obtained. ,in , According to the preset information entropy threshold With the structured degree scoring function Define the partitioning function: The Classified into basic training corpus and citationable evidence library ,in Indicates basis and A parameterized data flow routing function that determines the trainability and citationability of samples and completes the aggregation. S1-2, Regarding the aforementioned citation evidence library Long text collections in For any long text A slicing algorithm based on sliding window and semantic boundary markers is adopted. Semantic slicing is performed to obtain evidence slice sequences. And summarize them to form a collection of evidence slices. ,in, The number of slices generated from this long text. Define the total number of slices; define the multidimensional metadata feature space. ,in , and These represent the stages of grass production, suitable regions, and grass species, respectively; metadata extraction functions are used. Slice each piece of evidence Generate the corresponding metadata vector And assign a globally unique identifier to each piece of evidence. Based on this, construct a set of evidence text slices with metadata tags. ; S1-3, Regarding the aforementioned citation evidence library Multimodal literature For any multimodal literature Using a layout analysis model Parse it into a collection of page text blocks With image collection ,in and These represent the number of text blocks and images parsed from the document, respectively; for any image object Based on their spatial adjacency relationship from Selecting a set of adjacent text blocks And describe the function through a visual language model. Generate structured image descriptions To satisfy the dimensionality reduction constraints of matrix operations, the local image sets and their image descriptions corresponding to all multimodal literature are flattened sequentially to construct a globally citationable image sequence. With global image description sequence ,in Total number of images; Define a global image object. With evidence slices Association discriminant function If and only if there is a contextual referential relationship between the two. Otherwise, it is 0; a bidirectional mapping correlation matrix is ​​constructed accordingly. Its elements satisfy Construct a multimodal mapping set This enables the decoupling of text and graphics and the bidirectional mapping association of multimodal documents.

3. The method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process as described in claim 1, is characterized in that, Step S2 specifically includes the following steps: S2-1, Set the general large language model as... With the aforementioned citation evidence library The text paragraphs and multimodal document content serve as contextual input for prompts, based on a preset set of prompt word templates. trigger Generate initial instruction question-and-answer pairs ,in , The initial total number of instructions; for each initial problem , using the Semantic rewriting and scenario expansion are performed to generate a set of question variants that include colloquial and scenario-based features. And define the multimodal interaction intent tag as ,in For the number of variants generated, The number of intent categories and For the corresponding multi-hot or single-hot encoding; Based on this, a draft set of instruction questions and answers was constructed. , For the traversal index of the query variant; S2-2, The draft set of instructions and questions. We will import a customized questionnaire system to receive feedback evaluation matrices from domain experts for each draft. ;in, To represent the accuracy of factual audit scores, The standard answer after adversarial error correction. To align the corrected multimodal intent with references; set an acceptance threshold. Eliminate those that meet the requirements The draft entries, regarding the reserved entries replace ,use renew And convert the standardized fine-tuning serialization template into instruction fine-tuning sample pairs. To form a high-quality instruction fine-tuning corpus ,in For input prompt sequence, The target output sequence after expert verification. For sample size; S2-3. Initialize the network weight parameters of the multimodal large model of the base as follows: ; The first stage adopts the above Continue pre-training in the domain, assuming a single training sequence is... , For sequence time step index, The pre-training sequence length is used, and an autoregressive language modeling loss is employed: Pre-trained parameters are obtained through gradient descent. Loading in the second phase Using the Perform supervised fine-tuning, assuming the first... The target output sequence of the sample is , For the length of the target output sequence, a modified supervised fine-tuning loss function is used: After fine-tuning the parameters of the model for alignment, the final convergence yields a large-scale model parameter set for the grassland industry that possesses professional interaction capabilities. .

4. The method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process as described in claim 1, is characterized in that, Step S3 specifically includes the following steps: S3-1. Define the referable evidence database as follows: A joint extraction model based on a pre-defined grassland domain dictionary and a large language model. Extract entity sets from them Attribute set and relation sets ,in , , These represent the total number of entities, attributes, and relationships extracted, respectively. , , For the corresponding traversal index; construct a domain ontology for grassland production. , of which Concept set, A set of hierarchical relationships of concepts. Define a set of conceptual attributes; with the aforementioned For the pattern layer, the extracted , , Mapping generates a set of triples in the form of a resource description framework. ,in For the head entity, For tail entities, To establish the current relationship between the two, a hierarchical knowledge graph of the entire grassland industry process is constructed. ; S3-2, Extract the cited evidence library Text slice collection With image feature set The dimension of the shared semantic space is set to be Call the pre-trained multimodal embedding model To obtain the text vector Image vector representation A hierarchical navigable small-world indexing algorithm is used to construct a vector index. And store the vector and its identifier in the multimodal vector library: in A globally unique identifier assigned to each type of modal object; S3-3, Define the cross-modal alignment mapping function Its input is a globally unique identifier. The output is an aligned pair of knowledge graph objects and vector library objects. ,in Indicates that The corresponding map nodes or map elements, Indicates that The corresponding vector library entries; based on this, a dual-track knowledge base is constructed. It is used to support simultaneous retrieval of structured and unstructured data.

5. The method for constructing a large-scale domain model and providing question-answering services for the entire grass industry production process as described in claim 1, characterized in that, Step S4 specifically includes the following steps: S4-1. Initialize the workflow architecture ,in As a central routing intelligent agent, To cover different production stages Domain expert sub-agents The total number of domain expert sub-agents; receiving multimodal query requests from users. ,in For text query, For image queries; by Calling the intent recognition and planning function complex query requests Parsed as a directed acyclic dependent sequence ,in The total number of subtasks after decomposition, and for each subtask Generate task feature vectors ; Define the scheduling mapping matrix And calculate the allocation probability. Subtasks Scheduled until satisfied Domain expert sub-intelligent agents ; S4-2, For the activated domain expert sub-agent Extract user memory profile ,in This serves as a historical window for the recent rounds of dialogue. It includes long-term personalized characteristics such as region and planting scale; Define the rewrite function Generate standard search instructions Parallel-triggered dual-path retrieval: one path is based on the knowledge graph. Inference functions based on rule traversal and graph constraints Obtain the set of logical relationship paths ,in The total number of paths for recall; another path will Encoded as query vector and in the Normalized cosine similarity between China and Israel: Compared with the entry vectors in the vector library Calculate the matching degree and recall the desired outcome. The former A set of cross-modal semantic entries ,in For similarity threshold, This refers to the number of items recalled. S4-3, Constructing a Knowledge Integration Module The precise logical relationship path set is processed through a cross-attention mechanism. With cross-modal semantic slice set Deduplication and weight reallocation are performed to generate a multimodal retrieval enhancement context sequence that integrates structured and unstructured data. Define the concatenation operator This indicates that the sequence is concatenated according to a preset template to construct a combined input prompt. And input the mounting parameters The large-scale model in the grassland industry is subjected to autoregressive decoding to output the final solution. in For length is The candidate output sequence, The time step decoding index for the output sequence, the This serves as the final professional answer for the entire grass industry production process, returned to the user.