Method and System for Dynamic Expression and Interaction of Cabbage Knowledge Based on Multimodal Fusion

Through multimodal data fusion and dynamic knowledge graph construction, the problem of insufficient multimodal information integration in traditional agricultural knowledge management systems is solved, and accurate dynamic management and efficient decision-making support of the kale planting process are realized.

CN119862954BActive Publication Date: 2025-07-29BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510340192.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-29
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Traditional agricultural knowledge management systems lack the ability to integrate multimodal information, cannot reflect changes in the field environment in real time, lack intelligent adjustment and efficient decision-making capabilities, and have a single interaction method, which is difficult to meet the flexible needs of users.

Method used

Through multimodal data acquisition and fusion, a dynamic knowledge graph for cabbage planting management is constructed, and a cross-modal Transformer model is used for feature alignment and knowledge reasoning is realized to achieve accurate dynamic management and decision support of the cabbage planting process.

Benefits of technology

It improves the accuracy and intuitiveness of the expression of kale cultivation knowledge, provides farmers with efficient decision-making support, and breaks the limitations of single text or table output of traditional cultivation guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862954B_ABST
    Figure CN119862954B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for dynamic expression and interaction of cabbage knowledge based on multimodal fusion, which relates to the technical field of agricultural planting management. The method includes: acquiring multimodal data of cabbage planting management, extracting features from the multimodal data to obtain multimodal features, and performing feature alignment to obtain final fusion features; constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features; for the cabbage planting interaction data to be processed, invoking the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results. This application realizes the precise dynamic management and decision-making support of the cabbage planting process by implementing processes such as multimodal data collection and fusion, dynamic knowledge graph construction, and intelligent interaction functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural planting management, and particularly to a method and system for dynamic expression and interaction of cabbage knowledge based on multimodal fusion. Background Art

[0002] As an important economic vegetable, the planting management of cabbage covers multiple key aspects such as growth monitoring, environmental regulation, pest and disease control, and optimization of fertilization and irrigation. However, traditional agricultural knowledge management systems usually present knowledge in a single form of text or tables, lacking the ability to effectively integrate multimodal information in the actual field scenario, resulting in significant limitations in practical applications.

[0003] On the one hand, the construction method of the static knowledge base is difficult to reflect the dynamic changes of the field environment in real time, and it is impossible to intelligently adjust the management strategy according to real-time sensor data or user feedback, resulting in lack of pertinence in the guiding opinions; on the other hand, these systems lack efficient problem reasoning and decision-making capabilities, and cannot solve cross-domain or unstructured problems, such as the identification and prevention suggestions for newly occurring pests and diseases, affecting the decision-making efficiency of farmers. In addition, the traditional interaction methods are also relatively single, mostly presented in the form of fixed-format retrieval interfaces, which are difficult to meet the flexible questioning needs of users and lack intuitiveness and interactivity, especially with significant gaps in multimodal interaction functions such as image recognition and voice input. Summary of the Invention

[0004] The present invention provides a method and system for dynamic expression and interaction of cabbage knowledge based on multimodal fusion, which realizes precise dynamic management and decision support for the cabbage planting process through multimodal data collection and fusion, construction of a dynamic knowledge graph, and intelligent interaction functions.

[0005] The present invention provides a method for dynamic expression and interaction of cabbage knowledge based on multimodal fusion, including:

[0006] Obtaining multimodal data for cabbage planting management, where the multimodal data includes image data of cabbage growth, environmental time-series data, and annotated text data of cabbage planting management;

[0007] Extracting features from the multimodal data to obtain multimodal features, and calling a cross-modal Transformer model to align the multimodal features to obtain final fusion features;

[0008] Constructing a dynamic knowledge graph for cabbage planting management based on the final fusion features;

[0009] For the cabbage planting interaction data to be processed, calling the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results.

[0010] In some embodiments, the multimodal features include the image features of the image data, the temporal features of the environmental temporal data, and the text features of the annotated text data. The feature extraction of the multimodal data to obtain multimodal features includes:

[0011] Invoking a VisionTransformer model to perform a disease classification task and a leaf health status estimation task on the image data to obtain the image features of the image data. Among them, the disease classification task is used to classify cabbage diseases for the image data, and the leaf health status estimation task is used to estimate the health status of cabbage leaves for the image data. The VisionTransformer model is trained through a disease classification loss and a status estimation loss;

[0012] Invoking a temporal Transformer model based on causal attention to extract features from the environmental temporal data to obtain temporal features;

[0013] Invoking a Tokenizer model to encode the annotated text data into a Token ID sequence;

[0014] Obtaining prompt words of specific domain knowledge, and inputting the prompt words and the Token ID sequence into a RoBERTa model for context semantic feature extraction to obtain the text features of the annotated text data.

[0015] In some embodiments, the invocation of a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain the final fused features includes:

[0016] Performing cross-modal interaction calculations on the multimodal features through the cross-modal attention mechanism of the cross-modal Transformer model to obtain image-temporal interaction features, temporal-text interaction features, and image-text interaction features respectively;

[0017] Fusing the image-temporal interaction features, the temporal-text interaction features, and the image-text interaction features to obtain a joint feature representation;

[0018] Performing feature dimension transformation processing on the joint feature representation to obtain the final fused features.

[0019] In some embodiments, the cross-modal Transformer model is trained by constructing an overall loss function. The construction process of the overall loss function includes:

[0020] Calculate the feature similarity between modalities for the multi-modal features, and construct a modality contrast loss based on the feature similarity. The modality contrast loss includes an image-text modality contrast loss, an image-temporal modality contrast loss, and a temporal-text modality contrast loss;

[0021] Reconstruct the final fusion features through the decoder of the cross-modal Transformer model to obtain the feature representations of each modality;

[0022] Construct a modality reconstruction loss based on the feature representations of each modality and the multi-modal features;

[0023] Weight the modality contrast loss and the modality reconstruction loss respectively through preset contrast weights and reconstruction weights, and sum the weighted results to obtain an overall loss function.

[0024] In some embodiments, constructing a dynamic knowledge graph for cabbage planting management based on the final fusion features includes:

[0025] Align the multi-modal data of cabbage planting management with the final fusion features to obtain structured multi-modal input data;

[0026] Extract entities from the multi-modal features to obtain corresponding entity nodes and form an entity set;

[0027] Extract relationships from the multi-modal input data to obtain the relationships between the entity nodes and form a relationship set;

[0028] Construct triples based on the entity set and the relationship set to obtain a primary knowledge graph;

[0029] Dynamically update the primary knowledge graph to obtain a dynamic knowledge graph for cabbage planting management.

[0030] In some embodiments, the dynamic update process of the primary knowledge graph includes:

[0031] Obtain real-time environmental data during the cabbage planting process, and aggregate the real-time environmental data through a time window to obtain window environmental features;

[0032] Determine the feature change rate of the window environmental features between time windows. When the feature change rate is greater than the feature change threshold, trigger an update operation on the primary knowledge graph;

[0033] Among them, the update operation on the primary knowledge graph includes:

[0034] Construct a relationship change matrix before and after the update according to the window environmental features;

[0035] Within the time window, merge the relationship change matrix with the original relationship set in the primary knowledge graph to obtain a new relationship set.

[0036] In some embodiments, after constructing the relationship change matrix before and after the update, the method further includes:

[0037] Determine the feature similarity between the relationship set before the update and the relationship set after the update;

[0038] Respectively determine the exclusive conflict score, attribute conflict score, and reasoning conflict score existing between the relationship set before the update and the relationship set after the update;

[0039] Sum up the exclusive conflict score, the attribute conflict score, and the reasoning conflict score to obtain a conflict degree score;

[0040] Take the difference between the feature similarity and the conflict degree score as the update conflict score. When the update conflict score is greater than a preset conflict threshold, determine that there is a conflict between the relationship set before the update and the relationship set after the update;

[0041] Adjust the relationship weights in the relationship set according to the conflict degree score to correct the conflict existing between the relationship set before the update and the relationship set after the update.

[0042] In some embodiments, the method further includes:

[0043] Determine the feature similarity between any two entity nodes in the primary knowledge graph. When the feature similarity is greater than a preset similarity threshold, merge the two entity nodes;

[0044] Extract the entity nodes and edges in the dynamic knowledge graph, and call a graph convolutional neural network to perform relationship prediction on the entity nodes and the edges to obtain predicted nodes and predicted relationships;

[0045] Merge the predicted nodes and predicted relationships with the entity set and relationship set in the dynamic knowledge graph.

[0046] In some embodiments, for the to-be-processed interactive data of cabbage planting, call the dynamic knowledge graph to retrieve the to-be-processed interactive data of cabbage planting, and obtain corresponding cabbage planting feedback results, including:

[0047] When the to-be-processed interactive data of cabbage planting is a cabbage image that needs to be diagnosed for diseases, call the VisionTransformer model to extract the disease characteristics of the cabbage image, and retrieve the corresponding disease nodes in the dynamic knowledge graph according to the disease characteristics;

[0048] Performing knowledge reasoning on the disease nodes through the dynamic knowledge graph to obtain the disease diagnosis result of the cabbage image as the cabbage planting feedback result;

[0049] When the cabbage planting interaction data to be processed is a prompt word for consulting cabbage planting, calling the dynamic knowledge graph to perform knowledge reasoning on the prompt word to obtain cabbage planting data as the cabbage planting feedback result;

[0050] When the cabbage planting interaction data to be processed represents the execution of personalized recommendation for cabbage planting, calling the dynamic knowledge graph to perform knowledge reasoning on the historical cabbage planting data to obtain the recommended content for cabbage planting as the cabbage planting feedback result.

[0051] The present invention also provides a multi-modal fusion-based dynamic cabbage knowledge expression and interaction system, including:

[0052] A feature extraction module for obtaining multi-modal data of cabbage planting management, where the multi-modal data includes image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management;

[0053] A feature fusion module for extracting features from the multi-modal data to obtain multi-modal features, and calling a cross-modal Transformer model to perform feature alignment on the multi-modal features to obtain the final fusion features;

[0054] A graph construction module for constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features;

[0055] A dynamic interaction module for retrieving the cabbage planting interaction data by calling the dynamic knowledge graph for the cabbage planting interaction data to be processed to obtain the corresponding cabbage planting feedback result.

[0056] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the multi-modal fusion-based dynamic cabbage knowledge expression and interaction method as described in any one of the above when executing the computer program.

[0057] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-modal fusion-based dynamic cabbage knowledge expression and interaction method as described in any one of the above.

[0058] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the multi-modal fusion-based dynamic cabbage knowledge expression and interaction method as described in any one of the above.

[0059] The method and system for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the present invention extract multimodal features of multimodal data for cabbage planting management, perform feature alignment to obtain final fusion features, and then construct a dynamic knowledge graph for cabbage planting management based on the final fusion features. In this way, multimodal data is uniformly embedded into the dynamic knowledge graph framework, improving the accuracy of cabbage planting knowledge expression. In addition, the present invention also provides an interaction function. For cabbage planting interaction data, it can be retrieved using the dynamic knowledge graph for feedback, breaking the limitations of the traditional single text or table output for planting guidance, and providing more intuitive and efficient decision-making support for farmers. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0061] Figure 1 is a schematic flowchart of the method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the present invention.

[0062] Figure 2 is a schematic principle framework diagram of the method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the present invention.

[0063] Figure 3 is a schematic structural diagram of the system for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the present invention.

[0064] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0066] The following describes the method and system for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion of the present invention with reference to the drawings. Figure 1 is a schematic flowchart of the method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the present invention, as Figure 1As shown, the method includes the following steps 101 to 104, which are specifically described below.

[0067] Step 101: Obtain multi-modal data for cabbage planting management.

[0068] During the process of cabbage planting management, generally, sensor devices are used to monitor the growth status and environment of cabbages in real time, and the collected data is uploaded. Here, the uploaded multi-modal data for cabbage planting management is obtained. The multi-modal data includes image data of cabbage growth, environmental time-series data, and labeled text data for cabbage planting management.

[0069] Regarding the image data of cabbage growth, the data content is the growth status of cabbage leaves, disease characteristics (such as the shape and color of downy mildew spots), and the surface condition of the soil. The data format is in JPEG and PNG formats. During the data collection process, the image collection frequency is dynamically adjusted according to the cabbage growth cycle. For example, it is collected once a week during the sowing period and daily during the high disease incidence period.

[0070] The data content of the environmental time-series data includes key environmental variables such as the temperature, humidity, light intensity, soil humidity, concentration, etc. of cabbage planting. The data format is standardized time-series stream data (such as JSON format). The data update frequency is automatically recorded once every 15 minutes to ensure real-time performance.

[0071] For the labeled text data of cabbage planting management, it is mainly used as expert data. Its data content is the disease description, prevention and control suggestions, and planting management knowledge of cabbages. The data feature is that its label is a pairing of multi-modal data of images and texts, which is used to construct a high-quality basic knowledge graph.

[0072] Step 102: Extract features from the multi-modal data to obtain multi-modal features, and call the cross-modal Transformer model to align the multi-modal features to obtain the final fused features.

[0073] Next, extract features from the multi-modal data to obtain multi-modal features. Specifically, for the image data in the image modality, call image feature extraction models such as convolution, residual, or Transformer to extract features and obtain image features. For the environmental time-series data in the time-series modality, call long short-term memory networks or Transformer models based on causal attention mechanisms to extract time-series features and obtain time-series features. For the labeled text data in the text modality, perform text feature extraction through Transformer models or BERT models to obtain text features.

[0074] Next, the cross-modal Transformer model will be called to align the multi-modal features to obtain the final fused features. Here, the cross-modal Transformer model is used as the feature fusion layer, and then the image features, text features, and temporal features are input into the feature fusion layer. Through feature fusion, the final fused features are obtained, which contain multi-modal information and are used to construct the knowledge graph subsequently.

[0075] Step 103: Construct a dynamic knowledge graph for cabbage planting management based on the final fused features.

[0076] The construction of the dynamic knowledge graph is a key process for the conversion of multi-modal data into structured knowledge, involving four processes: data preprocessing, entity relationship extraction, graph generation and update, knowledge fusion, and reasoning. Through these four processes, the multi-modal data is gradually converted into a dynamic and inferable graph, providing support for intelligent decision-making in agricultural application scenarios.

[0077] Here, entity sets and relationship sets are constructed respectively through multi-modal features such as the final fused features, image features, text features, and temporal features, and then triples are constructed to form a primary knowledge graph. And the primary knowledge graph is dynamically updated according to the real-time input cabbage planting environment data to obtain the dynamic knowledge graph of cabbage planting management.

[0078] Step 104: For the to-be-processed cabbage planting interaction data, call the dynamic knowledge graph to retrieve the cabbage planting interaction data and obtain the corresponding cabbage planting feedback result.

[0079] Here, the dynamic knowledge graph is used as the database for cabbage planting management. When performing dynamic interaction, for the to-be-processed cabbage planting interaction data input by the user, the dynamic knowledge graph can be called to retrieve the cabbage planting interaction data and obtain the corresponding cabbage planting feedback result, which is fed back to the user for decision-making.

[0080] The cabbage planting interaction data input by the user can be a cabbage image for disease diagnosis, a prompt word for consulting cabbage planting, or a personalized recommendation for performing cabbage planting. The dynamic knowledge graph can complete the interaction for these cabbage planting interaction data, perform knowledge reasoning, and give feedback content such as disease diagnosis results, cabbage planting data, cabbage planting suggestions, and recommended content for cabbage planting, facilitating the user to make decisions.

[0081] In an embodiment of the present invention, by extracting multi-modal features of multi-modal data for cabbage planting management and performing feature alignment, the final fused features are obtained, and then a dynamic knowledge graph for cabbage planting management is constructed based on the final fused features. In this way, the multi-modal data is uniformly embedded into the dynamic knowledge graph framework, improving the accuracy of cabbage planting knowledge expression. In addition, the present invention also provides an interaction function. For the interactive data of cabbage planting, the dynamic knowledge graph can be used for retrieval to provide feedback, breaking the limitation of the traditional planting guidance with a single text or table output, and providing more intuitive and efficient decision-making support for farmers.

[0082] In some embodiments, the multi-modal features include image features of image data, temporal features of environmental time-series data, and text features of annotated text data. The following will introduce the specific process of extracting features from multi-modal data to obtain multi-modal features in combination with Figure 2 to introduce.

[0083] As Figure 2 shown, in the process of feature extraction and fusion, for the image features of the image data, the VisionTransformer model is called to perform a disease classification task and a leaf health status estimation task on the image data, and the image features of the image data are obtained.

[0084] Here, before extracting the image features, the image data is divided into multiple image patches of a fixed size. Each patch is regarded as a sequence and input into the VisionTransformer model. The image patches will be encoded through multiple layers of Transformer, and each layer includes multi-head self-attention and a feed-forward network. Finally, a global feature vector is output, which synthesizes the information of the disease spot shape, color, and leaf health status.

[0085] In the process of extracting the image features of the VisionTransformer model, a disease classification task and a leaf health status estimation task will be performed. Among them, the disease classification task is used to classify cabbage diseases for the image data, and the leaf health status estimation task is used to estimate the health status of cabbage leaves for the image data. The VisionTransformer model uses a shared feature extraction layer and generates two output paths at the same time. One is the classification path, which is connected to a classifier (including a fully connected layer and a softmax function), and the other is the regression path, which is connected to a regression module (including a fully connected layer and a sigmoid function). When performing the disease classification task, the global feature vector is input into the fully connected layer of the classification path for mapping processing to obtain the probability distribution of the disease categories , Represents the predicted probability corresponding to the nth disease category. And the leaf health status estimation task belongs to a regression task. In the regression task, the global feature vector is input into the fully connected layer of the regression path for mapping processing, and the leaf health status estimation value is output, such as [0, 1]. The higher the estimated value, the healthier the cabbage leaf.

[0086] Finally, the probability distribution of the disease category and the leaf health status estimation value are combined to form the image features of the image data, denoted as .

[0087] In addition, the VisionTransformer model in the embodiments of the present invention is trained through the disease classification loss and the status estimation loss. Multi-task learning is introduced during the training process. The specific training process is introduced below.

[0088] The training samples input of the VisionTransformer model are cabbage leaf disease images, and at the same time, the annotation information of the disease area of each image is input, including the disease category label (used to perform the disease classification task for disease classification) and the leaf health status label (used to perform the leaf health status estimation task to define the leaf health status level based on the leaf area and color distribution).

[0089] During the training process of multi-task learning, disease classification is performed through the disease classification task, such as powdery mildew, downy mildew, etc., and leaf health status estimation is performed through the leaf health status estimation task, such as healthy, slightly damaged, severely damaged. When constructing the final loss function L, the loss functions of the two tasks are combined, and the formula is as follows:

[0090] (1)

[0091] Among them, is the disease classification loss function, using cross-entropy loss for disease classification, is the health status estimation loss function, using mean square error loss for leaf health status estimation. is a dynamically adjusted weight factor, which can be dynamically adjusted according to the task importance.

[0092] During the training process, the final loss function L is used for backpropagation in the VisionTransformer model to update the parameters of the model. When the maximum number of training epochs is reached or the final loss function L converges, the training is stopped.

[0093] For environmental time series data, here the time series Transformer model based on causal attention is called to extract features from the environmental time series data to obtain time series features.

[0094] Traditional long short-term memory networks have performance bottlenecks in long-term dependence modeling. In the embodiments of the present invention, a temporal Transformer model based on causal attention (Temporal Causal Transformer, TCT) is adopted to perform temporal feature extraction. During the temporal feature extraction process, the model input is environmental temporal data, such as temperature, humidity, light intensity, soil humidity, etc. Assuming that each piece of environmental temporal data contains t time steps and the feature dimension of each time step is D, the dimension of the input matrix can be denoted as [t, D]. The core idea of the causal attention mechanism is to use the causal attention mechanism to ensure that the current time step can only depend on previous time steps, ensuring that the temporal modeling conforms to physical causality. The formula is as follows:

[0095] (2)

[0096] In the above formula (2), is the causal mask matrix, ensuring that the current time step only focuses on past time steps. are the query vector, key vector, and value vector respectively, is the vector dimension, and T represents the transpose of the matrix. The input environmental temporal data [t, D] undergoes a linear transformation to map the features of each time step to dimensions. Then, after the calculation of the causal attention mechanism, a hidden layer feature matrix [t, H] is generated, where H is the hidden layer dimension. Each hidden layer contains a causal attention module and a feed-forward network, stacked M layers.

[0097] When the causal attention mechanism performs calculations, dynamic window processing is performed on the input temporal data, specifically by dynamically adjusting the time window size according to the fluctuation amplitude of the temporal data. For example: when environmental parameters such as temperature and humidity fluctuate slightly, a smaller window is used. When the environment changes violently, the window is enlarged to capture more context information. In addition, feature weighting is also performed, that is, the feature weights of key time points are assigned through attention, redundant data is ignored, and the feature output at the last time step is used as the temporal feature , and the formula is as follows:

[0098] (3)

[0099] In the above formula (3), represents the output result of the hidden layer at the last time step.

[0100] It is a feature vector with a fixed length, containing information on the dynamic impact of environmental data on crop growth. This feature can be directly input into the subsequent feature fusion model. For the growth environment data of cabbage (temperature, humidity, light intensity, soil moisture), the model can extract key time points (such as the outbreak of downy mildew caused by continuous low temperature) and generate features reflecting the dynamic changes.

[0101] For the labeled text data, the Tokenizer model is called to encode the labeled text data into a sequence of Token IDs. During the text feature extraction process, the labeled text data is encoded into a sequence of Token IDs by the Tokenizer model, represented as [CLS]+T+[SEP]. The sequence of Token IDs can directly perform semantic feature extraction. However, in order to guide the model to better generate relevant semantic information in a specific domain (such as the agricultural domain or the cabbage planting domain), the embodiments of the present invention construct a domain-specific Prompt template, such as "The best control method for the current description of cabbage diseases is", to improve the understanding ability of domain knowledge.

[0102] Here, the prompt words for obtaining specific domain knowledge are obtained, and the prompt words and the sequence of Token IDs are input into the RoBERTa model for context semantic feature extraction to obtain the text features of the labeled text data. Specifically, the input data of the RoBERTa model is in the form of "text encoding + Prompt template". First, the Prompt template is added to the text encoding (i.e., the sequence of Token IDs), so the input data is represented as [prompt]+[CLS]+T+[SEP], and then the input data is input into the RoBERTa model for context semantic feature extraction to obtain the text features , where f represents the feature fusion function that combines the vector of the [CLS] token and the vector of [prompt] . The text features can reflect the semantic information of the cabbage leaf symptoms, such as "the leaves are yellow and have brown spots".

[0103] In the embodiments of the present invention, for the multi-modal data of the image modality, text modality, and time series modality, different feature extraction models are called to perform the extraction of multi-modal features. And during the feature extraction process, the diagnosis of cabbage diseases, the estimation of the leaf health state, the determination of the dynamic changes of the environmental time series data on the growth state of cabbage, and the construction of semantic information specific to the cabbage planting domain are realized, increasing the knowledge richness of the subsequent constructed knowledge graph.

[0104] In some embodiments, considering agricultural scenarios, feature fusion of multimodal data (image data, time series data, and text data) is a core step in achieving intelligent analysis and decision-making. To address the issue of insufficient information exchange between modalities, a multimodal feature fusion method is proposed. This method uses a cross-modal Transformer model to align multimodal features to obtain the final fused features. This is described in detail below.

[0105] First, the cross-modal attention mechanism of the cross-modal Transformer model is used to perform inter-modal interaction calculations on multimodal features, and image-time sequence interaction features, time sequence-text interaction features, and image-text interaction features are obtained respectively.

[0106] like Figure 2 As shown, in an embodiment of the present invention, a cross-modal Transformer (CMT) model is used as a feature fusion model, and a multi-head attention mechanism is used on the feature fusion layer of the model to establish inter-modal interaction, thereby generating a corresponding joint feature representation. The CMT model can combine self-supervised learning optimization to improve the semantic consistency and alignment capabilities between modalities.

[0107] The cross-modal attention mechanism aims to achieve deep interaction between image, time series data, and text features, and to mine the correlation between modalities. During the calculation process, the multi-head attention mechanism uses the features of each modality as query, key, and value input respectively to calculate cross-modal interaction information. The formula is as follows:

[0108] (4)

[0109] In the above formula (4), the query vector The feature vector representing the source mode. Key vector and value vector Both represent the eigenvectors of the target mode. is the vector dimension, and T represents the transpose of the matrix.

[0110] In the process of intermodal interaction, there are three forms of interaction. One is image-time interaction, which combines image features and timing characteristics Perform association to obtain image-time interaction features , which combines visual and temporal dynamic change information to explore the relationship between the environment and crop growth status. In the cross-modal attention mechanism, the image features As the feature vector of the source modality, that is, the query vector , time series characteristics is the eigenvector of the target mode, that is, the key vector Sum vector , the formula is as follows:

[0111] (5)

[0112] In the above formula (5), is the calculation function of the multi-head attention mechanism. The following formulas (6) and (7) can be referred to and will not be elaborated.

[0113] The second is the temporal-text interaction, which correlates the temporal feature with the text feature to obtain the temporal-text interaction feature . This feature connects the temporal dynamics with the semantic expression of domain-specific knowledge, enhancing the understanding of process descriptions. In the cross-modal attention mechanism, the temporal feature is used as the feature vector of the source modality, that is, the query vector , and the text feature is the feature vector of the target modality, that is, the key vector K and the value vector . The formula is as follows:

[0114] (6)

[0115] The third is the image-text interaction, which correlates the image feature with the text feature to obtain the image-text interaction feature , representing a multi-dimensional scene description. In the cross-modal attention mechanism, the image feature is used as the feature vector of the source modality, that is, the query vector , and the text feature is the feature vector of the target modality, that is, the key vector K and the value vector . The formula is as follows:

[0116] (7)

[0117] Next, combining the three interaction results of the features, the image-temporal interaction feature, the temporal-text interaction feature, and the image-text interaction feature are fused to obtain the joint feature representation . The formula is expressed as:

[0118] (8)

[0119] In the above formula (8), is the processing function of feature fusion.

[0120] Finally, the joint feature representation is processed by feature dimension transformation to obtain the final fusion feature . Here, the joint feature representation is adjusted through linear transformation and non-linear activation function The feature dimension is expressed as:

[0121] (9)

[0122] In the above formula (9), represents a nonlinear activation function, and Represents the trainable parameters in the cross-modal Transformer model.

[0123] The embodiments of the present invention perform feature interaction between modalities through cross-modal attention, forming multiple feature interactions, which can represent more knowledge information, and feature fusion of multiple interaction features can effectively enhance the semantic consistency and robustness of cross-modal features.

[0124] In addition, in order to enhance the feature alignment capability of the cross-modal Transformer model as a feature fusion model, the generalization capability of the cross-modal Transformer model for feature fusion is provided. The embodiment of the present invention provides a training method for a cross-modal Transformer model. Specifically, the cross-modal Transformer model is obtained by training by constructing an overall loss function. The overall loss function is introduced below. The construction process.

[0125] The Transformer model is trained through self-supervised learning to achieve feature alignment of self-supervised learning. Its overall loss function is divided into two parts: modality contrast loss and modality reconstruction loss. Modality contrast loss is used to constrain the consistency of interactions between modalities in the semantic space, while modality reconstruction loss is used to constrain the model to ensure information integrity and prevent feature information loss when reconstructing the original modality features.

[0126] Calculate the feature similarity between modalities for multimodal features, and construct modal contrast loss based on feature similarity. It consists of three parts, corresponding to the interaction between three modalities, including image-text modality contrast loss , image-temporal modality contrast loss Time series and text modality contrast loss .

[0127] Image-text modality contrast loss The construction formula is:

[0128] (10)

[0129] Image-temporal modality contrast loss The construction formula is:

[0130] (11)

[0131] Temporal-Text Modal Contrast Loss The construction formula is as follows:

[0132] (12)

[0133] In the above formulas (10), (11), and (12), N represents the batch size of training samples, j represents the j-th training sample, represents the temperature coefficient, sim represents the function for calculating the feature similarity between modalities, represents the temporal feature of the i-th training sample, represents the text feature of the i-th training sample, represents the image feature of the i-th training sample, and exp represents the exponential function.

[0134] For the modal reconstruction loss, the final fused feature is reconstructed through the decoder of the cross-modal Transformer model to obtain the feature representations of each modality, and then the modal reconstruction loss is constructed based on the feature representations of each modality and the multi-modal features , and the formula is as follows:

[0135] (13)

[0136] In the above formula (13), represents the final fused feature, , , respectively represent the functions for performing reconstruction processing by the decoders of the Transformer model for the image modality, text modality, and temporal modality, represents the image feature, represents the text feature, represents the temporal feature.

[0137] Finally, the modal contrast loss and the modal reconstruction loss are weighted by the preset contrast weight and reconstruction weight respectively, and the weighted results are summed to obtain the overall loss function , which is expressed as the following formula:

[0138] (14)

[0139] In the above formula (14), is the contrast weight, is the reconstruction weight, represents the modal contrast loss, and its specific value is the image-text modal contrast loss 、image-temporal modal contrast loss Temporal-Text Modal Contrastive Loss The sum of the three, is the modal reconstruction loss.

[0140] During the training process, through the overall loss function perform backpropagation in the cross-modal Transformer model to update the model's parameters. When reaching the maximum number of training epochs or when the overall loss function converges, stop the training.

[0141] In the embodiments of the present invention, the cross-modal Transformer model adopts a self-supervised learning method and performs model training by constructing modal contrastive loss and modal reconstruction loss, which can enable the cross-modal Transformer model to have the ability to align different modal features when performing multi-modal feature fusion, improve the generalization ability of feature fusion, avoid the loss of feature information, and enhance the semantic consistency and robustness of the fused features.

[0142] The following introduces the process of constructing a dynamic knowledge graph for cabbage planting management based on the final fused features As shown, the graph construction process of the dynamic knowledge graph includes processes such as data preprocessing, entity-relationship extraction, knowledge graph generation and update, knowledge fusion, and knowledge reasoning. Figure 2 Specifically, first, align the multi-modal data of cabbage planting management with the final fused features to obtain structured multi-modal input data.

[0143] Here, data preprocessing needs to be performed before data alignment, which can be achieved through data processing algorithms. By inputting the final fused features

[0144] and multi-modal data (i.e., image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management), use data processing algorithms to perform data cleaning and filter out noise information, such as removing low-quality images or abnormal environmental data. In addition, data formatting also needs to be performed to convert image data, labeled text data, and environmental time-series data into a standard format suitable for model input. After that, perform data alignment. Data alignment is also the alignment of modal features. Here, based on the timestamps in the final fused features

[0145] align the image data, labeled text data, and environmental time-series data respectively. For the image data, labeled text data, and environmental time-series data corresponding to each timestamp, a final fused feature needs to be aligned to ensure the synchronization between multi-modal data. The data processing algorithm finally outputs structured multi-modal input data, providing a basis for entity-relationship extraction.

[0146] Next, entity-relationship extraction is performed. The core of constructing a knowledge graph lies in extracting entities (such as "downy mildew") and relationships (such as "causes") from multi-modal data. Specifically, entity extraction is performed on multi-modal features to obtain corresponding entity nodes and form an entity set. Multi-modal features are also image features , text features , temporal features , and the RoBERTa-Transformer-CRF model is used in the entity extraction process and combined with the final fusion features , to identify entity names and their attributes, such as "humidity = 85%" or "downy mildew risk", etc., and finally output an entity set, for example: {too high humidity, downy mildew, leaf disease}.

[0147] At the same time, relationship extraction is also performed on the multi-modal input data here to obtain the relationships between entity nodes and form a relationship set. The multi-modal input data are the features after data alignment. When performing relationship extraction, a relationship classification model of cross-modal Transformer is called to capture the relationships between entities for the multi-modal input data, through the cross-modal attention mechanism and the final fusion features , to identify explicit and implicit relationships in the relationships, and finally obtain the corresponding relationship set, for example: {too high humidity, causes, downy mildew}.

[0148] Finally, knowledge graph generation and update are carried out. First, triples are constructed based on the entity set and the relationship set to obtain a primary knowledge graph. Here, the Resource Description Framework RDF is used to construct triples (Entity1, Relation, Entity2), for example: (too high humidity, causes, downy mildew), and the generated primary knowledge graph is stored using the graph database Neo4j

[0149] Then, dynamic update is performed on the primary knowledge graph to obtain a dynamic knowledge graph for cabbage planting management. Through the input real-time environmental data, it is necessary to update the entity nodes and relationships in the primary knowledge graph, including the addition, deletion, and modification of relationships, etc

[0150] In the embodiments of the present invention, multi-modal data and final fusion features are uniformly embedded into the dynamic knowledge graph framework, increasing the knowledge richness of the knowledge graph and improving the accuracy of cabbage planting knowledge expression, and solving the problem of the single expression method of the dynamic knowledge expression system in the prior art

[0151] The following introduces the dynamic update process of the primary knowledge graph

[0152] The primary knowledge graph needs to be updated in real time and should be able to accept the input of real-time environmental data. When performing real-time updates, obtain the real-time environmental data during the cabbage planting process, and aggregate the real-time environmental data through a time window to obtain window environmental features.

[0153] The real-time environmental data specifically includes the environmental data of cabbage planting and the corresponding structured data. The data types of the real-time environmental data are: meteorological data: temperature, humidity, light intensity, etc.; soil data: pH value, water content, etc.; crop status data: leaf color, pest and disease characteristics, etc.; sensor data: real-time environmental change data collected in the cabbage planting area through Internet of Things devices.

[0154] The structured data is to transform all the real-time environmental data into standardized feature vectors through normalization preprocessing, denoted as , where represents a certain feature item in the real-time environmental data (such as humidity or light intensity), is the time point, is the standardized feature vector after preprocessing, that is, the structured data.

[0155] Then perform the extraction of the temporal features of the real-time environmental data, and aggregate the real-time environmental data through the time window to obtain window environmental features to reflect the dynamic changes of the real-time environmental data. The window environmental features , and the calculation formula is as follows:

[0156] (15)

[0157] In the above formula (15), k is the size of the time window, represents the structured data corresponding to the real-time environmental data at time t, and the explanations of other parameters can be referred to.

[0158] Next, it is necessary to determine whether to update the primary knowledge graph. The judgment criterion is to determine the feature change rate of the window environmental features between time windows. The calculation formula of the feature change rate is as follows:

[0159] (16)

[0160] In the above formula (16), k is the size of the time window, represents the structured data corresponding to the real-time environmental data at time t, represents the structured data corresponding to the real-time environmental data at time t - 1, and N represents the number of time windows.

[0161] When the feature change rate is greater than the feature change threshold When the feature change rate triggers an update operation on the primary knowledge graph. When the feature change rate is less than or equal to the feature change threshold it indicates that there is no need to perform an update on the knowledge graph. The following describes the update operation on the primary knowledge graph.

[0162] The update operation of the primary knowledge graph mainly updates the relationships in the relationship set, the corresponding relationship weights, and the attribute values. First, based on the rule base and the graph reasoning results, the relationships to be updated are determined. Specifically, it includes removing relationships with unchanged feature values for a long time or exceeding the association range, modifying the relationship weights or attribute values according to the latest environmental data, and introducing new relationships related to the current environment.

[0163] Specifically, first, according to the window environment features a relationship change matrix before and after the update is constructed Based on the window environment features, it can be determined whether new relationships need to be added, outdated relationships need to be moved, or relationships remain unchanged, and thus mapped to different feature matrices as the relationship change matrix which is expressed as:

[0164] (17)

[0165] In the above formula (17), i represents the relationship before the update in the relationship set of the primary knowledge graph, that is, the relationship that still existed in the previous time window t - 1, and j represents the relationship after the update in the relationship set of the primary knowledge graph, that is, the new relationship calculated in the current time window t.

[0166] Finally, within the time window, the relationship change matrix is merged with the original relationship set R in the primary knowledge graph to obtain a new relationship set which is expressed as:

[0167] (18)

[0168] In the embodiment of the present invention, by calculating the feature change rate corresponding to the real-time environmental data through the time window, it can accurately determine whether to update the relationships in the primary knowledge graph. During the update, the relationship change matrix is determined through the calculated window environment features to achieve the update of the relationships, and finally the construction of the dynamic knowledge graph is completed. In this way, the dynamic adjustment of the knowledge graph is realized, ensuring the accuracy and real-time nature of the knowledge graph, and making up for the static and lagging problems of the traditional agricultural planting knowledge base.

[0169] In some embodiments, after constructing the relationship change matrix before and after the update, a conflict detection process needs to be performed. This process occurs before the relationship update and is mainly used to detect whether the newly added knowledge incorporated through real-time environmental data conflicts with the existing primary knowledge graph, and then perform conflict correction to ensure the consistency of the knowledge graph.

[0170] In the embodiments of the present invention, determining whether there is a conflict is based on calculating the update conflict score C as the judgment criterion. Specifically, first, determine the relationship set before the update and the relationship set after the update of feature similarity . Here, the cosine similarity between the two features can be calculated. Then, respectively determine the exclusive conflict score , attribute conflict score , and reasoning conflict score existing between the relationship set before the update and the relationship set after the update.

[0171] Here, the rules for detecting conflicts are determined as three conflict types: exclusive conflict, attribute conflict, and reasoning conflict. Among them, exclusive conflict means that two or more relationships are logically exclusive (such as "high humidity" and "low humidity"). The value of the exclusive conflict score is 1 when there is an exclusive relationship between two relationships, and 0 otherwise. Attribute conflict means that the attribute value of the same entity exceeds the rule-defined range. The attribute conflict score is calculated as the deviation degree and the attribute value is normalized to [0, 1] when the attribute value exceeds the preset range, otherwise there is no need to calculate. Reasoning conflict means that the conclusion drawn through graph reasoning conflicts with the existing knowledge. The value of the reasoning conflict score is 1 when the reasoning result conflicts with the existing knowledge, and 0 otherwise.

[0172] Next, combine these three parts to obtain the conflict degree score conf. Here, add the exclusive conflict score , attribute conflict score , and reasoning conflict score to obtain the conflict degree score conf, which is expressed as:

[0173] (19)

[0174] Then, take the difference between the feature similarity and the conflict degree score as the update conflict score , which is expressed as:

[0175] (20)

[0176] In the above formula (20), represents the set of relationships before the update, and represents the set of relationships after the update.

[0177] After determining the update conflict score for the current relationship update, it can be used as a judgment criterion to determine whether the updated relationship conflicts this time. When the update conflict score is greater than the preset conflict threshold it is determined that there is a conflict between the set of relationships before the update and the set of relationships after the update; otherwise, there is no conflict.

[0178] When it is determined that there is a conflict, adjust the relationship weights in the set of relationships according to the conflict degree score to correct the conflict existing between the set of relationships before the update and the set of relationships after the update.

[0179] Here, the present invention implements a conflict correction strategy. By adjusting the relationship weights in the set of relationships, specifically, adjusting the set of relationships before the update and the set of relationships after the update to balance the influence of feature similarity and conflict degree score. The adjustment formula is expressed as:

[0180] (21)

[0181] In the above formula (21), represents the adjusted relationship weight, represents the relationship weight before adjustment, is the weight coefficient, represents the set of relationships before the update and the set of relationships after the update of the conflict degree score.

[0182] In some embodiments, if the conflict still cannot be corrected by adjusting the relationship weights, then remove the conflict relationship from the set of relationships in the current initial knowledge graph and add a replacement relationship, which is expressed as:

[0183] (22)

[0184] In the above formula (22), represents the updated relationship set (new knowledge graph relationship set) after replacing the conflict relationship; represents the original relationship set, that is, the knowledge graph relationship set before conflict detection; represents the detected conflict relationship, that is, the relationship that needs to be removed, such as mutually exclusive, attribute, and inference conflict relationships; represents the replacement relationship, that is, the new relationship consistent with the current environment data of the knowledge graph, specifically generated based on the rule library or inference result.

[0185] In the embodiments of the present invention, before updating the relationships of the initial knowledge graph, a conflict detection method and a conflict correction method are proposed, and finally the construction of the dynamic knowledge graph is completed, which can ensure the knowledge consistency between the newly added knowledge and the current knowledge graph. Thus, a dynamic update method for cabbage planting knowledge is realized, which makes up for the problems of static and lagging traditional agricultural knowledge bases.

[0186] In addition, during the construction of the dynamic knowledge graph, knowledge fusion and knowledge reasoning are also involved. The goal of knowledge fusion is to unify the graph content through knowledge from different sources and improve the integrity and consistency of the graph, while knowledge reasoning is to complete data-driven knowledge reasoning through the dynamic knowledge graph, which can provide support for agricultural decision-making in cabbage planting. The following is a specific description.

[0187] For knowledge fusion, first, it is necessary to determine the feature similarity between any two entity nodes in the primary knowledge graph, denoted as , and the calculation formula is as follows: [[ID=1st]]

[0188] (23)

[0189] In the above formula (23), respectively represent any two entity nodes in the dynamic knowledge graph, then respectively represent any two entity nodes 's final fusion features.

[0190] When the feature similarity is greater than the preset similarity threshold, the two entity nodes are merged to achieve knowledge fusion, and when the feature similarity is less than or equal to the preset similarity threshold, no knowledge fusion is required.

[0191] For knowledge reasoning, the entity nodes and edges in the dynamic knowledge graph are extracted here, and the graph convolutional neural network (GCN) is called to predict the relationships between the entity nodes and edges, obtaining the predicted nodes and predicted relationships, and realizing the prediction of future relationships through the graph convolutional neural network. For example, according to "leaf humidity", it is predicted that "downy mildew" may be caused. The predicted nodes are modal entities, and the examples are "too high humidity" and "downy mildew", and the predicted edges are relationships, and the examples are "caused by" and "related to", and the attribute values are the information updated in time sequence.

[0192] Finally, the predicted nodes and predicted relationships are merged with the entity set and relationship set in the dynamic knowledge graph. Here, the inference knowledge predicted by the graph convolutional neural network is also supplemented into the knowledge graph. In this way, through knowledge fusion and knowledge reasoning, the integrity and consistency of the dynamic knowledge graph are ensured, and it can provide agricultural decision-making support for cabbage planting management.

[0193] In some embodiments, after constructing a dynamic knowledge graph, it is possible to provide decision support for cabbage planting management and offer efficient, precise, and personalized agricultural planting support services to farmers.

[0194] First, for the cabbage planting interaction data to be processed, the dynamic knowledge graph is called to retrieve the cabbage planting interaction data, and the corresponding cabbage planting feedback results are obtained.

[0195] Here, the dynamic knowledge graph can be modeled as a decision-making platform to achieve dynamic interaction with farmers. As Figure 2 shown, this decision-making platform can implement three functions: disease diagnosis, intelligent question answering, and personalized recommendation. When farmers plant cabbage and need decision support, the cabbage planting interaction data to be processed input by farmers can be obtained in real time. At this time, the dynamic knowledge graph is called to retrieve the cabbage planting interaction data, and the corresponding cabbage planting feedback results are obtained.

[0196] For the three functions of disease diagnosis, intelligent question answering, and personalized recommendation, the cabbage planting interaction data to be processed also falls into three situations, and the final cabbage planting feedback results also have three situations, which will be described one by one below.

[0197] When the cabbage planting interaction data to be processed is a cabbage image that needs to be diagnosed for diseases, the VisionTransformer model is called to extract the disease features of the cabbage image, and the corresponding disease nodes are retrieved in the dynamic knowledge graph according to the disease features. Then, knowledge reasoning is performed on the disease nodes through the dynamic knowledge graph to obtain the disease diagnosis result of the cabbage image as the cabbage planting feedback result.

[0198] After farmers upload a cabbage image that needs to be diagnosed for diseases, it means that the disease diagnosis function of the dynamic knowledge graph needs to be called. Here, the cabbage image is first preprocessed by an image recognition model to extract the corresponding image features, and then the image features are input into the dynamic knowledge graph to retrieve the corresponding entity nodes as disease nodes, and reasoning is performed in combination with regional environmental data (such as climate, soil conditions) to output the most likely disease type and its associated features.

[0199] Specifically, first, image features are extracted from the uploaded cabbage images through a multi-task learning model and Vision Transformer as disease characteristics (such as spot shape, color, etc.) and mapped to the image feature nodes in the dynamic knowledge graph. Then, in combination with the labeled text data of cabbage planting management, the RoBERTa model is called to extract context semantic features to obtain disease descriptive features and establish semantic associations with the disease nodes in the knowledge graph. At the same time, according to the environmental time series data of the uploaded cabbage images, time series features are extracted through the TCT model as disease causal change features to represent time nodes and environmental impact factors (such as climate mutations, seasonal changes).

[0200] Next, disease feature matching is performed. First is image matching. Through the disease features generated by the above-mentioned Vision Transformer, in the dynamic knowledge graph, the disease nodes closest to the disease features are found using vector retrieval. Then, through the disease descriptive features generated by RoBERTa, in the dynamic knowledge graph, the corresponding disease description nodes and prevention and control measure nodes are matched through semantic similarity calculation in the embedding space. Next is multi-modal feature joint matching. Using the cross-modal Transformer model, the disease features, disease causal change features, and disease descriptive features are jointly calculated to generate the final fusion feature, and in the dynamic knowledge graph, through the feature vector alignment of the multi-modal disease description nodes, comprehensive matching of disease types is achieved.

[0201] During the above matching process, a path search algorithm based on the knowledge graph can be called to perform reasoning in the dynamic knowledge graph in combination with the Bayesian network and the multi-modal fusion model. By traversing the disease nodes and their associated nodes through the path and combining the environmental causal relationship, the most likely disease type and its transmission path are output. Finally, the diagnosis result is returned, including the disease name, the cause of the disease, the transmission path, specific prevention and control measures, the system recommends suitable pesticides and usage methods, and at the same time, feedback results for cabbage planting such as medication time suggestions are provided according to climate prediction data.

[0202] When the cabbage planting interaction data to be processed is a prompt for consulting cabbage planting, the dynamic knowledge graph is called to perform knowledge reasoning on the prompt to obtain cabbage planting data as the cabbage planting feedback result.

[0203] Here, when a farmer inputs planting-related questions through text as prompts for consulting cabbage planting, such as "What crops are suitable for this piece of land?" or "When should fertilization be carried out?", it indicates that the dynamic knowledge graph needs to be called to execute the intelligent question-answering function. First, through the natural language processing module, the semantic parsing and intention recognition of the cabbage planting prompts are completed. Then, according to the question intention, the corresponding entity nodes (such as crop types, fertilization rules) are located from the dynamic knowledge graph, and knowledge reasoning is carried out in combination with the real-time data input of multimodal perception (such as soil nutrient data, regional planting records). Finally, accurate answers are generated as cabbage planting data, such as the types of crops suitable for planting, the types and dosages of fertilizers, and relevant scientific bases (such as historical data, experimental results) are attached, etc., as the cabbage planting feedback results. In addition, the farmer can append questions, such as "Why is this fertilizer recommended?", etc., and at this time, the dynamic knowledge graph is continuously called to further execute knowledge reasoning to obtain an explanatory answer.

[0204] When the cabbage planting interaction data to be processed represents the execution of personalized recommendations for cabbage planting, the dynamic knowledge graph is called to perform knowledge reasoning on the historical cabbage planting data to obtain the recommended content for cabbage planting as the cabbage planting feedback result.

[0205] Here, when a farmer inputs the cabbage planting interaction data to be processed and needs to obtain personalized recommendations for cabbage planting, at this time, a farmer portrait will be dynamically constructed based on the historical cabbage planting data such as the farmer's historical planting behavior, regional environmental characteristics, and real-time sensor data, such as the types and time of planted crops, environmental characteristics such as regional soil and climate, the farmer's operation habits, commonly used planting equipment, and concerned planting problems, etc. Then, the dynamic knowledge graph is called to perform knowledge reasoning on the farmer portrait to generate the corresponding recommended content for cabbage planting, such as fertilization, irrigation, or disease prevention and control plans for specific crops, recommended best planting times according to meteorological forecasts, and providing relevant technical guidance or learning resources for the user's common problems, etc., as the cabbage planting feedback results.

[0206] In the embodiment of the present invention, through the knowledge reasoning ability of the dynamic knowledge graph for cabbage planting management, three interaction functions of disease diagnosis, intelligent question-answering, and personalized recommendation are completed, realizing the dynamic interaction of cabbage planting interaction data, ensuring the timeliness and scientific nature of planting management knowledge, breaking the limitations of the traditional single text or table output of planting guidance, and effectively providing agricultural decision-making for cabbage planting.

[0207] The following describes the cabbage knowledge dynamic expression and interaction system based on multimodal fusion provided by the present invention. The cabbage knowledge dynamic expression and interaction system based on multimodal fusion described below can be mutually corresponding and referred to the multimodal fusion-based cabbage knowledge dynamic expression and interaction method described above. As Figure 3 shown, the cabbage knowledge dynamic expression and interaction system based on multimodal fusion includes:

[0208] A feature extraction module 301 is configured to obtain multimodal data for cabbage planting management. The multimodal data includes image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management. A feature fusion module 302 is configured to extract features from the multimodal data to obtain multimodal features, and call a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain final fusion features. A graph construction module 303 is configured to construct a dynamic knowledge graph for cabbage planting management based on the final fusion features. A dynamic interaction module 304 is configured to, for the to-be-processed cabbage planting interaction data, call the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results.

[0209] It should be noted that the beneficial effects of the dynamic expression and interaction system of cabbage knowledge based on multimodal fusion here correspond to those of the dynamic expression and interaction method of cabbage knowledge based on multimodal fusion in the above text. Therefore, the beneficial effects of the dynamic expression and interaction system of cabbage knowledge based on multimodal fusion are not elaborated here.

[0210] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the dynamic expression and interaction method of cabbage knowledge based on multimodal fusion. The method includes: obtaining multimodal data for cabbage planting management, where the multimodal data includes image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management; extracting features from the multimodal data to obtain multimodal features, and calling a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain final fusion features; constructing a dynamic knowledge graph for cabbage planting management based on the final fusion features; for the to-be-processed cabbage planting interaction data, calling the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results.

[0211] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0212] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the above-mentioned various methods. The method includes: obtaining multimodal data for cabbage planting management, where the multimodal data includes image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management; extracting features from the multimodal data to obtain multimodal features, and invoking a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain final fusion features; constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features; and for the cabbage planting interaction data to be processed, invoking the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results.

[0213] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion provided by the above-mentioned various methods. The method includes: obtaining multimodal data for cabbage planting management, where the multimodal data includes image data of cabbage growth, environmental time-series data, and labeled text data of cabbage planting management; extracting features from the multimodal data to obtain multimodal features, and invoking a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain final fusion features; constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features; and for the cabbage planting interaction data to be processed, invoking the dynamic knowledge graph to retrieve the cabbage planting interaction data to obtain corresponding cabbage planting feedback results.

[0214] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0215] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for dynamic expression and interaction of cabbage knowledge based on multimodal fusion, characterized in that, Including: Obtain multi-modal data of cabbage planting management, where the multi-modal data includes image data of cabbage growth, environmental time-series data, and annotated text data of cabbage planting management; Extract features from the multi-modal data to obtain multi-modal features, and call a cross-modal Transformer model to align the multi-modal features to obtain final fused features; Construct a dynamic knowledge graph of cabbage planting management based on the final fused features; For the cabbage planting interaction data to be processed, call the dynamic knowledge graph to retrieve the cabbage planting interaction data and obtain the corresponding cabbage planting feedback results; The constructing a dynamic knowledge graph of cabbage planting management based on the final fused features includes: Align the multi-modal data of cabbage planting management with the final fused features to obtain structured multi-modal input data; Extract entities from the multi-modal features to obtain corresponding entity nodes and form an entity set; Extract relationships from the multi-modal input data to obtain the relationships between the entity nodes and form a relationship set; Construct triples based on the entity set and the relationship set to obtain a primary knowledge graph; Perform dynamic updates on the primary knowledge graph to obtain a dynamic knowledge graph of cabbage planting management The dynamic update process of the primary knowledge graph includes: Obtain real-time environmental data during the cabbage planting process, and aggregate the real-time environmental data through a time window to obtain window environmental features; Determine the feature change rate of the window environmental features between time windows. When the feature change rate is greater than the feature change threshold, trigger an update operation on the primary knowledge graph; Among them, the update operation on the primary knowledge graph includes: Construct a relationship change matrix before and after the update according to the window environmental features; Within the time window, merge the relationship change matrix with the original relationship set in the primary knowledge graph to obtain a new relationship set.

2. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, wherein, The multi-modal features include image features of the image data, time-series features of the environmental time-series data, and text features of the annotated text data. The extracting features from the multi-modal data to obtain multi-modal features includes: Call a Vision Transformer model to perform a disease classification task and a leaf health status estimation task on the image data to obtain image features of the image data. Among them, the disease classification task is used to classify cabbage diseases for the image data, and the leaf health status estimation task is used to estimate the health status of cabbage leaves for the image data. The Vision Transformer model is trained through a disease classification loss and a status estimation loss; Call a time-series Transformer model based on causal attention to extract features from the environmental time-series data to obtain time-series features; Call a Tokenizer model to encode the annotated text data into a Token ID sequence; Obtain prompt words for specific domain knowledge, and input the prompt words and the Token ID sequence into the RoBERTa model for context semantic feature extraction to obtain the text features of the labeled text data.

3. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, wherein The calling of the cross-modal Transformer model to perform feature alignment on the multi-modal features to obtain the final fusion features includes: Performing inter-modal interaction calculations on the multi-modal features through the cross-modal attention mechanism of the cross-modal Transformer model to obtain image-temporal interaction features, temporal-text interaction features, and image-text interaction features respectively; Fusing the image-temporal interaction features, the temporal-text interaction features, and the image-text interaction features to obtain a joint feature representation; Performing feature dimension transformation processing on the joint feature representation to obtain the final fusion features.

4. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, characterized in that The cross-modal Transformer model is trained by constructing an overall loss function. The construction process of the overall loss function includes: Calculating the feature similarity between modalities for the multi-modal features, and constructing a modality contrast loss based on the feature similarity. The modality contrast loss includes an image-text modality contrast loss, an image-temporal modality contrast loss, and a temporal-text modality contrast loss; Reconstructing the final fusion features through the decoder of the cross-modal Transformer model to obtain the feature representations of each modality; Constructing a modality reconstruction loss based on the feature representations of each modality and the multi-modal features; Weighting the modality contrast loss and the modality reconstruction loss respectively through preset contrast weights and reconstruction weights, and summing the weighted results to obtain the overall loss function.

5. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, wherein After constructing the relationship change matrix before and after the update, the method further includes: Determining the feature similarity between the relationship set before the update and the relationship set after the update; Respectively determining the mutual exclusion conflict score, attribute conflict score, and inference conflict score existing between the relationship set before the update and the relationship set after the update; Adding the mutual exclusion conflict score, the attribute conflict score, and the inference conflict score to obtain a conflict degree score; Taking the difference between the feature similarity and the conflict degree score as the update conflict score. When the update conflict score is greater than a preset conflict threshold, it is determined that there is a conflict between the relationship set before the update and the relationship set after the update; Adjusting the relationship weights in the relationship set according to the conflict degree score to correct the conflict existing between the relationship set before the update and the relationship set after the update.

6. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, characterized in that The method further includes: Determining the feature similarity between any two entity nodes in the primary knowledge graph. When the feature similarity is greater than a preset similarity threshold, the two entity nodes are merged; Extracting entity nodes and edges from the dynamic knowledge graph, and calling a graph convolutional neural network to perform relationship prediction on the entity nodes and the edges to obtain predicted nodes and predicted relationships; Merging the predicted nodes and predicted relationships with the entity set and relationship set in the dynamic knowledge graph.

7. The method for dynamically expressing and interacting with cabbage knowledge based on multimodal fusion according to claim 1, characterized in that, For the to-be-processed interactive data of cabbage planting, call the dynamic knowledge graph to retrieve the to-be-processed interactive data of cabbage planting, and obtain the corresponding cabbage planting feedback result, including: When the to-be-processed interactive data of cabbage planting is a cabbage image that needs disease diagnosis, call the VisionTransformer model to extract the disease characteristics of the cabbage image, and retrieve the corresponding disease node in the dynamic knowledge graph according to the disease characteristics; Perform knowledge reasoning on the disease node through the dynamic knowledge graph to obtain the disease diagnosis result of the cabbage image as the cabbage planting feedback result; When the to-be-processed interactive data of cabbage planting is a prompt word for consulting cabbage planting, call the dynamic knowledge graph to perform knowledge reasoning on the prompt word to obtain cabbage planting data as the cabbage planting feedback result; When the to-be-processed interactive data of cabbage planting represents the execution of personalized recommendation for cabbage planting, call the dynamic knowledge graph to perform knowledge reasoning on the historical cabbage planting data to obtain the recommended content for cabbage planting as the cabbage planting feedback result.

8. A multi-modal fusion-based dynamic knowledge expression and interaction system for cabbages, characterized in that, Including: A feature extraction module for obtaining multimodal data of cabbage planting management, where the multimodal data includes image data of cabbage growth, environmental time-series data, and annotated text data of cabbage planting management; A feature fusion module for extracting features from the multimodal data to obtain multimodal features, and calling a cross-modal Transformer model to perform feature alignment on the multimodal features to obtain the final fusion features; A graph construction module for constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features; A dynamic interaction module for, for the to-be-processed interactive data of cabbage planting, calling the dynamic knowledge graph to retrieve the to-be-processed interactive data of cabbage planting to obtain the corresponding cabbage planting feedback result; The constructing a dynamic knowledge graph of cabbage planting management based on the final fusion features includes: Align the multimodal data of cabbage planting management with the final fusion features to obtain structured multimodal input data; Perform entity extraction on the multimodal features to obtain the corresponding entity nodes and form an entity set; Perform relation extraction on the multimodal input data to obtain the relations between the entity nodes and form a relation set; Construct triples based on the entity set and the relation set to obtain a primary knowledge graph; Perform dynamic update on the primary knowledge graph to obtain a dynamic knowledge graph of cabbage planting management; The dynamic update process of the primary knowledge graph includes: Obtain real-time environmental data during the cabbage planting process, and aggregate the real-time environmental data through a time window to obtain window environmental features; Determine the feature change rate of the window environmental features between time windows, and when the feature change rate is greater than the feature change threshold, trigger an update operation on the primary knowledge graph; Among them, the update operation on the primary knowledge graph includes: Construct a relationship change matrix before and after the update according to the window environmental features; Within the time window, merge the relationship change matrix with the original relationship set in the primary knowledge graph to obtain a new relationship set.

Citation Information

Patent Citations

  • Temperature prediction method and system for multi-mode AIGC cold-chain logistics monitoring platform

    CN119477139A

  • Multi-modal perception-driven vegetable knowledge graph construction method and multi-modal perception-driven vegetable knowledge graph construction device

    CN119599104A

  • Knowledge graph dynamic updating method and device based on crop growth process

    CN119646102A