Vegetable multi-modal knowledge graph construction method and device and storage medium

By collecting and processing multimodal data of vegetable knowledge, a multimodal knowledge graph of vegetable is constructed, which solves the problem of lack of multimodal knowledge graph in the existing technology, and realizes accurate vegetable knowledge Q&A and personalized knowledge services.

CN120046707APending Publication Date: 2025-05-27BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411967309.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The lack of multimodal knowledge graphs for vegetables in the prior art, resulting in the inability to conduct accurate vegetable knowledge questions and answers.

Method used

By collecting multimodal data on vegetable knowledge, including text, file and picture modal data, multimodal semantic fusion features are obtained, and entities are identified using entity extraction models, entity relationship extraction models are extracted to obtain entity relationships, and finally a vegetable multimodal knowledge graph is constructed.

Benefits of technology

It has realized the construction of a more accurate and comprehensive vegetable multimodal knowledge graph, can provide personalized knowledge services, support real-time data combined with reasoning, conduct accurate production knowledge questions and answers, guide vegetable production, and provide key technical means for unmanned operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046707A_ABST
    Figure CN120046707A_ABST
Patent Text Reader

Abstract

The invention provides a vegetable multi-modal knowledge graph construction method and device and a storage medium. The vegetable multi-modal knowledge graph construction method comprises the following steps: collecting vegetable knowledge multi-modal data; the vegetable knowledge multi-modal data comprises data of a text modal, data of a file modal and data of a picture modal in a vegetable field; obtaining a multi-modal semantic fusion feature based on the vegetable knowledge multi-modal data; based on the multi-modal semantic fusion features, utilizing an entity extraction model to identify entities in the vegetable knowledge multi-modal data, and utilizing an entity relationship extraction model to obtain relationships among the entities; and according to the entities and the relationship between the entities, obtaining the vegetable multi-modal knowledge graph. Semantic fusion is performed on the basis of multi-modal data in the field of vegetable knowledge, so that entity recognition and relation extraction are performed, and a more accurate and comprehensive vegetable multi-modal knowledge graph is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and in particular to a method, device and storage medium for constructing a vegetable multimodal knowledge graph. Background Art

[0002] Since the introduction of knowledge graphs, a wave of research has been launched in the industrial and academic fields. Benefiting from the structured semantic expression of knowledge graphs and convenient entity and relationship structure features, they are increasingly widely used in vertical fields such as medicine, law, agriculture, and industrial production. However, due to the limitations of the modality, although traditional text semantic knowledge graphs have made great achievements in data collation, knowledge services, and semantic reasoning, their service capabilities are slightly insufficient for complex agricultural scenarios. They are often unable to provide more accurate services because the semantic expression lacks semantic evidence from images or other patterns. Therefore, there is an urgent need to expand the modality of knowledge graphs to more fully express semantics.

[0003] Multimodal knowledge graphs have been a hot topic in the field of artificial intelligence in recent years. Thanks to their entity-relationship graph database structure and rich semantic expressions in multiple modalities of "image, vision, audio, and text", they have unique advantages in the agricultural field because the data presents multimodal, multi-scale, multi-source heterogeneous and cross-media characteristics. In terms of diseases, pests, growth monitoring, fertilizer and water irrigation analysis, agricultural technical knowledge services, etc., in order to accurately express semantics, it is usually necessary to apply more than one modality. The use of multimodal knowledge graphs is one of the important means to comprehensively express vegetable production knowledge collection, technical services, and knowledge-driven agricultural operations.

[0004] The current multimodal knowledge graph construction technology can provide analysis methods for specific tasks to a certain extent. However, in agricultural scenarios, due to the various characteristics of agriculture, including complex scenarios, large regional differences, multi-source and heterogeneous data, and technical service characteristics, there is a lack of multimodal knowledge graphs that can provide personalized knowledge services. Summary of the invention

[0005] The present invention provides a method, device and storage medium for constructing a multimodal knowledge graph of vegetables, so as to solve the technical problem that there is no multimodal knowledge graph for vegetables in the prior art, and thus accurate vegetable knowledge questions and answers cannot be performed.

[0006] In a first aspect, the present invention provides a method for constructing a vegetable multimodal knowledge graph, comprising the following steps: Collecting multimodal data of vegetable knowledge; the multimodal data of vegetable knowledge includes text modal data, file modal data and picture modal data in the field of vegetables; Acquiring multimodal semantic fusion features based on the vegetable knowledge multimodal data; Based on the multi-modal semantic fusion features, use an entity extraction model to identify entities in the multi-modal data of vegetable knowledge, and use an entity relationship extraction model to obtain the relationships between the entities; Obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0007] In some embodiments, the obtaining of the multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge includes: Based on the data of the text modality and the data of the document modality in the vegetable field, obtain vegetable knowledge text data, and use the data of the picture modality in the vegetable field as vegetable knowledge picture data; Use the pre-trained language model Roberta to process the vegetable knowledge text data to obtain a text semantic segmentation result, use the semantic segmentation model Yolov8-seg to process the vegetable knowledge picture data to obtain a picture semantic segmentation result, and use a structure pre-trained model to obtain the text semantic structure corresponding to the vegetable knowledge text data and the image semantic structure corresponding to the vegetable knowledge picture data to obtain structure information; Map the text semantic segmentation result, the picture semantic segmentation result, and the structure information into the same semantic subspace to obtain multi-modal semantic fusion features.

[0008] In some embodiments, the obtaining of the vegetable knowledge text data based on the data of the text modality and the data of the document modality in the vegetable field includes: Use optical character recognition (OCR) technology to obtain the text in the data of the document modality; Use the data of the text modality and the text in the data of the document modality as vegetable knowledge text data.

[0009] In some embodiments, the knowledge stored in the schema layer of the vegetable multi-modal knowledge graph is divided into vegetable variety schema, vegetable cultivation schema, pest and disease schema, and agricultural machinery schema.

[0010] In some embodiments, the ontology library for managing the schema layer includes vegetable operation knowledge classification, vegetable operation object attributes, and vegetable data attributes; Among them, the vegetable operation knowledge classification includes vegetable production personnel, vegetable production environment, vegetable operation equipment, vegetable varieties, and vegetable pests and diseases; the vegetable production personnel include personnel information and personnel behavior.

[0011] In some embodiments, the obtaining of the relationships between the entities by using an entity relationship extraction model includes: Use the entity relationship extraction model CasRel to annotate the position information of the head entity and the tail entity for the entities; Classify the relationships between the entities according to the annotation results to obtain the relationships between the entities.

[0012] In a second aspect, the present invention provides a device for constructing a vegetable multi-modal knowledge graph, including the following modules: A collection module, configured to collect multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in document modality, and data in picture modality in the field of vegetables; A first acquisition module, configured to obtain multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge; A second acquisition module, configured to identify entities in the multi-modal data of vegetable knowledge by using an entity extraction model based on the multi-modal semantic fusion features, and obtain the relationships between the entities by using an entity relationship extraction model; A third acquisition module, configured to obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the method for constructing a vegetable multi-modal knowledge graph as described in any one of the above is implemented.

[0014] In a fourth aspect, a non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing a vegetable multi-modal knowledge graph as described in any one of the above is implemented.

[0015] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for constructing a vegetable multi-modal knowledge graph as described in any one of the above is implemented.

[0016] The method, device, and storage medium for constructing a vegetable multi-modal knowledge graph provided by the present invention collect multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in document modality, and data in picture modality in the field of vegetables; obtain multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge; identify entities in the multi-modal data of vegetable knowledge by using an entity extraction model based on the multi-modal semantic fusion features, and obtain the relationships between the entities by using an entity relationship extraction model; obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities. Based on semantic fusion of multi-modal data in the field of vegetable knowledge, entity recognition and relationship extraction are performed, so as to construct a more accurate and comprehensive vegetable multi-modal knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0018] Figure 1 It is one of the schematic flowcharts of the method for constructing a multi-modal knowledge graph of vegetables provided by the present invention.

[0019] Figure 2 It is the second schematic flowchart of the method for constructing a multi-modal knowledge graph of vegetables provided by the present invention.

[0020] Figure 3 It is the schematic structural diagram of the multi-modal knowledge graph of vegetables in the example scenario provided by the present invention.

[0021] Figure 4 It is the schematic structural diagram of the sub-graph of vegetable diseases in the example scenario provided by the present invention.

[0022] Figure 5 It is the schematic structural diagram of the sub-graph of vegetable agricultural machinery operations in the example scenario provided by the present invention.

[0023] Figure 6 It is the schematic diagram of the entity relationship extraction process provided by the present invention.

[0024] Figure 7 It is the schematic structural diagram of the multi-modal vegetable production knowledge reasoning and answering in the example scenario provided by the present invention.

[0025] Figure 8 It is the schematic structural diagram of the device for constructing a multi-modal knowledge graph of vegetables provided by the present invention.

[0026] Figure 9 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0027] Currently, for the multi-modal knowledge graph construction technology, generally, good results can be obtained for specific structure matching, knowledge path legends, and relationship extraction, and the application scenarios also require more regular data paradigms and business processes accordingly. However, for agricultural scenarios such as vegetable production, there are lack of good technical means, and there is an urgent need for a method that can integrate multiple modal semantics and support knowledge reasoning, and provide matching data services for different types, scenarios, and basic situations.

[0028] Based on the above technical problems, the present invention proposes a method for constructing a vegetable multi-modal knowledge graph, which collects multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in document modality, and data in picture modality in the field of vegetables; based on the multi-modal data of vegetable knowledge, multi-modal semantic fusion features are obtained; based on the multi-modal semantic fusion features, an entity extraction model is used to identify entities in the multi-modal data of vegetable knowledge, and an entity relationship extraction model is used to obtain the relationships between the entities; according to the entities and the relationships between the entities, a vegetable multi-modal knowledge graph is obtained. By performing semantic fusion on multi-modal data in the field of vegetable knowledge, entity recognition and relationship extraction are carried out, so as to construct a more accurate and comprehensive vegetable multi-modal knowledge graph, thereby enabling real-time data combination reasoning through the vegetable multi-modal knowledge graph, performing accurate production knowledge Q&A, better guiding vegetable production, and providing a key technical means for subsequent unmanned operations.

[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0030] Figure 1 is one of the flow schematic diagrams of the method for constructing a vegetable multi-modal knowledge graph provided by the present invention. As Figure 1 shown, the present invention provides a method for constructing a vegetable multi-modal knowledge graph. The method includes: Step 101, collect multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in document modality, and data in picture modality in the field of vegetables.

[0031] Specifically, the multi-modal data of vegetable knowledge refers to multi-modal data in the field of vegetables, including data in text modality, data in document modality, and data in picture modality, etc. Among them, the data in document modality can be PDF or WORD documents, etc.

[0032] In addition to traditional knowledge collation and manual induction, the data collection method can also use web crawlers to collect vegetable knowledge books, documents, network resources, etc. on the Internet respectively, so as to obtain data in three modalities of text, document, and image.

[0033] Step 102, obtain multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge.

[0034] Specifically, based on the multimodal data of vegetable knowledge, semantic features in the data are extracted, and then data of different modalities are semantically aligned or fused to obtain multimodal semantic fusion features.

[0035] Step 103: Based on the multimodal semantic fusion features, an entity extraction model is used to identify entities in the vegetable knowledge multimodal data, and an entity relationship extraction model is used to obtain the relationships between the entities.

[0036] Specifically, based on the fusion of semantics of data of each modality, an entity extraction model is used to identify fused entities, and an entity relationship extraction model is used to extract entity relationships.

[0037] For example, for the text "Chen planted 12 acres of broccoli in City A", the entity extraction model is used to extract entities, and the entities obtained are "Chen", "City A", and "Broccoli". The entity relationship extraction model is used to obtain the relationship between the entities, and the relationship between the entities "Chen", "City A" and "Broccoli" is obtained as person-planting location-planting variety.

[0038] Step 104: Obtain a vegetable multimodal knowledge graph based on the entities and the relationships between the entities.

[0039] Specifically, after obtaining the entities and entity relationships in the vegetable knowledge multimodal data, the triples of entities, entity relationships and knowledge graph structures are established to form a vegetable multimodal knowledge graph that supports Neo4j structure, and GraphXR can be used for convenient human-computer interactive visualization.

[0040] In an embodiment of the present application, based on graph construction, question entity embedding, graph module structure embedding learning, text representation learning are performed respectively for multimodal knowledge question and answer, especially picture and text question and answer, semantic enhancement is performed through a multi-head attention mechanism, and answer sorting and output are performed, and vegetable production knowledge question and answer are realized with the help of reasoning of the model in the knowledge graph.

[0041] The vegetable multimodal knowledge graph construction method provided in the embodiment of the present application, through the data organization method of multimodal data semantic alignment, enables data of different modalities to be combined to construct a more comprehensive and accurate vegetable multimodal knowledge graph, and can provide personalized knowledge services for different vegetable production scenarios. Through the vegetable multimodal knowledge graph, real-time data combination reasoning is realized, accurate production knowledge Q&A is performed, vegetable production is better guided, and key technical means are provided for subsequent unmanned operations.

[0042] In some embodiments, the acquiring of multimodal semantic fusion features based on the vegetable knowledge multimodal data includes: Obtain vegetable knowledge text data based on the data in the text modality and the data in the file modality in the vegetable field, and use the data in the image modality in the vegetable field as vegetable knowledge image data; Use the pre-trained language model Roberta to process the vegetable knowledge text data to obtain a text semantic segmentation result, use the semantic segmentation model Yolov8-seg to process the vegetable knowledge image data to obtain an image semantic segmentation result, and use a structure pre-trained model to obtain the text semantic structure corresponding to the vegetable knowledge text data and the image semantic structure corresponding to the vegetable knowledge image data to obtain structure information; Map the text semantic segmentation result, the image semantic segmentation result, and the structure information into the same semantic subspace to obtain a multi-modal semantic fusion feature.

[0043] Specifically, after collecting data in different modalities in the vegetable field, the data in the collected image modality is the vegetable knowledge image data, and the text is extracted from the collected data in the file modality and used together with the data in the text modality as the vegetable knowledge text data.

[0044] In some embodiments, the obtaining of the vegetable knowledge text data based on the data in the text modality and the data in the file modality in the vegetable field includes: Use optical character recognition (OCR) technology to obtain the text in the data in the file modality; Use the data in the text modality and the text in the data in the file modality as the vegetable knowledge text data.

[0045] Specifically, after collecting data in three modalities of text, file, and image, the data in the file modality needs to be converted into text. For example, use optical character recognition (OCR) technology and other methods to process the data in the file modality to obtain the text in the file or document, obtain text at the chapter level, or segment it according to paragraphs to form text processing units with relatively unified lengths. Then use the processed text together with the collected data in the text modality (including paragraph and sentence text data, etc.) as the vegetable knowledge text data.

[0046] Figure 2 It is the second flow diagram of the method for constructing a vegetable multi-modal knowledge graph provided by the present invention, as Figure 2As shown in the figure, after obtaining the vegetable knowledge text data and vegetable knowledge picture data, these data are preprocessed in parallel, including processing the vegetable knowledge text data with a pre-trained language model to obtain a text semantic segmentation result, processing the vegetable knowledge picture data with a semantic segmentation model to obtain a picture semantic segmentation result, and using a structure pre-trained model to obtain the text semantic structure corresponding to the vegetable knowledge text data and the image semantic structure corresponding to the vegetable knowledge picture data, thereby obtaining structure information.

[0047] Among them, the pre-trained language model can be a Roberta (Robustly Optimized BERT Pretraining Approach) model, which is used to perform text segmentation and vectorization on the text to form a processable vector sequence, that is, the text semantic segmentation result.

[0048] The semantic segmentation model can be a Yolov8-seg model, which is a semantic segmentation version of the YOLO (You Only Look Once) model and is used to perform semantic segmentation on the image to obtain a picture semantic segmentation result.

[0049] The structure pre-trained model is used to extract the semantic structure (i.e., structure information) in the text or image. For example, it extracts the semantic structure in the vegetable knowledge text data to obtain a text structure, and extracts the semantic structure in the vegetable knowledge picture data to obtain an image structure.

[0050] The text semantic segmentation result, picture semantic segmentation result, and structure information obtained by preprocessing the vegetable knowledge text data and vegetable knowledge picture data are mapped into the same semantic subspace to obtain multi-modal semantic fusion features.

[0051] In the vegetable multi-modal knowledge graph construction method provided by the embodiments of the present application, after data collection and preprocessing and local construction, for different mode ontologies in vegetable production, there are cross-modal related data items. When preprocessing, while performing semantic segmentation on the text and image, the text structure and image structure are also parsed synchronously to more comprehensively express relevant semantic information. By mapping the results obtained from preprocessing into the same semantic subspace to obtain multi-modal semantic fusion features, the dual alignment of semantics and graph structure of common text, documents, and images in the vegetable knowledge field is realized. Through parallel semantic processing, the processing efficiency of multi-modal data is improved structurally, and a good foundation is provided for constructing an accurate and effective vegetable multi-modal knowledge graph.

[0052] In some embodiments, the knowledge stored in the mode layer of the vegetable multi-modal knowledge graph is divided into vegetable variety mode, vegetable cultivation mode, pest and disease mode, and agricultural machinery mode.

[0053] Specifically, in the process of constructing the knowledge graph, the pattern layer structure is first constructed by combining the expert knowledge of vegetable varieties, vegetable cultivation, pests and diseases, and agricultural machinery. That is, the knowledge stored in the pattern layer includes four major patterns: vegetable variety pattern, vegetable cultivation pattern, pest and disease pattern, and agricultural machinery pattern. Based on the pattern layer as the supporting basis, the construction of the vegetable multi-modal knowledge graph is guided.

[0054] In some embodiments, the ontology library for managing the pattern layer includes vegetable operation knowledge classification, vegetable operation object attributes, and vegetable data attributes; Among them, the vegetable operation knowledge classification includes vegetable production personnel, vegetable production environment, vegetable operation equipment, vegetable varieties, and vegetable pests and diseases; the vegetable production personnel include personnel information and personnel behavior.

[0055] Specifically, for the four major patterns of vegetable varieties, vegetable cultivation, pests and diseases, and agricultural machinery and the interaction scenarios of personnel, environment, and agronomy after the completion of the vegetable multi-modal knowledge, especially the process of guiding vegetable intelligent operations through the knowledge graph, three major sections of vegetable operation knowledge classification, vegetable operation object attributes, and data attributes are proposed to construct the ontology library of the vegetable multi-modal knowledge graph.

[0056] Figure 3 It is a schematic structural diagram of the vegetable multi-modal knowledge graph of the example scenario provided by the present invention, as Figure 3 shown. Owl is the ontology. In terms of vegetable knowledge classification, it includes vegetable production personnel (Producer), vegetable production environment (Environment), vegetable operation equipment (EqService), vegetable varieties (Vegetable), and vegetable pests and diseases (DiseasePests), etc. The vegetable production personnel include personnel information (person) and personnel behavior (behavior), etc. The vegetable production environment includes location (location) and planting object (object). The vegetable operation equipment includes operation (operate) and operation type (work). The vegetable varieties include class (variety) and genus (category). The vegetable pests and diseases include onset conditions (condition) and onset characteristics (feature).

[0057] The vegetable operation object attributes include having behavior (hasBehavior), being in a location (islocation), having an impact (effect), doing (todo), having conditions (hasCondition), adapting to (fitOn), etc., and can connect the relationships between locals.

[0058] Vegetable data attributes include status attribute, color attribute, shape attribute, work type attribute, temperature attribute, humidity attribute, illumination attribute, CO 2 concentration (CO 2 attribute), etc.

[0059] The method for constructing a vegetable multi-modal knowledge graph provided by the embodiments of the present application structurally organizes domain expert knowledge according to varieties, pests and diseases, vegetable machinery, and cultivation, providing professional knowledge support for subsequent entity recognition and relationship extraction.

[0060] In some embodiments, in order to improve the efficiency of the vegetable multi-modal knowledge graph, the entity recognition part parses text entities and entities that align with image descriptions, focuses on processing text information, and performs relevant semantic matching on this basis. The entity extraction model used is the RoBerta+wwm+CRF model, which is used for named entity recognition and can identify elements such as content diseases, pests, causes of diseases, prevention and control methods, vegetable varieties, planting regions, phenological periods, operation links, operation machinery, plot environments, planting personnel, and environmental data, providing a basis for subsequent relationship extraction processing.

[0061] Among them, in the RoBerta+wwm+CRF model, the text passes through the initial vector obtained by the RoBerta-wwm pre-trained model, dynamically fuses dictionary information for lexical enhancement, generates a text semantic vector, learns sequence dependencies through a downstream bidirectional long short-term memory (BiLSTM), and finally extracts entities through conditional random field (CRF) decoding. Among them, RoBerta-wwm is the RoBerta model using the whole word masking strategy (Whole Word Masking, wwm).

[0062] Figure 4 It is a schematic structural diagram of the vegetable disease sub-graph of the example scenario provided by the present invention. As Figure 4 shown, in the vegetable disease sub-graph of the vegetable multi-modal knowledge graph, it includes various concepts such as disease symptoms, disease locations, and disease types. Each concept is associated with multiple sub-concepts, and the sub-concepts can correspond to one or more entities. Figure 5 It is a schematic structural diagram of the vegetable agricultural machinery operation sub-graph of the example scenario provided by the present invention. As Figure 5As shown, in the sub-graph of vegetable agricultural machinery operations in the vegetable multi-modal knowledge graph, there are several concepts including equipment, control, and sensors. Each concept is associated with multiple sub-concepts, and each sub-concept corresponds to multiple entities.

[0063] In some embodiments, obtaining the relationships between the entities by using the entity relationship extraction model includes: Using the entity relationship extraction model CasRel to annotate the position information of the head entity and the tail entity for the entities; Classifying the relationships of the entities according to the annotation results to obtain the relationships between the entities.

[0064] Specifically, in terms of relationship extraction, the entity relationship extraction model CasRel is used to annotate the position information of the head entity and the tail entity for the entities.

[0065] In the embodiments of the present application, an improved CasRel is used for entity relationship extraction, that is, the generative discriminative pre-training model ELECTRA with higher recognition and generation efficiency is used for head entity annotation, the attention mechanism is used to analyze the head entity and the tail entity, and the relationship is parsed and classified through weight calculation to obtain the relationships between the entities. Attention mechanism embedding is performed at the head and tail of the entity position to fully learn the position information, and the self-attention mechanism is introduced in the entity recognition process for semantic learning to improve the overall learning efficiency.

[0066] For example, Figure 6 is a schematic diagram of the entity relationship extraction process provided by the present invention. As Figure 6 shown, for the text "Chen San planted 12 mu of broccoli in a certain city", the entities are "Chen San", "a certain city", and "broccoli". The CasRel model is used for entity relationship extraction, the head entity attention mechanism is used to analyze the head entity, and the tail entity attention mechanism is used to analyze the tail entity, and the relationships between the entities "Chen San", "a certain city", and "broccoli" are obtained as person - planting location - planting variety.

[0067] The method for constructing a vegetable multi-modal knowledge graph provided by the embodiments of the present application simplifies the process of text information processing and reduces the computational complexity by fusing the entity and relationship extraction process based on CasRel combined with the attention mechanism in the vegetable entity relationship extraction stage, and ensures the accuracy by introducing the attention mechanism.

[0068] In some embodiments, Figure 7 is a schematic diagram of the structure of multi-modal vegetable production knowledge reasoning and answering in an example scenario provided by the present invention. As Figure 7As shown in the figure, based on the construction of the vegetable multi-modal knowledge graph, for multi-modal knowledge Q&A, especially picture-text Q&A, question entity embedding, graph module structure embedding learning, and text representation learning are respectively carried out. Semantic enhancement is performed through the multi-head attention mechanism, and answer ranking is output. The production knowledge Q&A of vegetables is realized by means of the reasoning of the model in the knowledge graph.

[0069] The vegetable multi-modal knowledge graph construction method provided by the embodiments of the present application is oriented to the vegetable production scenario, including the structured introduction of domain expert knowledge, ensuring the professionalism of vegetable knowledge services. Data processing includes not only fixed production professional knowledge, but also sensor data and agricultural machinery working condition data, integrating the knowledge of the three sectors of "agricultural machinery - agronomy - information", and with the help of the graph-structured processing and analysis ability of the knowledge graph, it can perform a processing mode of real-time data combination and reasoning, providing a key technical means for subsequent unmanned operations.

[0070] Figure 8 The structural schematic diagram of the vegetable multi-modal knowledge graph construction device provided by the present invention is as Figure 8 As shown in the figure, the present invention provides a vegetable multi-modal knowledge graph construction device, including a generation module 801, a first determination module 802, a second determination module 803, and a third acquisition module 804.

[0071] The generation module 801 is used to collect vegetable knowledge multi-modal data; the vegetable knowledge multi-modal data includes text-modal data, document-modal data, and picture-modal data in the vegetable field; The first determination module 802 is used to obtain multi-modal semantic fusion features based on the vegetable knowledge multi-modal data; The second determination module 803 is used to identify entities in the vegetable knowledge multi-modal data based on the multi-modal semantic fusion features by using an entity extraction model, and obtain the relationships between the entities by using an entity relationship extraction model; The third acquisition module 804 is used to obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0072] Specifically, the above-mentioned vegetable multi-modal knowledge graph construction device provided by the present invention can implement all the method steps implemented by the embodiments of the above-mentioned vegetable multi-modal knowledge graph construction method, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described herein.

[0073] It should be noted that the division of units / modules in the above embodiments of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0074] Figure 9 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 9 shown, the electronic device may include: a processor 901, a communication interface 902, a memory 903, and a communication bus 904. Among them, the processor 901, the communication interface 902, and the memory 903 complete communication with each other through the communication bus 904. The processor 901 can call the logical instructions in the memory 903 to execute the method for constructing a vegetable multi-modal knowledge graph, and the method includes: Collecting multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in document modality, and data in picture modality in the field of vegetables; Obtaining multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge; Based on the multi-modal semantic fusion features, using an entity extraction model to identify entities in the multi-modal data of vegetable knowledge, and using an entity relationship extraction model to obtain the relationships between the entities; Obtaining a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0075] In some embodiments, the obtaining of the multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge includes: Obtaining vegetable knowledge text data based on the data in text modality and data in document modality in the field of vegetables, and using the data in picture modality in the field of vegetables as vegetable knowledge picture data; Using the pre-trained language model Roberta to process the vegetable knowledge text data to obtain a text semantic segmentation result, using the semantic segmentation model Yolov8-seg to process the vegetable knowledge picture data to obtain a picture semantic segmentation result, and using a structure pre-trained model to obtain the text semantic structure corresponding to the vegetable knowledge text data and the image semantic structure corresponding to the vegetable knowledge picture data to obtain structure information; Mapping the text semantic segmentation result, the picture semantic segmentation result, and the structure information into the same semantic subspace to obtain multi-modal semantic fusion features.

[0076] In some embodiments, obtaining vegetable knowledge text data from the data in the text modality and the data in the document modality in the vegetable field includes: Using optical character recognition (OCR) technology to obtain the text in the data of the document modality; Taking the data in the text modality and the text in the data of the document modality as vegetable knowledge text data.

[0077] In some embodiments, the knowledge stored in the schema layer of the vegetable multi-modal knowledge graph is divided into vegetable variety schema, vegetable cultivation schema, pest and disease schema, and agricultural machinery schema.

[0078] In some embodiments, the ontology library for managing the schema layer includes vegetable operation knowledge classification, vegetable operation object attributes, and vegetable data attributes; Among them, the vegetable operation knowledge classification includes vegetable production personnel, vegetable production environment, vegetable operation equipment, vegetable varieties, and vegetable pests and diseases; the vegetable production personnel include personnel information and personnel behavior.

[0079] In some embodiments, using the entity relationship extraction model to obtain the relationships between the entities includes: Using the entity relationship extraction model CasRel to annotate the position information of the head entity and the tail entity of the entity; Classifying the relationships of the entities according to the annotation results to obtain the relationships between the entities.

[0080] Specifically, the processor 901 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD), and the processor may also adopt a multi-core architecture.

[0081] When the logical instructions in the memory 903 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0082] In some embodiments, a computer program product is also provided. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the vegetable multi-modal knowledge graph construction method provided in each of the above method embodiments. The method includes: Collect multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in file modality, and data in picture modality in the field of vegetables; Obtain multi-modal semantic fusion features based on the multi-modal data of vegetable knowledge; Based on the multi-modal semantic fusion features, use an entity extraction model to identify entities in the multi-modal data of vegetable knowledge, and use an entity relationship extraction model to obtain the relationships between the entities; Obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0083] Specifically, the above computer program product provided by the embodiments of this application can implement all the method steps implemented by each of the above method embodiments, and can achieve the same technical effects. Here, the same parts and beneficial effects as those in the method embodiments will not be specifically described again.

[0084] In some embodiments, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and the computer program is used to enable a computer to execute the vegetable multi-modal knowledge graph construction method provided in each of the above method embodiments. The method includes: Collect multi-modal data of vegetable knowledge; the multi-modal data of vegetable knowledge includes data in text modality, data in file modality, and data in picture modality in the field of vegetables; Obtain multi-modal semantic fusion features based on the vegetable knowledge multi-modal data; Based on the multi-modal semantic fusion features, use an entity extraction model to identify entities in the vegetable knowledge multi-modal data, and use an entity relationship extraction model to obtain the relationships between the entities; Obtain a vegetable multi-modal knowledge graph according to the entities and the relationships between the entities.

[0085] Specifically, the computer-readable storage medium provided by the present invention can implement all the method steps implemented by the above method embodiments, and can achieve the same technical effects. Here, the same parts and beneficial effects as those in the method embodiments will not be specifically described again.

[0086] It should be noted that: the computer-readable storage medium can be any available medium or data storage device that can be accessed by a processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROM, EPROM, EEPROM, non-volatile memories (NAND FLASH), solid-state drives (SSD)).

[0087] In addition, it should be noted that: the terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple.

[0088] The term "multiple" in the present invention means two or more, and other quantifiers are similar.

[0089] "Determining B based on A" in the present invention means that the factor A should be considered when determining B. It is not limited to "determining B only based on A", but also should include: "determining B based on A and C", "determining B based on A, C, and E", "determining C based on A, and further determining B based on C", etc. In addition, it can also include using A as a condition for determining B. For example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B"; for another example, "when A meets the third condition, determine B based on the first parameter". Of course, it can also be a condition for using A as a factor for determining B. For example, "when A meets the first condition, use the first method to determine C, and further determine B based on C", etc.

[0090] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0091] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0092] These processor-executable instructions can also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the processor-readable memory produce a manufactured article including an instruction means that implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0093] These processor-executable instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0094] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for constructing a vegetable multimodal knowledge graph, characterized in that: include: Collecting multimodal data of vegetable knowledge; the multimodal data of vegetable knowledge includes text modal data, file modal data and picture modal data in the field of vegetables; Acquiring multimodal semantic fusion features based on the vegetable knowledge multimodal data; Based on the multimodal semantic fusion features, an entity extraction model is used to identify entities in the vegetable knowledge multimodal data, and an entity relationship extraction model is used to obtain the relationship between the entities; A vegetable multimodal knowledge graph is obtained based on the entities and the relationships between the entities.

2. The method for constructing a vegetable multimodal knowledge graph according to claim 1, characterized in that: The step of acquiring multimodal semantic fusion features based on the vegetable knowledge multimodal data includes: Acquire vegetable knowledge text data based on the text modality data and the document modality data in the vegetable field, and use the image modality data in the vegetable field as vegetable knowledge image data; The vegetable knowledge text data is processed using the pre-trained language model Roberta to obtain a text semantic segmentation result, the vegetable knowledge picture data is processed using the semantic segmentation model Yolov8-seg to obtain a picture semantic segmentation result, and the text semantic structure corresponding to the vegetable knowledge text data and the image semantic structure corresponding to the vegetable knowledge picture data are obtained using a structural pre-training model to obtain structural information; The text semantic segmentation result, the image semantic segmentation result and the structural information are mapped into the same semantic subspace to obtain a multimodal semantic fusion feature.

3. The method for constructing a vegetable multimodal knowledge graph according to claim 2, characterized in that: The method of obtaining vegetable knowledge text data based on text modality data and document modality data in the vegetable field includes: Using optical character recognition (OCR) technology to obtain text in the data of the file modality; The data in the text mode and the characters in the data in the file mode are used as vegetable knowledge text data.

4. The method for constructing a vegetable multimodal knowledge graph according to claim 1, characterized in that: The knowledge stored in the pattern layer of the vegetable multimodal knowledge graph is divided into vegetable variety patterns, vegetable cultivation patterns, pest and disease patterns, and agricultural machinery patterns.

5. The method for constructing a vegetable multimodal knowledge graph according to claim 4, characterized in that: The ontology library used to manage the model layer includes vegetable operation knowledge classification, vegetable operation object attributes, and vegetable data attributes; Among them, the vegetable operation knowledge classification includes vegetable production personnel, vegetable production environment, vegetable operation equipment, vegetable varieties and vegetable pests and diseases; the vegetable production personnel includes personnel information and personnel behavior.

6. The method for constructing a vegetable multimodal knowledge graph according to claim 1, characterized in that: The method of obtaining the relationship between the entities by using the entity relationship extraction model includes: The entity relationship extraction model CasRel is used to annotate the location information of the head entity and the tail entity of the entity; The entities are classified according to the labeling results to obtain the relationships between the entities.

7. A device for constructing a vegetable multimodal knowledge graph, characterized in that: include: A collection module, used to collect multimodal data of vegetable knowledge; The vegetable knowledge multimodal data includes text modal data, file modal data and picture modal data in the field of vegetables; A first acquisition module is used to acquire multimodal semantic fusion features based on the vegetable knowledge multimodal data; A second acquisition module is used to identify entities in the vegetable knowledge multimodal data using an entity extraction model based on the multimodal semantic fusion feature, and to obtain the relationship between the entities using an entity relationship extraction model; The third acquisition module is used to acquire a vegetable multimodal knowledge graph based on the entities and the relationships between the entities.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the method for constructing a multimodal knowledge graph of vegetables as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing a multimodal knowledge graph of vegetables as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for constructing a multimodal knowledge graph of vegetables as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Power entity joint relation extraction method and system based on multi-modal large model

    CN121524910A