Deep learning-based cultivated land protection text knowledge extraction method and system

By processing multimodal farmland protection data based on deep learning, a six-element directed graph of text knowledge was constructed, which solved the problem of poor extraction of text knowledge of traditional methods in professional fields, and achieved efficient and accurate knowledge extraction and data organization.

CN120011559AActive Publication Date: 2025-05-16HAINAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510111079.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The traditional artificial intelligence neural network method is not effective in extracting text knowledge of cultivated land protection in professional fields, and the computing power requirements for model training and deployment are high, resulting in a lack of correlation and organization of the extracted data.

Method used

Using a deep learning-based method, by obtaining multimodal farmland protection data and cleaning data, using a fusion rule prior and expansion gate convolution neural network for text data processing, combining the graph neural network to process remote sensing images, an unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is constructed for entity recognition and relationship construction, and finally constructing a six-element directed graph for text knowledge for farmland protection.

Benefits of technology

It realizes efficient and accurate relationship extraction and attribute extraction, provides high-quality text knowledge of cultivated land protection, and supports data, technical and service support for cultivated land protection and governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011559A_ABST
    Figure CN120011559A_ABST
Patent Text Reader

Abstract

The invention relates to a cultivated land protection text knowledge extraction method and system based on deep learning. The method comprises the steps of performing data preprocessing based on a hierarchical clustering algorithm and a data cleaning method of information entropy to obtain to-be-extracted multi-modal data; text data in the to-be-extracted multi-modal data is processed through the fusion rule prior and the expansion gate convolutional neural network; processing a remote sensing image in the to-be-extracted multi-modal data through the graph neural network; obtaining geographic information in the to-be-extracted multi-modal data and constructing a relationship between an image and a text; the method comprises the following steps: constructing an unstructured cultivated land text semantic extraction network to perform entity recognition, and extracting an entity-attribute-attribute value triple and an entity-relation-attribute information triple according to entity information and cultivated land attribute information; the fusion rule prior and the expansion gate convolutional neural network are used for entity extraction and relation construction, and efficient and accurate relation extraction and attribute extraction can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of farmland protection text knowledge extraction, and in particular to a farmland protection text knowledge extraction method and system based on deep learning. Background Art

[0002] The content of cultivated land protection data is diverse, including text, images, multimedia data, sensor data and other data in different modes, with the characteristics of complex form, diverse modes and huge data volume. These data sources are scattered, and there are great differences in expression methods, knowledge dimensions, data structures, etc. There is no fixed form of expression, and the writing is highly professional, making it difficult to apply existing models to extract knowledge from the content of cultivated land protection texts.

[0003] In terms of knowledge extraction, in 2019, Liu et al. connected entities in three knowledge graphs that all contain rich digital text and image information through the SameAs predicate, and realized relational reasoning between different resources. In 2020, Wang et al. proposed the multimodal graph Richpedia, which uses rule-based relational extraction templates and hyperlink information in Wikipedia image descriptions to generate multimodal semantic relationships between image entities. In 2022, Bai C proposed a new model, namely E-commerce Knowledge Extraction Based on Multimodal Machine Reading Comprehension (EKE-MRC), to better extract e-commerce product attributes.

[0004] However, the traditional artificial intelligence neural network method has achieved good knowledge extraction results under the drive of large-scale, high-quality data, but the extraction effect is not satisfactory in professional fields, and the training and deployment of numerical models require high computing power. Therefore, the data extracted by the traditional method of extracting knowledge from farmland protection texts often lacks relevance and organization, and the relevant knowledge is difficult to be used efficiently. Summary of the invention

[0005] Based on this, in order to solve the above technical problems, a deep learning-based text knowledge extraction method and system for farmland protection is provided, which can realize efficient and accurate relationship extraction and attribute extraction, thereby providing data, technology and service support for farmland protection and governance.

[0006] A method for extracting textual knowledge of farmland protection based on deep learning, the method comprising:

[0007] Acquire multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted;

[0008] The text data in the multimodal data to be extracted is processed by fusing rule priors and dilation gate convolutional neural networks, and feature vectors are extracted to obtain entity information of the text;

[0009] The remote sensing image in the multimodal data to be extracted is processed by a graph neural network to obtain entity information of the image; the geographic information in the multimodal data to be extracted is obtained, and the relationship between the image and the text is constructed according to the spatial information in the geographic information; and the cultivated land attribute information is determined according to the geographic information;

[0010] Constructing an unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF, using the unstructured cultivated land text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extracting entity-attribute-attribute value triples according to the entity information and cultivated land attribute information; extracting entity-relationship-attribute information triples according to the entity information, relationship, and cultivated land attribute information;

[0011] According to the entity-attribute-attribute value triples and entity-relationship-attribute information triples, a six-element directed graph of text knowledge on farmland protection is constructed.

[0012] In one embodiment, multimodal farmland protection data is obtained, and a data cleaning method based on a hierarchical clustering algorithm and information entropy is used to pre-process the multimodal farmland protection data, including:

[0013] Collecting remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and storing the multimodal farmland protection data in a database;

[0014] When performing knowledge extraction, the multimodal farmland protection data is extracted from the database;

[0015] The multimodal farmland protection data is cleaned and format converted using a data cleaning method based on a hierarchical clustering algorithm and information entropy.

[0016] In one embodiment, the text data in the multimodal data to be extracted is processed by fusing rule priors and dilation gate convolutional neural networks, and feature vectors are extracted to obtain entity information of the text, including:

[0017] The expansion gate convolutional neural network performs character segmentation and word segmentation processing on the text data in the multimodal data to be extracted to obtain processed text data;

[0018] Inputting the processed text data into the dilation gate convolution layer in the dilation gate convolution neural network, and extracting feature vectors through the dilation gate convolution layer;

[0019] Capturing dependencies in the text data through a self-attention mechanism in the dilation gate convolutional neural network;

[0020] The entity information of the text is obtained based on the feature vector and the dependency relationship.

[0021] In one embodiment, the method further comprises:

[0022] Obtaining a word mixture vector according to the text data in the multimodal data to be extracted, and inputting the word mixture vector into the expansion gate convolutional neural network;

[0023] Relationships are captured through the self-attention mechanism in the dilation gate convolutional neural network and verified in combination with prior knowledge to identify the subject in the multimodal data to be extracted.

[0024] In one embodiment, the method further comprises:

[0025] Obtaining a position vector according to the geographic information in the multimodal data to be extracted, and inputting the position vector into the expansion gate convolutional neural network;

[0026] Relationships are captured through the self-attention mechanism in the dilation gate convolutional neural network, and feature extraction is performed through the dilation gate convolutional layer, which is verified in combination with prior knowledge to identify the objects in the multimodal data to be extracted.

[0027] In one embodiment, the unstructured farmland text semantic extraction network is used to perform entity recognition on the multimodal data to be extracted, and entity-attribute-attribute value triples are extracted according to the entity information and farmland attribute information, including:

[0028] Capturing the word dependency relationship within the sentence in the multimodal data to be extracted through the multi-layer encoder of Bert in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF;

[0029] The BiLSTM part in the unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF performs contextual analysis of text semantics on the multimodal data to be extracted, identifies the semantic structure of the cultivated land protection text, and extracts text features according to the semantic structure;

[0030] The entity-attribute-attribute value triple of farmland protection is extracted according to the word dependency relationship and text features through the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF.

[0031] In one of the embodiments, the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is trained using an end-to-end method.

[0032] In one embodiment, the method further comprises:

[0033] Collecting unstructured data on farmland protection and manually labeling the unstructured data;

[0034] The labeled data is divided into training set and validation set, and the entities and relationships are obtained and stored in the database;

[0035] The unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is verified using the verification set, and the entity recognition effect is evaluated using F1.

[0036] A text knowledge extraction system for farmland protection based on deep learning, the system comprising:

[0037] A data processing module is used to obtain multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted;

[0038] A first extraction module is used to process the text data in the multimodal data to be extracted by fusing rule priors and dilation gate convolutional neural networks, and extract feature vectors to obtain entity information of the text;

[0039] The second extraction module is used to process the remote sensing image in the multimodal data to be extracted through a graph neural network to obtain entity information of the image; obtain geographic information in the multimodal data to be extracted, and construct a relationship between the image and the text according to the spatial information in the geographic information; and determine the cultivated land attribute information according to the geographic information;

[0040] A knowledge extraction module is used to construct an unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF, use the unstructured cultivated land text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extract entity-attribute-attribute value triples according to the entity information and cultivated land attribute information; extract entity-relationship-attribute information triples according to the entity information, relationship, and cultivated land attribute information;

[0041] The directed graph construction module is used to construct a six-tuple directed graph of text knowledge on farmland protection based on the entity-attribute-attribute value triples and the entity-relationship-attribute information triples.

[0042] In one of the embodiments, the data processing module is also used to collect remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and store the multimodal farmland protection data in a database; when performing knowledge extraction, the multimodal farmland protection data is extracted from the database; and the multimodal farmland protection data is cleaned and format converted using a data cleaning method based on a hierarchical clustering algorithm and information entropy.

[0043] The above-mentioned deep learning-based farmland protection text knowledge extraction method and system can effectively eliminate the interference of invalid and redundant data on subsequent text knowledge extraction by processing data using a data cleaning method based on a hierarchical clustering algorithm and information entropy; it uses a fusion rule prior and an expansion gate convolutional neural network for entity extraction and relationship construction, so that knowledge can complement and inherit each other, and can achieve efficient and accurate relationship extraction and attribute extraction, thereby providing data, technology and service support for farmland protection and governance. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is an application environment diagram of a method for extracting textual knowledge about farmland protection based on deep learning in one embodiment;

[0045] Figure 2 A schematic diagram of a flow chart of a method for extracting textual knowledge on farmland protection based on deep learning in one embodiment;

[0046] Figure 3 A schematic diagram of a convolutional neural network model framework integrating rule priors and dilation gates in one embodiment;

[0047] Figure 4 A schematic diagram of a method for extracting knowledge from multimodal farmland big data in one embodiment;

[0048] Figure 5 A schematic diagram of the Bert-BiLSTM-CRF network structure in one embodiment;

[0049] Figure 6 A schematic diagram of an interface of knowledge extraction results in one embodiment;

[0050] Figure 7 It is a structural block diagram of a farmland protection text knowledge extraction system based on deep learning in one embodiment;

[0051] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe extraction modules, but these extraction modules are not limited by these terms. These terms are only used to distinguish a first extraction module from another extraction module. For example, without departing from the scope of this application, a first extraction module may be referred to as a second extraction module, and similarly, a second extraction module may be referred to as a first extraction module. Both the first extraction module and the second extraction module are extraction modules, but they are not the same extraction module.

[0054] The method for extracting knowledge from farmland protection text based on deep learning provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment includes a computer device 110. The computer device 110 can obtain multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted; the computer device 110 can process the text data in the multimodal data to be extracted by fusing rule priors and expansion gate convolutional neural networks, and extract feature vectors to obtain entity information of the text; the computer device 110 can process the remote sensing images in the multimodal data to be extracted by a graph neural network to obtain entity information of the image; obtain geographic information in the multimodal data to be extracted, and construct image and spatial information according to the spatial information in the geographic information. The computer device 110 can construct an unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, use the unstructured farmland text semantic extraction network to perform entity recognition on the extracted multimodal data, and extract entity-attribute-attribute value triples according to entity information and farmland attribute information; extract entity-relationship-attribute information triples according to entity information, relationship, and farmland attribute information; the computer device 110 can construct a six-element directed graph of text knowledge of farmland protection according to entity-attribute-attribute value triples and entity-relationship-attribute information triples. Among them, the computer device 110 can be, but not limited to, various personal computers, laptops, smart phones, robots and other devices.

[0055] In one embodiment, Figure 2 As shown, a method for extracting text knowledge of farmland protection based on deep learning is provided, which includes the following steps:

[0056] Step 202, obtaining multimodal farmland protection data, and preprocessing the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted.

[0057] Among them, the multimodal farmland protection data may include data related to farmland protection, such as remote sensing images, text information, and geographic information. The computer equipment may use a data cleaning method based on a hierarchical clustering algorithm and information entropy to clean and convert the multimodal farmland protection data in order to extract its entity information.

[0058] In one embodiment, a deep learning-based farmland protection text knowledge extraction method is provided, which may also include a data preprocessing process, the specific process including: collecting remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and storing the multimodal farmland protection data in a database; when performing knowledge extraction, extracting the multimodal farmland protection data from the database; using a data cleaning method based on a hierarchical clustering algorithm and information entropy to clean the multimodal farmland protection data and perform format conversion processing.

[0059] Unstructured texts such as policies and regulations related to farmland protection and land data reports are the core of farmland protection knowledge learning. In this embodiment, data collection and database establishment can be completed, which is convenient for data processing.

[0060] Step 204, the text data in the multimodal data to be extracted is processed by fusing the rule prior and the dilation gate convolutional neural network, and the feature vector is extracted to obtain the entity information of the text.

[0061] A farmland protection entity extraction and relationship construction model that integrates rule priors and dilation gate convolutional neural networks can be designed in computer equipment. The receptive field of the output unit can be increased through dilation convolution.

[0062] In one embodiment, a method for extracting farmland protection text knowledge based on deep learning is provided, which may include a process of data processing by fusing rule priors and an expansion gate convolutional neural network. The specific process includes: the expansion gate convolutional neural network performs character and word segmentation processing on the text data in the extracted multimodal data to obtain processed text data; the processed text data is input into the expansion gate convolutional layer in the expansion gate convolutional neural network, and feature vectors are extracted through the expansion gate convolutional layer; the dependency relationship in the text data is captured through the self-attention mechanism in the expansion gate convolutional neural network; and the entity information of the text is obtained based on the feature vector and the dependency relationship.

[0063] Specifically, Figure 3As shown, in one embodiment, a method for extracting text knowledge on farmland protection based on deep learning is provided, which may also include a process of identifying a subject, and the specific process includes: obtaining a word mixture vector based on text data in the multimodal data to be extracted, and inputting the word mixture vector into a dilation gate convolutional neural network; capturing relationships through the self-attention mechanism in the dilation gate convolutional neural network, and verifying it in combination with prior knowledge, so as to identify the subject in the multimodal data to be extracted.

[0064] In one embodiment, Figure 3 As shown, a method for extracting text knowledge on farmland protection based on deep learning can also include a process of identifying an object. The specific process includes: obtaining a position vector based on geographic information in the multimodal data to be extracted, and inputting the position vector into an expansion gate convolutional neural network; capturing relationships through the self-attention mechanism in the expansion gate convolutional neural network, and extracting features through the expansion gate convolutional layer, and verifying with prior knowledge to identify the object in the multimodal data to be extracted.

[0065] like Figure 3 As shown, in this embodiment, after the computer device obtains the multimodal data to be extracted, the expansion gate convolutional neural network model first performs character segmentation and word segmentation on the text data in the multimodal data to be extracted, and then inputs the text data in the training set into the expansion gate convolution layer for feature vector extraction. When processing text data, it can expand the text context width and extract the global information of the text, thereby better understanding the text semantics. Among them, the formula for processing data by the expansion gate one-dimensional convolution layer can be expressed as: y = ; Where Conv1D(·) represents one-dimensional convolution, X represents the vector sequence to be processed, represents point-wise multiplication, and σ(·) represents the gating function.

[0066] Step 206, the remote sensing image in the multimodal data to be extracted is processed through a graph neural network to obtain entity information of the image; the geographic information in the multimodal data to be extracted is obtained, and the relationship between the image and the text is constructed according to the spatial information in the geographic information; and the cultivated land attribute information is determined according to the geographic information.

[0067] In order to make full use of the rich information structure brought by multimodal data, entities, relations, and attributes are extracted in a differentiated manner, such as Figure 4 As shown. Specifically, for remote sensing images, text information and geographic information, convolutional neural networks and graph neural networks are used to extract them respectively, so that more suitable feature information can be obtained. In addition, the spatial information contained in geographic information data is fully utilized to accurately construct the relationship between images and texts. Geographic location information can also be used to obtain cultivated land information in images and texts, providing accurate data support for cultivated land protection.

[0068] Step 208, construct an unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, use the unstructured farmland text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extract entity-attribute-attribute value triples based on entity information and farmland attribute information; extract entity-relationship-attribute information triples based on entity information, relationship, and farmland attribute information.

[0069] In view of the different forms of cultivated land documents, the variety of land attribute types, and the lack of a unified cultivated land attribute extraction model in the field of cultivated land protection, an unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF is constructed. In one embodiment, the unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF is trained using an end-to-end method, such as Figure 5 As shown in FIG. 1 , the constructed unstructured farmland text semantic extraction network Bert-BiLSTM-CRF may include a Bert module, a BiLSTM module, and a CRF module.

[0070] In one embodiment, a method for extracting knowledge from farmland protection text based on deep learning is provided, which may also include a process of using an unstructured farmland text semantic extraction network Bert-BiLSTM-CRF to process data. The specific process includes: capturing word dependencies within sentences in multimodal data to be extracted through the multi-layer encoder of Bert in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF; performing contextual analysis of text semantics on the multimodal data to be extracted through the BiLSTM part in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, identifying the semantic structure of the farmland protection text, and extracting text features based on the semantic structure; extracting entity-attribute-attribute value triples of farmland protection based on word dependencies and text features through the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF.

[0071] Specifically, Figure 5 As shown in the figure, the pre-trained Bert multi-layer encoder can capture the word dependency relationship within the sentence, identify the semantic structure of the cultivated land protection text by parsing the context of the text semantics, and extract text features covering long-distance semantics. Secondly, the Bert-BiLSTM-CRF network is trained using an end-to-end method to perform word segmentation and semantic recognition on the cultivated land protection text. The words are divided according to Chinese grammar and supplemented by cultivated land protection entities and concepts. By establishing the long-distance dependency between cultivated land protection entities and sentences, natural language understanding technology is used to perform semantic understanding and identify text intent, and the entity-attribute-attribute value triples of cultivated land protection are extracted.

[0072] Step 210, constructing a six-element directed graph of textual knowledge on farmland protection based on entity-attribute-attribute value triples and entity-relationship-attribute information triples.

[0073] Computer equipment can construct a six-element directed graph G of text knowledge for farmland protection based on entity-attribute-attribute value triples and entity-relationship-attribute information triples. T =(E,R,A,V,T r ,T a ), where E = {e1, e2, …, e |e|}, R = {r1, r2, ..., r |R|}、A={a1,a2,…,a |A|}、V={v1,v2,…,v |V|} refer to the set of cultivated land protection text knowledge entities, relations, attributes and values, respectively. r ={(h,r,t)|h,r∈E,r∈R},T a ={(e,a,v)|e∈E,a∈A,v∈V} is a set of relation triples and attribute triples.

[0074] In one embodiment, a method for extracting knowledge from farmland protection text based on deep learning is provided, which may also include a process for evaluating the effect of entity recognition. The specific process includes: collecting unstructured data on farmland protection and manually annotating the unstructured data; dividing the annotated data into a training set and a verification set, obtaining entities and relationships and storing them in a database; using the verification set to verify the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, and using F1 to evaluate the entity recognition effect.

[0075] A method for extracting farmland protection text knowledge based on deep learning provided in this application is used to extract farmland protection related data, and unstructured data such as technical regulations, planning texts, reports, and literature materials for farmland protection are manually annotated. 80% of the data is used as a training set and 20% as a validation set, and 1518 entities and 1716 relationships are obtained. The entity and relationship data are stored in the Neo4j database. ; F1 is used to evaluate the entity recognition effect. The calculation method formula can be expressed as: TP is a correct positive example, FP is a wrong positive example, and FN is a wrong negative example. The labeling accuracy is shown in the following table:

[0076] Labeling accuracy of various entity types

[0077]

[0078] In this embodiment, after extracting farmland protection related data using a farmland protection text knowledge extraction method based on deep learning provided by this application, the corresponding knowledge extraction display interface is as follows: Figure 6 As shown, the input multimodal data to be extracted can be "land economics originates from basic land terms", and the extraction results are "land", "economics", "originates from", "land", "basic", and "terms"; among them, the subject is land economics; the relationship is derived from; and the other subject is land terms.

[0079] It should be understood that, although the various steps in the above-mentioned flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flow chart may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0080] In one embodiment, Figure 7 As shown, a system for extracting knowledge from text on farmland protection based on deep learning is provided, comprising: a data processing module 710, a first extraction module 720, a second extraction module 730, a knowledge extraction module 740 and a directed graph construction module 750, wherein:

[0081] The data processing module 710 is used to obtain multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted;

[0082] A first extraction module 720 is used to process text data in the multimodal data to be extracted by fusing rule priors and dilation gate convolutional neural networks, and extract feature vectors to obtain entity information of the text;

[0083] The second extraction module 730 is used to process the remote sensing image in the multimodal data to be extracted through a graph neural network to obtain entity information of the image; obtain geographic information in the multimodal data to be extracted, and construct a relationship between the image and the text according to the spatial information in the geographic information; and determine the cultivated land attribute information according to the geographic information;

[0084] The knowledge extraction module 740 is used to construct an unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, use the unstructured farmland text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extract entity-attribute-attribute value triples according to entity information and farmland attribute information; extract entity-relationship-attribute information triples according to entity information, relationship, and farmland attribute information;

[0085] The directed graph construction module 750 is used to construct a six-tuple directed graph of text knowledge on farmland protection based on entity-attribute-attribute value triples and entity-relationship-attribute information triples.

[0086] In one embodiment, the data processing module 710 is also used to collect remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and store the multimodal farmland protection data in a database; when performing knowledge extraction, the multimodal farmland protection data is extracted from the database; and the multimodal farmland protection data is cleaned and format converted using a data cleaning method based on a hierarchical clustering algorithm and information entropy.

[0087] In one embodiment, the first extraction module 720 is also used for the expansion gate convolutional neural network to perform character and word segmentation on the text data in the extracted multimodal data to obtain processed text data; the processed text data is input into the expansion gate convolutional layer in the expansion gate convolutional neural network, and the feature vector is extracted through the expansion gate convolutional layer; the dependency relationship in the text data is captured through the self-attention mechanism in the expansion gate convolutional neural network; and the entity information of the text is obtained based on the feature vector and the dependency relationship.

[0088] In one embodiment, the first extraction module 720 is also used to obtain a word mixture vector based on the text data in the multimodal data to be extracted, and input the word mixture vector into the dilation gate convolutional neural network; the relationship is captured through the self-attention mechanism in the dilation gate convolutional neural network, and verified in combination with prior knowledge to identify the subject in the multimodal data to be extracted.

[0089] In one embodiment, the first extraction module 720 is also used to obtain a position vector based on the geographic information in the multimodal data to be extracted, and input the position vector into the dilation gate convolutional neural network; the relationship is captured through the self-attention mechanism in the dilation gate convolutional neural network, and feature extraction is performed through the dilation gate convolutional layer, which is verified in combination with prior knowledge to identify the object in the multimodal data to be extracted.

[0090] In one embodiment, the knowledge extraction module 740 is also used to capture the word dependency relationship within the sentence in the multimodal data to be extracted through the multi-layer encoder of Bert in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF; perform text semantic context analysis on the multimodal data to be extracted through the BiLSTM part in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, identify the semantic structure of the farmland protection text, and extract text features based on the semantic structure; extract the entity-attribute-attribute value triple of farmland protection based on word dependency and text features through the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF.

[0091] In one embodiment, the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is trained using an end-to-end method.

[0092] In one embodiment, the knowledge extraction module 740 is also used to collect unstructured data on farmland protection and manually annotate the unstructured data; divide the annotated data into a training set and a verification set, obtain entities and relationships and store them in a database; use the verification set to verify the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF, and use F1 to evaluate the entity recognition effect.

[0093] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for extracting text knowledge of farmland protection based on deep learning is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0094] Those skilled in the art will understand that Figure 8The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0095] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of a method for extracting text knowledge of farmland protection based on deep learning are implemented.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for extracting text knowledge of farmland protection based on deep learning are implemented.

[0097] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0098] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0099] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for extracting text knowledge of farmland protection based on deep learning, characterized in that: The method comprises: Acquire multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted; The text data in the multimodal data to be extracted is processed by fusing rule priors and dilation gate convolutional neural networks, and feature vectors are extracted to obtain entity information of the text; The remote sensing image in the multimodal data to be extracted is processed by a graph neural network to obtain entity information of the image; the geographic information in the multimodal data to be extracted is obtained, and the relationship between the image and the text is constructed according to the spatial information in the geographic information; and the cultivated land attribute information is determined according to the geographic information; Constructing an unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF, using the unstructured cultivated land text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extracting entity-attribute-attribute value triples according to the entity information and cultivated land attribute information; extracting entity-relationship-attribute information triples according to the entity information, relationship, and cultivated land attribute information; According to the entity-attribute-attribute value triples and entity-relationship-attribute information triples, a six-element directed graph of text knowledge on farmland protection is constructed.

2. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 1 is characterized in that: Acquire multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy, including: Collecting remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and storing the multimodal farmland protection data in a database; When performing knowledge extraction, the multimodal farmland protection data is extracted from the database; The multimodal farmland protection data is cleaned and format converted using a data cleaning method based on a hierarchical clustering algorithm and information entropy.

3. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 1 is characterized in that: The text data in the multimodal data to be extracted is processed by fusing rule priors and dilation gate convolutional neural networks, and feature vectors are extracted to obtain entity information of the text, including: The expansion gate convolutional neural network performs character segmentation and word segmentation processing on the text data in the multimodal data to be extracted to obtain processed text data; Inputting the processed text data into the dilation gate convolution layer in the dilation gate convolution neural network, and extracting feature vectors through the dilation gate convolution layer; Capturing dependencies in the text data through a self-attention mechanism in the dilation gate convolutional neural network; The entity information of the text is obtained based on the feature vector and the dependency relationship.

4. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 3 is characterized in that: The method further comprises: Obtaining a word mixture vector according to the text data in the multimodal data to be extracted, and inputting the word mixture vector into the expansion gate convolutional neural network; Relationships are captured through the self-attention mechanism in the dilation gate convolutional neural network and verified in combination with prior knowledge to identify the subject in the multimodal data to be extracted.

5. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 3 is characterized in that: The method further comprises: Obtaining a position vector according to the geographic information in the multimodal data to be extracted, and inputting the position vector into the expansion gate convolutional neural network; Relationships are captured through the self-attention mechanism in the dilation gate convolutional neural network, and feature extraction is performed through the dilation gate convolutional layer, which is verified in combination with prior knowledge to identify the objects in the multimodal data to be extracted.

6. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 1 is characterized in that: The unstructured cultivated land text semantic extraction network is used to perform entity recognition on the multimodal data to be extracted, and entity-attribute-attribute value triples are extracted according to the entity information and cultivated land attribute information, including: Capturing the word dependency relationship within the sentence in the multimodal data to be extracted through the multi-layer encoder of Bert in the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF; The BiLSTM part in the unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF performs contextual analysis of text semantics on the multimodal data to be extracted, identifies the semantic structure of the cultivated land protection text, and extracts text features according to the semantic structure; The entity-attribute-attribute value triple of farmland protection is extracted according to the word dependency relationship and text features through the unstructured farmland text semantic extraction network Bert-BiLSTM-CRF.

7. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 6 is characterized in that: The unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is trained using an end-to-end method.

8. The method for extracting textual knowledge of farmland protection based on deep learning according to claim 1 is characterized in that: The method further comprises: Collecting unstructured data on farmland protection and manually labeling the unstructured data; The labeled data is divided into training set and validation set, and entities and relationships are obtained and stored in the database; The unstructured farmland text semantic extraction network Bert-BiLSTM-CRF is verified using the verification set, and the entity recognition effect is evaluated using F1.

9. A text knowledge extraction system for farmland protection based on deep learning, characterized in that: The system comprises: A data processing module is used to obtain multimodal farmland protection data, and pre-process the multimodal farmland protection data using a data cleaning method based on a hierarchical clustering algorithm and information entropy to obtain multimodal data to be extracted; A first extraction module is used to process the text data in the multimodal data to be extracted by fusing rule priors and dilation gate convolutional neural networks, and extract feature vectors to obtain entity information of the text; The second extraction module is used to process the remote sensing image in the multimodal data to be extracted through a graph neural network to obtain entity information of the image; obtain geographic information in the multimodal data to be extracted, and construct a relationship between the image and the text according to the spatial information in the geographic information; and determine the cultivated land attribute information according to the geographic information; A knowledge extraction module is used to construct an unstructured cultivated land text semantic extraction network Bert-BiLSTM-CRF, use the unstructured cultivated land text semantic extraction network to perform entity recognition on the multimodal data to be extracted, and extract entity-attribute-attribute value triples according to the entity information and cultivated land attribute information; extract entity-relationship-attribute information triples according to the entity information, relationship, and cultivated land attribute information; The directed graph construction module is used to construct a six-tuple directed graph of text knowledge on farmland protection based on the entity-attribute-attribute value triples and the entity-relationship-attribute information triples.

10. The system for extracting farmland protection text knowledge based on deep learning according to claim 9 is characterized in that: The data processing module is also used to collect remote sensing images, text information, and geographic information from unstructured text as multimodal farmland protection data, and store the multimodal farmland protection data in a database; when performing knowledge extraction, the multimodal farmland protection data is extracted from the database; and the multimodal farmland protection data is cleaned and format converted using a data cleaning method based on a hierarchical clustering algorithm and information entropy.

Citation Information

Patent Citations

  • Document-level named entity recognition method based on hierarchical feature fusion of double graphs

    CN115906846A

  • Intelligent search method and system based on multi-source heterogeneous data

    CN116049454A

  • Ancient book text entity relationship extraction method and system based on deep learning

    CN118966226A

  • SVO entity information retrieval system

    US20230351111A1

  • Multimodal document information extraction method based on graph neural network, device and medium

    WO2023138023A1