Knowledge graph construction model training and knowledge graph construction method, device and equipment
Through multimodal data feature extraction and teacher-student model optimization, the problems of inconsistency and catastrophic forgetting of agricultural multimodal data fusion were solved, and a dynamic agricultural knowledge map was constructed, which improved the accuracy and efficiency of agricultural decision-making.
Patent Information
- Application Number
- CN202410747676.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-06-11
AI Technical Summary
The prior art is difficult to effectively process emerging entities and relationships when processing and integrating agricultural multimodal data, and ignores non-text data, resulting in catastrophic forgetting phenomena and modal data fusion inconsistency, affecting the performance of the model during continuous learning.
Multimodal data feature extraction and fusion method is used to extract image and text features through convolutional neural networks and BERT models, and the knowledge graph construction model is trained in combination with entity labels and relationship labels. The model training process is optimized using teacher-student model and gradient modulation technology.
A dynamic and comprehensive agricultural knowledge map was constructed, which improved the accuracy and efficiency of agricultural decision-making, and solved the problems of data type fusion inconsistency and catastrophic forgetting.
Smart Images

Figure CN118734948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of agriculture and artificial intelligence technology, in particular to the fields of computer natural language processing and computer vision technology, and specifically to a method, device and equipment for training a knowledge graph construction model and constructing a knowledge graph. Background Art
[0002] Agricultural data plays a vital role in agricultural production and management. As a key information source for agricultural systems, agricultural data plays a role in crop monitoring, pest and disease prevention, yield forecasting, and resource optimization management. Multimodal agricultural data, such as text and images, plays a crucial role in modern agricultural management. While these data contain a wealth of agricultural information, their diversity and complexity make extracting effective information from them challenging. Therefore, developing a method that can effectively and accurately process and integrate these multimodal data and construct a continuously updated agricultural knowledge graph is crucial for improving agricultural production efficiency and decision-making quality. Summary of the Invention
[0003] The present invention provides a method, device and equipment for training a knowledge graph construction model and constructing a knowledge graph to construct a dynamic and comprehensive agricultural knowledge graph.
[0004] According to one aspect of the present invention, a method for training a knowledge graph construction model is provided, the method comprising:
[0005] Acquire a target multimodal sample; the target multimodal sample includes a sample image and a sample text;
[0006] Performing feature extraction on the target multimodal sample to obtain sample graphic features of the target multimodal sample;
[0007] The knowledge graph construction model is trained based on the sample image and text features, entity labels and relationship labels.
[0008] According to another aspect of the present invention, a method for constructing a knowledge graph is provided, the method comprising:
[0009] Acquiring multimodal target data; the multimodal target data includes a target image and a target text;
[0010] A knowledge graph construction model is used to perform entity and relationship prediction on the multimodal target data to obtain a target knowledge graph of the multimodal target data; wherein, the knowledge graph construction model is trained according to the training method of the knowledge graph construction model described in any embodiment of the present invention.
[0011] According to another aspect of the present invention, a training device for a knowledge graph construction model is provided, the device comprising:
[0012] A target sample acquisition module is used to acquire a target multimodal sample; the target multimodal sample includes a sample image and a sample text;
[0013] A sample image and text feature determination module, configured to extract features from the target multimodal sample to obtain sample image and text features of the target multimodal sample;
[0014] The model training module is used to train the knowledge graph construction model based on the sample image and text features, entity labels and relationship labels.
[0015] According to another aspect of the present invention, a knowledge graph construction device is provided, the device comprising:
[0016] A target data acquisition module, configured to acquire multimodal target data; the multimodal target data includes a target image and a target text;
[0017] A knowledge graph determination module is used to use a knowledge graph construction model to perform entity and relationship prediction on the multimodal target data to obtain a target knowledge graph of the multimodal target data; wherein, the knowledge graph construction model is trained according to the training method of the knowledge graph construction model described in any embodiment of the present invention.
[0018] According to another aspect of the present invention, an electronic device is provided, comprising:
[0019] at least one processor; and
[0020] a memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the knowledge graph construction model or the knowledge graph construction method described in any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions, and the computer instructions are used to enable a processor to implement the training method for the knowledge graph construction model or the knowledge graph construction method described in any embodiment of the present invention when executed.
[0023] According to another aspect of the present invention, a computer program product is provided, characterized in that the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the training method of the knowledge graph construction model or the knowledge graph construction method according to any embodiment of the present invention.
[0024] The technical solution of an embodiment of the present invention obtains target multimodal samples; the target multimodal samples include sample images and sample text, and then performs feature extraction on the target multimodal samples to obtain sample image and text features of the target multimodal samples. The knowledge graph construction model is then trained based on the sample image and text features, entity labels, and relationship labels. This technical solution can effectively process and integrate multiple data types in the agricultural field, such as image and text data, to construct a dynamic and comprehensive knowledge graph, thereby enabling the knowledge graph to support agricultural decision-making and improve the accuracy and efficiency of agricultural management.
[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the first embodiment of the present invention;
[0028] Figure 2 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the second embodiment of the present invention;
[0029] Figure 3 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the third embodiment of the present invention;
[0030] Figure 4 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the fourth embodiment of the present invention;
[0031] Figure 5 This is a flowchart of a knowledge graph construction method provided in accordance with the fifth embodiment of the present invention;
[0032] Figure 6 2. It is a structural diagram of a training device for a knowledge graph construction model according to a sixth embodiment of the present invention;
[0033] Figure 7 This is a schematic diagram of the structure of a knowledge graph construction device provided according to the seventh embodiment of the present invention;
[0034] Figure 8 It is a structural diagram of an electronic device that implements the training method of the knowledge graph construction model or the knowledge graph construction method of an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0036] It should be noted that the terms "first", "second", "sample", "target", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0037] In addition, it should be noted that in the technical solution of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of relevant data such as target multimodal samples and multimodal target data are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0038] Due to the complexity and high dynamism of agricultural multimodal data, existing technologies face significant challenges in processing and integrating this data to construct effective continuous knowledge graphs. Existing technologies have the following problems: (1) In dynamic real-world scenarios, existing models have difficulty effectively processing newly emerging entities and relationships; (2) Existing knowledge graph construction mainly focuses on entity and relationship extraction from text data, while ignoring other multimodal resources; (3) In a continuous learning environment, existing technologies have difficulty effectively preventing the forgetting of old knowledge when learning new knowledge, leading to the so-called catastrophic forgetting phenomenon. (4) There are biases and inconsistencies in the fusion process of different modal data such as text and images, which leads to a decline in the performance of the model during continuous learning.
[0039] Example 1
[0040] Figure 1This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the first embodiment of the present invention. This embodiment is applicable to the case of how to construct a knowledge graph, and is particularly applicable to the case of how to construct an agricultural-related knowledge graph in the agricultural field. This method can be executed by a training device for a knowledge graph construction model, which can be implemented in the form of hardware and / or software and can be integrated into an electronic device that carries the training function of the knowledge graph construction model, such as a server. Figure 1 As shown, the method includes:
[0041] S110: Obtain target multimodal samples.
[0042] In this embodiment, the target multimodal samples refer to various types of data in the agricultural field, which may include sample images and sample texts; wherein, the sample images refer to image data in the agricultural field; and the sample texts refer to text data in the agricultural field.
[0043] Specifically, the target multimodal samples can be obtained from various channels, such as from the Internet.
[0044] S120: Extract features of the target multimodal sample to obtain sample graphic features of the target multimodal sample.
[0045] In this embodiment, the sample image-text features refer to the features of the target multimodal sample, including the visual features of the sample image and the text features of the sample text, and can be represented in the form of vectors or matrices.
[0046] As an optional method, the knowledge graph construction model may include a feature extraction sub-model, which may be composed of a convolutional neural network. Specifically, the feature extraction sub-model may be used to extract features from the sample image and sample text of the target multimodal sample respectively to obtain sample image features and sample text features, and then the sample image features and sample text features are spliced to obtain the sample image and text features of the target multimodal sample.
[0047] In another optional manner, the knowledge graph construction model includes a visual feature extraction sub-model and a text feature extraction sub-model; wherein the visual feature extraction sub-model is used to extract features of sample images, such as a ViT (Visual Transformer) model; and the text feature extraction sub-model is used to extract features of sample text, such as a BERT (Bidirectional Encoder Representation from Transformers) model. Accordingly, feature extraction is performed on the target multimodal sample to obtain sample image and text features of the target multimodal sample, including: using the visual feature extraction sub-model to perform visual feature extraction on the sample image to obtain sample image features; using the text feature extraction sub-model to perform text feature extraction on the sample text to obtain sample text features; performing attention operations on the sample image features and the sample text features to obtain sample image and text features of the target multimodal sample.
[0048] The sample image features refer to the features obtained after feature extraction of the sample image, which can be expressed in the form of a matrix or vector. The sample text features refer to the features obtained after feature extraction of the sample text, which can be expressed in the form of a matrix or vector.
[0049] It is understandable that due to the different forms of expression of images and texts, feature extraction of sample images and sample texts through the visual feature extraction sub-model and the text feature extraction sub-model respectively can more accurately extract the feature information of sample images and sample texts, thereby providing support for the training of the knowledge graph construction model.
[0050] S130. Train the knowledge graph construction model based on sample image and text features, entity labels, and relationship labels.
[0051] In this embodiment, entity labels refer to category labels of real agricultural entities, such as names, locations, organizations, etc. Relationship labels refer to labels of real relationships between different entities, such as located in, works at, belongs to, etc.
[0052] Specifically, the sample graph and text features can be input into the knowledge graph construction model, and after model prediction, the predicted entities and predicted relationships are obtained. Then, based on the predicted entities, predicted relationships, entity labels and relationship labels, the model training loss is determined. For example, based on a preset loss function such as the cross-entropy loss function, the entity prediction loss can be determined based on the predicted entities and entity labels; based on a preset loss function such as the cross-entropy loss function, the relationship prediction loss can be determined based on the predicted relationships and relationship labels; and then the entity prediction loss and the relationship prediction loss are added together to obtain the model training loss. The model training loss is used to iteratively train the knowledge graph construction model until the training stop condition is met and the model training is stopped. It should be noted that the training stop condition can be that the number of iterations meets the set number, or that the model training loss is stable within a set range; wherein, the set number and set range can be set by those skilled in the art according to the actual scenario requirements.
[0053] It should also be noted that entity prediction loss refers to the loss determined based on the predicted entity and entity label. Relationship prediction loss refers to the loss determined based on the predicted relationship and relationship label. Model training loss refers to the loss incurred in training the knowledge graph construction model. Predicted entities refer to the entities predicted by the knowledge graph construction model. Predicted relationships refer to the relationships between different entities predicted by the knowledge graph construction model, including two entities and the connection between them.
[0054] The technical solution of an embodiment of the present invention obtains target multimodal samples, including sample images and sample text, and then extracts features from the target multimodal samples to obtain sample image and text features. The knowledge graph construction model is then trained based on the sample image and text features, entity labels, and relationship labels. This technical solution can effectively process and integrate multiple data types in the agricultural field, such as image and text data, to construct a dynamic and comprehensive knowledge graph, enabling the knowledge graph to support agricultural decision-making and improve the accuracy and efficiency of agricultural management.
[0055] On the basis of the above embodiments, as an optional method of the present invention, the knowledge graph construction model training process also includes: determining the image contribution of sample image features to the prediction results; wherein the prediction results include predicted entities and predicted relationships; determining the text contribution of sample text features to the prediction results; determining the image-text contribution ratio based on the image contribution and the text contribution; determining the modulation coefficient based on the image-text contribution ratio and the hyperparameters of the gradient modulation; and using the modulation coefficient to modulate the gradient in the knowledge graph construction model training process.
[0056] Image contribution refers to the contribution of sample image features to the model's ability to correctly identify entities or extract relationships. Text contribution refers to the contribution of sample text features to the model's ability to correctly identify entities or extract relationships. The image-text contribution ratio refers to the ratio of the contribution of sample text features to sample image features. The modulation coefficient is used to modulate the gradient during model training.
[0057] Specifically, the image contribution of the sample image feature to the prediction result and the text contribution of the sample text feature to the prediction result can be determined based on the following formula. The image-text contribution ratio can be determined based on the image contribution and the text contribution:
[0058]
[0059]
[0060]
[0061] in, They represent the predicted probabilities of the sample image features and sample text features of the i-th target multimodal sample for the prediction results (entity labels or relationship labels), M represents the number of possible entity or label categories output by the knowledge graph construction model, that is, the number of entity categories, or the number of relationship categories; y represents the category label of the predicted entity or predicted relationship, y i is the true entity label or relation label of the target multimodal sample; is an indicator function, if y is equal to y i The result is 1, otherwise it is 0, which is used to select the predicted probability that matches the actual entity label or relationship label; the softmax function is used to normalize a vector into a probability distribution. The output of the visual feature extraction sub-model is converted into a probability distribution, where is the output function of the visual feature extraction sub-model, is the visual feature extraction sub-model weight matrix, θ v are the model parameters of the visual feature extraction sub-model, is the sample image of the i-th target multimodal sample. Similarly, The output of the text feature extraction sub-model is converted into a probability distribution, where is the output function of the text feature extraction sub-model, is the weight matrix of the text feature extraction sub-model, θ t are the model parameters or learning parameters of the text feature extraction sub-model, is a sample text. It represents the contribution ratio of sample text features to sample image features calculated on batch Bn. Bn is a random mini-batch of samples selected in step n during model training.
[0062] After that, the modulation coefficient can be determined based on the following formula according to the image-text contribution ratio and the hyperparameters of gradient modulation:
[0063]
[0064]
[0065] in, is the adjustment coefficient; α represents the hyperparameter affected by the adjustment modulation, i.e., the hyperparameter of the gradient modulation; G t(k) is the average modulation coefficient of the model trained after task k.
[0066] Furthermore, a modulation coefficient or an average modulation coefficient is used to modulate the gradient during the training process of the knowledge graph construction model.
[0067] It is understandable that the problem of unbalanced learning rhythm of the model can be solved by adaptively adjusting the gradients of the visual and textual modalities by evaluating their respective contributions to the learning objective.
[0068] Example 2
[0069] Figure 2 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the second embodiment of the present invention. In this embodiment, based on the above embodiment, the knowledge graph construction model includes an entity recognition sub-model and a relationship extraction sub-model; wherein the entity recognition sub-model includes a first teacher model and a first student model; the relationship extraction sub-model includes a second teacher model and a second student model; and further optimizes "training the knowledge graph construction model based on sample graphic features, entity labels and relationship labels" to provide an optional implementation plan. Figure 2 As shown, the training method of the knowledge graph construction model of this embodiment may include:
[0070] S210: Obtain target multimodal samples.
[0071] Among them, the target multimodal samples include sample images and sample texts.
[0072] S220: Extract features of the target multimodal sample to obtain sample graphic features of the target multimodal sample.
[0073] S230. Use the entity recognition sub-model to perform entity recognition on the sample image and text features to obtain the first predicted entity output by the first teacher model, the second predicted entity output by the first student model, the first image and text features output by the first teacher feature extraction layer in the first teacher model, and the second image and text features output by the first student feature extraction layer in the first student model.
[0074] In this embodiment, the entity recognition sub-model refers to a branch model for entity recognition. Optionally, the entity recognition sub-model includes a first teacher model and a first student model. The first teacher model is a pre-trained neural network model with a large network structure and parameters; the first student model is a lightweight model with a small network structure and parameters, and the first student model is also composed of a neural network; the first teacher model includes a first teacher feature extraction layer and a first entity recognition layer, wherein the first entity recognition layer can be a fully connected layer for outputting entity prediction results; the first teacher feature extraction layer can be at least one convolutional layer for further feature extraction of sample image and text features; the first student model includes a first student feature extraction layer and a second entity recognition layer, wherein the first student feature extraction layer can be at least one convolutional layer for further feature extraction of sample image and text features; the second entity recognition layer can be a fully connected layer for outputting entity prediction results.
[0075] The first predicted entity refers to the entity in the target multimodal sample predicted by the first teacher model. The second predicted entity refers to the entity in the target multimodal sample predicted by the first student model. The first graphic feature refers to the feature output by the first teacher feature extraction layer in the first teacher model, which can be represented in matrix or vector form. The second graphic feature refers to the feature output by the first student feature extraction layer in the first student model, which can be represented in matrix or vector form.
[0076] Specifically, the sample graphic features can be input into the first teacher model and the first student model of the entity recognition sub-model respectively to obtain the first graphic features output by the first teacher feature extraction layer in the first teacher model, the first predicted entity output by the first entity recognition layer in the first teacher model, the second graphic features output by the first student feature extraction layer in the first student model, and the second predicted entity output by the second entity recognition layer in the first student model.
[0077] S240: Determine entity training loss based on the first predicted entity, the second predicted entity, the first image-text feature, the second image-text feature, and the entity label.
[0078] In this embodiment, the entity training loss refers to the loss used to guide the optimization of the entity recognition sub-model parameters during the training process.
[0079] An optional method is to determine the first entity prediction loss based on a preset loss function such as a cross-entropy loss function according to the first predicted entity and the entity label; determine the second entity prediction loss based on the second predicted entity and the entity label based on a preset loss function such as a cross-entropy loss function; based on the first image and text features and the second image and text features, for example, the similarity between the first image and text features and the second image and text features can be calculated, and the similarity can be used as the image and text feature loss; and then the first entity prediction loss, the second entity prediction loss and the image and text feature loss are weightedly summed to obtain the entity training loss.
[0080] S250. Use the relationship extraction sub-model to perform relationship extraction on the sample image and text features to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0081] In this embodiment, the relationship extraction sub-model is a branch model for entity relationship extraction. Optionally, the relationship extraction sub-model includes a second teacher model and a second student model. The second teacher model is a pre-trained neural network model with a large network structure and parameters; the second student model is a lightweight model with a small network structure and parameters, and the second student model also consists of a neural network; the second teacher model includes a second teacher feature extraction layer and a first relationship extraction layer, wherein the second teacher feature extraction layer can be at least one convolutional layer for further feature extraction of sample image and text features; the first relationship extraction layer can be a fully connected layer for outputting relationship prediction results; the second student model includes a second student feature extraction layer and a second relationship extraction layer; wherein the second student feature extraction layer can be at least one convolutional layer for further feature extraction of sample image and text features; the second relationship extraction layer can be a fully connected layer for outputting relationship prediction results.
[0082] The first predicted relationship refers to the entity relationship in the target multimodal sample predicted by the second teacher model. The second predicted relationship refers to the entity relationship in the target multimodal sample predicted by the second student model. The first relationship representation refers to the features output by the second feature extraction layer in the second teacher model, which can be represented in matrix or vector form. The second relationship representation refers to the features output by the second student feature extraction layer in the second student model, which can be represented in matrix or vector form.
[0083] Specifically, the sample image and text features can be output to the second teacher model and the second student model of the relationship extraction sub-model respectively to obtain the first relationship representation output by the second teacher feature extraction layer in the second teacher model and the first predicted relationship output by the first relationship extraction layer in the second teacher model, as well as the second relationship representation output by the second student feature extraction layer in the second student model and the second predicted entity output by the second relationship extraction layer in the second student model.
[0084] S260: Determine the relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label.
[0085] In this embodiment, the target cluster refers to the cluster described by the target multimodal sample; each cluster represents an entity relationship; it should be noted that the cluster is obtained by clustering at least two multimodal samples. The so-called relationship training loss refers to the loss used to guide the optimization of the relationship extraction sub-model parameters during the training process.
[0086] An optional method is to determine the relationship prediction loss based on a preset loss function, such as a cross-entropy loss function, according to the second predicted relationship and the relationship label; then, based on the first relationship representation, the second relationship representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs, for example, determine the similarity between the first relationship representation, the second relationship representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs, and use the similarity as the relationship representation loss, and then perform weighted summation on the relationship prediction loss and the relationship representation loss to obtain the relationship training loss.
[0087] S270. Determine the model training loss based on the entity training loss and the relationship training loss, and use the model training loss to train the knowledge graph construction model.
[0088] Specifically, the entity training loss and the relationship training loss can be weighted and summed to obtain the model training loss, and the model training loss can be used to train the knowledge graph construction model.
[0089] The technical solution provided by the embodiment of the present invention can improve the accuracy of entity recognition and relationship extraction by adopting an entity recognition sub-model and a relationship extraction sub-model to perform entity recognition and relationship extraction on the target multimodal samples respectively; further adopting a student model and a pre-trained teacher model to perform entity recognition and relationship extraction, so that the complex teacher model can better guide the student model to perform entity recognition and relationship extraction, thereby further improving the accuracy of entity recognition and relationship extraction.
[0090] On the basis of the above embodiments, as an optional method of the present invention, the entity training loss can also be determined based on the first predicted entity, the second predicted entity, the first graphic feature, the second graphic feature and the entity label, including: determining the prediction difference loss based on the first predicted entity and the second predicted entity; determining the entity prediction loss based on the second predicted entity and the entity label; determining the first feature distillation loss based on the first graphic feature and the second graphic feature; determining the entity training loss based on the prediction difference loss, the entity prediction loss and the first feature distillation loss.
[0091] The prediction difference loss measures the difference between the entity prediction results output by the second teacher model and the second student model. The entity prediction loss guides parameter optimization in the second student model. The second feature distillation loss distills the first and second image-text features to reduce model forgetting.
[0092] Specifically, the KL divergence (Kullback–Leibler divergence, KLD for short) between the first predicted entity and the second predicted entity can be calculated as the prediction difference loss. Then, based on a preset loss function such as the cross-entropy loss function, the entity prediction loss is determined according to the second predicted entity and the entity label. Then, the similarity between the first image and text features and the second image and text features is used as the first feature distillation loss. Finally, the prediction difference loss, entity prediction loss and first feature distillation loss are weighted and summed to obtain the entity training loss.
[0093] It can be understood that introducing prediction difference loss and minimizing this loss during the training process can make the prediction results of the student model as close as possible to the prediction results of the teacher model, thereby helping the student model retain the old knowledge in the teacher model. Introducing feature distillation loss can reduce the loss of forgetting during model training. Introducing entity prediction loss can improve the accuracy of entity prediction of the student model. These three losses are combined to determine the model training loss to improve the accuracy of knowledge graph construction model training.
[0094] Example 3
[0095] Figure 3 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the third embodiment of the present invention. Based on the above embodiment, this embodiment further optimizes "using a relation extraction sub-model to extract relations from sample image and text features to obtain the first predicted relation output by the second teacher model, the second predicted relation output by the second student model, the first relation representation output by the second teacher feature extraction layer in the second teacher model, and the second relation representation output by the second student feature extraction layer in the second student model" and provides an optional solution, such as Figure 3As shown, the knowledge graph construction method of this embodiment may include:
[0096] S310: Obtain target multimodal samples.
[0097] Among them, the target multimodal samples include sample images and sample texts.
[0098] S320: Extract features of the target multimodal sample to obtain sample graphic features of the target multimodal sample.
[0099] S330. Use the entity recognition sub-model to perform entity recognition on the sample image and text features to obtain the first predicted entity output by the first teacher model, the second predicted entity output by the first student model, the first image and text features output by the first teacher feature extraction layer in the first teacher model, and the second image and text features output by the first student feature extraction layer in the first student model.
[0100] S340: Determine entity training loss based on the first predicted entity, the second predicted entity, the first image-text feature, the second image-text feature, and the entity label.
[0101] S350. Use the relationship extraction sub-model to perform relationship extraction on the sample image and text features to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0102] In this implementation, when using the relationship extraction sub-model to extract relationships, the entity relationship representation must first be determined based on the sample graph and text features, and then the relationship extraction sub-model must be used to perform relationship prediction on the entity relationship representation. Specifically, the following can be done: cluster at least two multimodal samples to obtain at least one cluster; each cluster represents an entity relationship; for the target cluster to which the target multimodal sample belongs, determine the distance between each multimodal sample point in the target cluster and the target cluster center in the target cluster based on the sample graph and text features of each multimodal sample point in the target cluster; based on the distance, select a first number of candidate multimodal sample points from the target cluster; fuse the sample graph and text features of the first number of candidate multimodal sample points to obtain the entity relationship representation of the target multimodal sample; use the relationship extraction sub-model to perform relationship extraction on the entity relationship representation to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0103] The entity relationship representation refers to the representation used to characterize the relationship between entities, which can be expressed in the form of a matrix or a vector. The candidate multimodal sample points refer to the multimodal sample points in the target cluster that are close to the cluster center.
[0104] Specifically, at least two multimodal samples are clustered to obtain at least one cluster, each cluster representing an entity relationship. Afterwards, for each multimodal sample point in the target cluster described by the target multimodal sample, the distance between the multimodal sample point and the cluster center is determined based on the sample graph and text features corresponding to the multimodal sample point and the sample graph and text features corresponding to the target cluster midpoint. Based on the distance, a first number of candidate multimodal sample points close to the center point are selected from the target cluster, and then the sample graph and text features of the first number of candidate multimodal sample points are fused, for example, the sample graph and text features of the first number of candidate multimodal sample points are averaged to obtain the entity relationship representation of the target multimodal sample. Finally, a relationship extraction sub-model is used to perform relationship extraction on the entity relationship representation to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0105] It can be understood that by representing the relationship based on the sample graph and text features and then using the relationship extraction sub-model to predict the relationship, the relationship extraction is made more accurate.
[0106] S360: Determine the relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label.
[0107] S370. Determine the model training loss based on the entity training loss and the relationship training loss, and use the model training loss to train the knowledge graph construction model.
[0108] The technical solution provided by the embodiment of the present invention can improve the accuracy of entity recognition and relationship extraction by adopting an entity recognition sub-model and a relationship extraction sub-model to perform entity recognition and relationship extraction on the target multimodal samples respectively; further adopting a student model and a pre-trained teacher model to perform entity recognition and relationship extraction, so that the complex teacher model can better guide the student model to perform entity recognition and relationship extraction, thereby further improving the accuracy of entity recognition and relationship extraction.
[0109] On the basis of the above embodiments, as an optional method of the present invention, data enrichment can also be performed on the target multimodal sample. Specifically, it can be: constructing a pseudo multimodal sample of the target multimodal sample based on the entity relationship representation of the target multimodal sample, the sample graphic features corresponding to the target cluster, and Gaussian noise.
[0110] Specifically, we first determine the root of the diagonal covariance of the sample graph and text features corresponding to the target cluster, that is, the sample graph and text features of all multimodal sample points in the target cluster. Then, we multiply the root of the diagonal covariance by Gaussian noise, and add the multiplication result to the entity relationship representation to obtain a pseudo-multimodal sample for constructing the target multimodal sample. This makes the generated pseudo-multimodal sample closer to the real sample, that is, the target multimodal sample.
[0111] It should be noted that the diagonal covariance consists of the variance of each feature dimension, which is used to describe the difference of multimodal samples in each dimension in the corresponding relationship of the target cluster.
[0112] It can be understood that by constructing pseudo multimodal samples of target multimodal samples, sample diversity can be increased.
[0113] Example 4
[0114] Figure 4 This is a flowchart of a training method for a knowledge graph construction model provided in accordance with the fourth embodiment of the present invention. Based on the above embodiment, this embodiment further optimizes "determining the relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label" and provides an optional implementation plan. Figure 4 As shown, the training method of the knowledge graph construction model of this embodiment may include:
[0115] S410: Obtain target multimodal samples.
[0116] Among them, the target multimodal samples include sample images and sample texts.
[0117] S420: Extract features of the target multimodal sample to obtain sample graphic features of the target multimodal sample.
[0118] S430. Use the entity recognition sub-model to perform entity recognition on the sample image and text features to obtain the first predicted entity output by the first teacher model, the second predicted entity output by the first student model, the first image and text features output by the first teacher feature extraction layer in the first teacher model, and the second image and text features output by the first student feature extraction layer in the first student model.
[0119] S440: Determine an entity training loss based on the first predicted entity, the second predicted entity, the first image-text feature, the second image-text feature, and the entity label.
[0120] S450. Use the relationship extraction sub-model to perform relationship extraction on the sample image and text features to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0121] S460: Determine the relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label.
[0122] An optional method is to determine the relationship prediction loss based on the second relationship representation and the relationship label; determine the contrastive distillation loss based on the first relationship representation, the second relationship representation, and the sample image and text features corresponding to the target cluster to which the target multimodal sample belongs; and determine the relationship training loss based on the relationship prediction loss and the contrastive distillation loss.
[0123] The relationship prediction loss can be determined based on a preset loss function, such as a cross-entropy loss function, according to the second predicted relationship and the relationship label; then, based on the first relationship representation, the second relationship representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs, for example, the similarity between the first relationship representation, the second relationship representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs is determined, and the similarity is used as the comparative distillation loss, and then the relationship prediction loss and the comparative distillation loss are weightedly summed to obtain the relationship training loss.
[0124] Furthermore, the comparative distillation loss is determined based on the first relational representation, the second relational representation, and the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, including: determining the second feature distillation loss based on the first relational representation and the second relational representation; determining the other multimodal samples closest to the target multimodal sample from the other multimodal samples that belong to the same relationship as the target multimodal sample based on the first relational representation, and using the first relational representation of the other multimodal samples as the maximum relational representation; determining the other multimodal samples farthest from the target multimodal sample from the other multimodal samples that do not belong to the same relationship as the target multimodal sample based on the first relational representation, and using the first relational representation of the other multimodal samples as the minimum relational representation; determining the distillation triplet loss based on the second relational representation, the maximum relational representation, the minimum relational representation, and the sample graphic features corresponding to the target cluster to which the target multimodal sample belongs; and determining the comparative distillation loss based on the second feature distillation loss and the distillation triplet loss.
[0125] Specifically, the second characteristic distillation loss may be determined based on the following formula according to the first relationship expression and the second relationship expression:
[0126]
[0127] Among them, L rd represents the second characteristic distillation loss, represents the first relation, represents the second relational representation, x represents the target multimodal sample, Indicates the sample image and text features corresponding to the target cluster to which the target multimodal sample belongs.
[0128] Then, according to the first relational representation, the multimodal sample closest to the target multimodal sample is determined from the other multimodal samples that have the same relation with the target multimodal sample, and the first relational representation of the other multimodal sample is used as the maximum relational representation, which is recorded as According to the first relational representation, from other multimodal samples that do not belong to the same relation as the target multimodal sample, determine the other multimodal sample that is farthest from the target multimodal sample, and use the first relational representation of the other multimodal sample as the minimum relational representation, which is recorded as
[0129] Furthermore, the distillation triplet loss L can be determined based on the following formula according to the second relation representation, the maximum relation representation, the minimum relation representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs: dtr :
[0130]
[0131] Finally, the second characteristic distillation loss and the distillation triplet loss are added together to obtain the comparative distillation loss.
[0132] It can be understood that by comparing the first relation representation with the second relation representation and introducing knowledge distillation, the relation extraction accuracy of the relation extraction sub-model can be improved.
[0133] S470. Determine the model training loss based on the entity training loss and the relationship training loss, and use the model training loss to train the knowledge graph construction model.
[0134] The technical solution provided by the embodiment of the present invention can improve the accuracy of entity recognition and relationship extraction by adopting an entity recognition sub-model and a relationship extraction sub-model to perform entity recognition and relationship extraction on the target multimodal samples respectively; further adopting a student model and a pre-trained teacher model to perform entity recognition and relationship extraction, so that the complex teacher model can better guide the student model to perform entity recognition and relationship extraction, thereby further improving the accuracy of entity recognition and relationship extraction.
[0135] Example 5
[0136] Figure 5 This is a flow chart of a knowledge graph construction method provided according to the fifth embodiment of the present invention. This embodiment is applicable to the situation of how to construct a knowledge graph, and is particularly applicable to the situation of how to construct an agricultural-related knowledge graph in the agricultural field. The method can be executed by a knowledge graph construction device, which can be implemented in the form of hardware and / or software and can be integrated into an electronic device that carries the knowledge graph construction function, such as a server. Figure 5 As shown, the knowledge graph construction method of this embodiment may include:
[0137] S510: Acquire multimodal target data.
[0138] In this embodiment, multimodal target data refers to multimodal data that requires knowledge graph construction, that is, data that requires knowledge graph construction for some new agricultural multimodal data; including target images and target texts; among which, target images refer to newly added agricultural image data; target texts refer to newly added agricultural text data.
[0139] Specifically, based on business needs, multimodal target data in business scenarios can be obtained.
[0140] S520. Use the knowledge graph to build a model to predict entities and relationships of the multimodal target data, and obtain a target knowledge graph of the multimodal target data.
[0141] In this embodiment, the knowledge graph construction model refers to a model for predicting entities and relationships based on multimodal target data, and can be trained using the knowledge graph construction model training method provided in any embodiment of the present invention. The target knowledge graph refers to a knowledge graph determined based on the multimodal target data, including target entities and entity relationships between different entities.
[0142] Specifically, the multimodal target data can be input into the knowledge graph construction model. After model prediction, the target entity of the multimodal target data and the entity relationship between different target entities can be obtained, thereby obtaining the target knowledge graph.
[0143] The technical solution provided by an embodiment of the present invention obtains multimodal target data, including target images and target text, and then uses a knowledge graph construction model to predict entities and relationships in the multimodal target data, thereby obtaining a target knowledge graph for the multimodal target data. This technical solution uses a pre-trained knowledge graph construction model to enrich the knowledge graph of the newly added multimodal data. This not only allows for the extraction of relationships between entities, but also enables continuous updating of the knowledge graph, thereby supporting more accurate agricultural monitoring and decision-making.
[0144] Example 6
[0145] Figure 6 This is a structural diagram of a training device for a knowledge graph construction model provided according to the sixth embodiment of the present invention. This embodiment is applicable to the case of how to construct a knowledge graph, and is particularly applicable to the case of how to construct an agricultural-related knowledge graph in the agricultural field. The device can be implemented in the form of hardware and / or software and can be integrated into an electronic device that carries the training function of the knowledge graph construction model, such as a server. Figure 6 As shown, the training device for the knowledge graph construction model of this embodiment may include:
[0146] The target sample acquisition module 610 is used to acquire a target multimodal sample; the target multimodal sample includes a sample image and a sample text;
[0147] The sample image and text feature determination module 620 is used to extract features of the target multimodal sample to obtain the sample image and text features of the target multimodal sample;
[0148] The model training module 630 is used to train the knowledge graph construction model based on sample image and text features, entity labels and relationship labels.
[0149] The technical solution of an embodiment of the present invention obtains target multimodal samples, including sample images and sample text, and then extracts features from the target multimodal samples to obtain sample image and text features. The knowledge graph construction model is then trained based on the sample image and text features, entity labels, and relationship labels. This technical solution can effectively process and integrate multiple data types in the agricultural field, such as image and text data, to construct a dynamic and comprehensive knowledge graph, enabling the knowledge graph to support agricultural decision-making and improve the accuracy and efficiency of agricultural management.
[0150] Optionally, the knowledge graph construction model includes a visual feature extraction sub-model and a text feature extraction sub-model; accordingly, the sample image and text feature determination module 620 is specifically used to:
[0151] Using the visual feature extraction sub-model, visual features of the sample image are extracted to obtain the sample image features;
[0152] Using the text feature extraction sub-model, extract text features from the sample text to obtain sample text features;
[0153] Attention operation is performed on the sample image features and sample text features to obtain the sample image and text features of the target multimodal sample.
[0154] Optionally, the knowledge graph construction model includes an entity recognition sub-model and a relationship extraction sub-model; wherein the entity recognition sub-model includes a first teacher model and a first student model; and the relationship extraction sub-model includes a second teacher model and a second student model. Accordingly, the model training module 630 includes:
[0155] an entity prediction information determination unit, configured to perform entity recognition on the sample image and text features using an entity recognition sub-model, to obtain a first predicted entity output by the first teacher model, a second predicted entity output by the first student model, a first image and text feature output by the first teacher feature extraction layer in the first teacher model, and a second image and text feature output by the first student feature extraction layer in the first student model;
[0156] An entity training loss determination unit, configured to determine an entity training loss based on the first predicted entity, the second predicted entity, the first image-text feature, the second image-text feature, and the entity label;
[0157] a relationship prediction information determination unit, configured to perform relationship extraction on sample image and text features using a relationship extraction sub-model, to obtain a first predicted relationship output by the second teacher model, a second predicted relationship output by the second student model, a first relationship representation output by the second teacher feature extraction layer in the second teacher model, and a second relationship representation output by the second student feature extraction layer in the second student model;
[0158] a relationship training loss determination unit, configured to determine a relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label;
[0159] The model training unit is used to determine the model training loss based on the entity training loss and the relationship training loss, and use the model training loss to train the knowledge graph construction model.
[0160] Optionally, the entity training loss determination unit is specifically used to:
[0161] determining a prediction difference loss based on the first prediction entity and the second prediction entity;
[0162] Determining an entity prediction loss based on the second predicted entity and the entity label;
[0163] determining a first characteristic distillation loss based on the first graphic feature and the second graphic feature;
[0164] The entity training loss is determined based on the prediction difference loss, entity prediction loss, and first feature distillation loss.
[0165] Optionally, the relationship prediction information determining unit is specifically configured to:
[0166] Clustering at least two multimodal samples to obtain at least one cluster; each cluster represents an entity relationship;
[0167] For the target cluster to which the target multimodal sample belongs, the distance between each multimodal sample point in the target cluster and the target cluster center in the target cluster is determined according to the sample graph and text features of each multimodal sample point in the target cluster;
[0168] Selecting a first number of candidate multimodal sample points from the target cluster according to the distance;
[0169] Fusing the sample image and text features of the first number of candidate multimodal sample points to obtain an entity relationship representation of the target multimodal sample;
[0170] A relationship extraction sub-model is used to perform relationship extraction on the entity relationship representation to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
[0171] Optionally, the model training module 630 further includes:
[0172] The pseudo sample determination unit is used to construct a pseudo multimodal sample of the target multimodal sample based on the entity relationship representation of the target multimodal sample, the sample image and text features corresponding to the target cluster, and Gaussian noise.
[0173] Optional, relational training loss determination unit, including:
[0174] a relation prediction loss determination subunit, configured to determine a relation prediction loss based on the second relation representation and the relation label;
[0175] a contrastive distillation loss determination subunit, configured to determine the contrastive distillation loss based on the first relational representation, the second relational representation, and sample image and text features corresponding to the target cluster to which the target multimodal sample belongs;
[0176] The relation training loss determination subunit is used to determine the relation training loss based on the relation prediction loss and the contrastive distillation loss.
[0177] Optionally, the comparative distillation loss determination subunit is specifically used to:
[0178] determining a second characteristic distillation loss based on the first relationship representation and the second relationship representation;
[0179] Determine, based on the first relational representation, from other multimodal samples that have the same relation as the target multimodal sample, the other multimodal samples that are closest to the target multimodal sample, and use the first relational representation of the other multimodal samples as the maximum relational representation;
[0180] Determine, based on the first relational representation, from other multimodal samples that do not have the same relation as the target multimodal sample, the other multimodal samples that are farthest away from the target multimodal sample, and use the first relational representation of the other multimodal samples as the minimum relational representation;
[0181] Determine the distillation triplet loss based on the second relation representation, the maximum relation representation, the minimum relation representation, and the sample graph and text features corresponding to the target cluster to which the target multimodal sample belongs;
[0182] Based on the second characteristic distillation loss and the distillation triplet loss, the comparative distillation loss is determined.
[0183] Optionally, the device further includes a gradient modulation module, configured to:
[0184] Determine the image contribution of sample image features to the prediction results during the knowledge graph construction model training process; wherein the prediction results include predicted entities and predicted relationships;
[0185] Determine the contribution of sample text features to the prediction results;
[0186] Determine the image-text contribution ratio based on the image contribution and text contribution;
[0187] Determine the modulation coefficient based on the image-text contribution ratio and the hyperparameters of gradient modulation;
[0188] The modulation coefficient is used to modulate the gradient during the training process of the knowledge graph construction model.
[0189] The training device for the knowledge graph construction model provided in an embodiment of the present invention can execute the training method for the knowledge graph construction model provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0190] Example 7
[0191] Figure 7This is a schematic diagram of the structure of a knowledge graph construction device provided according to the seventh embodiment of the present invention. This embodiment is applicable to the situation of how to construct a knowledge graph, and is particularly applicable to the situation of how to construct an agricultural-related knowledge graph in the agricultural field. The device can be implemented in the form of hardware and / or software and can be integrated into an electronic device that carries the knowledge graph construction function, such as a server. Figure 7 As shown, the knowledge graph construction device of this embodiment may include:
[0192] The target data acquisition module 710 is used to acquire multimodal target data; the multimodal target data includes a target image and a target text;
[0193] The knowledge graph determination module 720 is used to use the knowledge graph construction model to predict entities and relationships of multimodal target data to obtain a target knowledge graph of the multimodal target data; wherein, the knowledge graph construction model is trained according to the training method of the knowledge graph construction model of any embodiment of the present invention.
[0194] The technical solution provided by an embodiment of the present invention obtains multimodal target data, including target images and target text, and then uses a knowledge graph construction model to predict entities and relationships in the multimodal target data, thereby obtaining a target knowledge graph for the multimodal target data. This technical solution uses a pre-trained knowledge graph construction model to enrich the knowledge graph of the newly added multimodal data. This not only allows for the extraction of relationships between entities, but also enables continuous updating of the knowledge graph, thereby supporting more accurate agricultural monitoring and decision-making.
[0195] The knowledge graph construction device provided by the embodiment of the present invention can execute the knowledge graph construction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0196] Example 8
[0197] Figure 8 It is a structural diagram of an electronic device that implements the training method of the knowledge graph construction model or the knowledge graph construction method of an embodiment of the present invention. Figure 8A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0198] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0199] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0200] The processor 11 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the training method of the knowledge graph construction model or the knowledge graph construction method.
[0201] In some embodiments, the training method of the knowledge graph construction model or the knowledge graph construction method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method of the knowledge graph construction model or the knowledge graph construction method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the training method of the knowledge graph construction model or the knowledge graph construction method by any other appropriate means (for example, by means of firmware).
[0202] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0203] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0204] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0205] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0206] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0207] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0208] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0209] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A training method for a knowledge graph construction model, characterized in that: include: Acquire a target multimodal sample; the target multimodal sample includes a sample image and a sample text; Performing feature extraction on the target multimodal sample to obtain sample graphic features of the target multimodal sample; The knowledge graph construction model is trained based on the sample image and text features, entity labels, and relationship labels; wherein the knowledge graph construction model includes an entity recognition sub-model and a relationship extraction sub-model; the entity recognition sub-model includes a first teacher model and a first student model; and the relationship extraction sub-model includes a second teacher model and a second student model; The step of training the knowledge graph construction model based on the sample image and text features, entity labels, and relationship labels includes: Using the entity recognition sub-model, perform entity recognition on the sample graphic features to obtain a first predicted entity output by the first teacher model, a second predicted entity output by the first student model, a first graphic feature output by the first teacher feature extraction layer in the first teacher model, and a second graphic feature output by the first student feature extraction layer in the first student model; Determining an entity training loss based on the first predicted entity, the second predicted entity, the first graphic feature, the second graphic feature, and the entity label; Using the relationship extraction sub-model, perform relationship extraction on the sample image and text features to obtain a first predicted relationship output by the second teacher model, a second predicted relationship output by the second student model, a first relationship representation output by the second teacher feature extraction layer in the second teacher model, and a second relationship representation output by the second student feature extraction layer in the second student model; Determining a relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label; The model training loss is determined based on the entity training loss and the relationship training loss, and the knowledge graph construction model is trained using the model training loss.
2. The method according to claim 1, characterized in that The knowledge graph construction model includes a visual feature extraction sub-model and a text feature extraction sub-model; accordingly, feature extraction is performed on the target multimodal sample to obtain sample image and text features of the target multimodal sample, including: Using the visual feature extraction sub-model, extract visual features from the sample image to obtain sample image features; Using the text feature extraction sub-model, extracting text features from the sample text to obtain sample text features; An attention operation is performed on the sample image features and the sample text features to obtain sample image and text features of the target multimodal sample.
3. The method according to claim 1, characterized in that Determining an entity training loss according to the first predicted entity, the second predicted entity, the first image-text feature, the second image-text feature, and the entity label includes: determining a prediction difference loss based on the first prediction entity and the second prediction entity; Determining an entity prediction loss based on the second predicted entity and the entity label; determining a first characteristic distillation loss according to the first graphic feature and the second graphic feature; An entity training loss is determined based on the prediction difference loss, the entity prediction loss, and the first feature distillation loss.
4. The method according to claim 1, wherein Using the relationship extraction sub-model, performing relationship extraction on the sample image and text features to obtain a first predicted relationship output by the second teacher model, a second predicted relationship output by the second student model, a first relationship representation output by the second teacher feature extraction layer in the second teacher model, and a second relationship representation output by the second student feature extraction layer in the second student model, including: Clustering at least two multimodal samples to obtain at least one cluster; each cluster represents an entity relationship; For the target cluster to which the target multimodal sample belongs, determining the distance between each multimodal sample point in the target cluster and the target cluster center in the target cluster according to the sample graph and text features of each multimodal sample point in the target cluster; Selecting a first number of candidate multimodal sample points from the target cluster according to the distance; Fusing sample image and text features of a first number of candidate multimodal sample points to obtain an entity relationship representation of the target multimodal sample; The relationship extraction sub-model is used to perform relationship extraction on the entity relationship representation to obtain the first predicted relationship output by the second teacher model, the second predicted relationship output by the second student model, the first relationship representation output by the second teacher feature extraction layer in the second teacher model, and the second relationship representation output by the second student feature extraction layer in the second student model.
5. The method according to claim 4, characterized in that Also includes: A pseudo multimodal sample of the target multimodal sample is constructed according to the entity relationship representation of the target multimodal sample, the sample graphic features corresponding to the target cluster, and Gaussian noise.
6. The method according to claim 1, characterized in that Determining a relationship training loss according to the second predicted relationship, the first relationship representation, the second relationship representation, sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label includes: determining a relationship prediction loss based on the second relationship representation and the relationship label; determining a contrastive distillation loss based on the first relational representation, the second relational representation, and sample graphic features corresponding to the target cluster to which the target multimodal sample belongs; A relation training loss is determined according to the relation prediction loss and the contrastive distillation loss.
7. The method according to claim 6, characterized in that Determining a contrastive distillation loss according to the first relational representation, the second relational representation, and sample image and text features corresponding to the target cluster to which the target multimodal sample belongs includes: determining a second characteristic distillation loss according to the first relationship representation and the second relationship representation; Determining, based on the first relational representation, other multimodal samples that are closest to the target multimodal sample from other multimodal samples that have the same relation as the target multimodal sample, and using the first relational representation of the other multimodal sample as the maximum relational representation; Determining, based on the first relational representation, other multimodal samples that are not in the same relation as the target multimodal sample, the other multimodal samples that are farthest from the target multimodal sample, and using the first relational representation of the other multimodal samples as the minimum relational representation; Determining a distillation triplet loss according to the second relational representation, the maximum relational representation, the minimum relational representation, and sample graphic features corresponding to the target cluster to which the target multimodal sample belongs; A comparative distillation loss is determined based on the second characteristic distillation loss and the distillation triplet loss.
8. The method according to claim 2, characterized in that The knowledge graph construction model training process also includes: Determining the image contribution of the sample image feature to the prediction result; wherein the prediction result includes a predicted entity and a predicted relationship; Determining the text contribution of the sample text feature to the prediction result; Determining an image-text contribution ratio according to the image contribution and the text contribution; Determining a modulation coefficient according to the image-text contribution ratio and the hyperparameters of gradient modulation; The modulation coefficient is used to modulate the gradient during the training process of the knowledge graph construction model.
9. A knowledge graph construction method, characterized in that: include: Acquire multimodal target data; The multimodal target data includes a target image and a target text; A knowledge graph construction model is used to perform entity and relationship prediction on the multimodal target data to obtain a target knowledge graph of the multimodal target data; wherein, the knowledge graph construction model is trained according to the training method of the knowledge graph construction model described in any one of claims 1-8.
10. A training device for a knowledge graph construction model, characterized in that: include: A target sample acquisition module is used to acquire a target multimodal sample; the target multimodal sample includes a sample image and a sample text; A sample image and text feature determination module, configured to extract features from the target multimodal sample to obtain sample image and text features of the target multimodal sample; A model training module is used to train a knowledge graph construction model based on the sample image and text features, entity labels, and relationship labels; wherein the knowledge graph construction model includes an entity recognition sub-model and a relationship extraction sub-model; the entity recognition sub-model includes a first teacher model and a first student model; and the relationship extraction sub-model includes a second teacher model and a second student model; The model training module includes: an entity prediction information determination unit, configured to perform entity recognition on the sample image and text features using the entity recognition sub-model, to obtain a first predicted entity output by the first teacher model, a second predicted entity output by the first student model, a first image and text feature output by the first teacher feature extraction layer in the first teacher model, and a second image and text feature output by the first student feature extraction layer in the first student model; an entity training loss determining unit, configured to determine an entity training loss based on the first predicted entity, the second predicted entity, the first graphic feature, the second graphic feature, and the entity label; a relationship prediction information determination unit, configured to perform relationship extraction on the sample image and text features using the relationship extraction sub-model, to obtain a first predicted relationship output by the second teacher model, a second predicted relationship output by the second student model, a first relationship representation output by the second teacher feature extraction layer in the second teacher model, and a second relationship representation output by the second student feature extraction layer in the second student model; a relationship training loss determining unit, configured to determine a relationship training loss based on the second predicted relationship, the first relationship representation, the second relationship representation, sample graphic features corresponding to the target cluster to which the target multimodal sample belongs, and the relationship label; The model training unit is used to determine the model training loss based on the entity training loss and the relationship training loss, and use the model training loss to train the knowledge graph construction model.
11. A knowledge graph construction device, characterized in that: include: A target data acquisition module, used to acquire multimodal target data; The multimodal target data includes a target image and a target text; A knowledge graph determination module is used to use a knowledge graph construction model to perform entity and relationship prediction on the multimodal target data to obtain a target knowledge graph of the multimodal target data; wherein, the knowledge graph construction model is trained according to the training method of the knowledge graph construction model described in any one of claims 1-8.
12. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method for the knowledge graph construction model described in any one of claims 1-8, or the knowledge graph construction method described in claim 9.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable the processor to implement the training method for the knowledge graph construction model described in any one of claims 1 to 8, or the knowledge graph construction method described in claim 9 when executed.
14. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the training method for the knowledge graph construction model according to any one of claims 1 to 8, or the knowledge graph construction method according to claim 9.
Citation Information
Patent Citations
Data processing method and device
CN116109979A