Knowledge base construction method based on automatic extraction, classification and association of multimode entities and storage medium
By adopting the automatic extraction, classification and association methods of multimodal entities in the construction of enterprise knowledge bases, the accuracy and efficiency of entity identification and extraction in the existing technology are solved, efficient extraction and association of knowledge is achieved, and the knowledge management capabilities of enterprises are improved.
Patent Information
- Application Number
- CN202510056304.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The accuracy and efficiency of entity identification and extraction in the construction of enterprise knowledge bases is low, and the potential of knowledge graphs has not been fully utilized.
The knowledge base construction method based on automatic extraction, classification and association of multimodal entities is adopted, including preprocessing of text data and image data, entity extraction and semantic relationship acquisition, feature fusion, construction of multimodal deep learning models and introduction of attention mechanisms, as well as the construction of triplets and the optimization of knowledge bases.
It improves the accuracy and efficiency of the knowledge base, realizes effective extraction, association and integration of knowledge, and enhances the company's knowledge management level.
Smart Images

Figure CN120014327A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a knowledge base construction method and storage medium based on automatic extraction, classification and association of multi-mode entities. Background Art
[0002] Enterprise knowledge graphs accurately represent and associate knowledge resources through structured elements such as entities, attributes, and relationships, making knowledge easier to understand and use. It covers all kinds of knowledge resources inside and outside the enterprise, including employees, products, services, processes, documents, customers, markets, competitors, etc., forming a complete, accurate, and related knowledge network. This not only helps enterprises achieve comprehensive coverage of knowledge and reduce information islands and redundancy, but also promotes centralized management and classification of knowledge, making it easier for employees to query, browse, and obtain the knowledge they need. However, although the existing knowledge base construction methods have recognized the importance of enterprise knowledge graphs to a certain extent and tried to apply them to knowledge management practices, they still face many challenges and limitations and have failed to fully tap the potential of knowledge graphs.
[0003] Existing knowledge base construction methods often rely on traditional text processing technologies for entity recognition and extraction. The accuracy and efficiency of these technologies are greatly reduced when faced with complex and changeable enterprise data. Summary of the invention
[0004] In order to solve the technical problem of low accuracy of knowledge base in the prior art, the present invention provides the following technical solution:
[0005] On the one hand, a method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities is provided, the method comprising:
[0006] S1, preprocessing text data and image data respectively;
[0007] S2, extracting entities in the text data based on a bidirectional encoder and a random field model, and obtaining semantic relationships in the text data through a sentence analysis algorithm;
[0008] S3, acquiring entities in the image data based on a target detection model, and acquiring semantic relationships in the image data through an image description generation algorithm;
[0009] S4, performing feature fusion on the entities in the text data and the entities in the image data to obtain a multimodal entity;
[0010] S5. Construct a multimodal deep learning model and introduce an attention mechanism into the multimodal deep learning model; classify the multimodal entities based on the multimodal deep learning model to obtain classified multimodal entities;
[0011] S6, obtaining a parallel relationship between entities in the text data and entities in the image data, and constructing a triple based on the parallel relationship, a semantic relationship in the text data, and a semantic relationship in the image data;
[0012] S7, constructing a knowledge base based on the classified multimodal entities and the triples;
[0013] S8. Perform a quality assessment on the knowledge base, and optimize the knowledge base based on the assessment result.
[0014] As an optional embodiment of the present invention, optionally, in step S2, extracting entities in the text data based on a bidirectional encoder and a random field model, and obtaining semantic relationships in the text data through a sentence analysis algorithm includes:
[0015] S201, encoding the preprocessed text data using the bidirectional encoder to capture context information;
[0016] S202, using a random field model to perform sequence annotation on the encoded text data to identify entity boundaries;
[0017] S203, determining entity types based on the context information and entity boundaries, and extracting entities from the text data;
[0018] S204: parse the sentence structure in the text data using a dependency syntax analysis algorithm to obtain the grammatical relationship between entities in the sentence.
[0019] As an optional embodiment of the present invention, optionally, in step S3, entities in the image data are obtained based on the target detection model, and semantic relationships in the image data are obtained through an image description generation algorithm:
[0020] S301, using a target detection model to locate an object in the image data, and identify the object, and taking the object as an entity;
[0021] S302, classifying the objects by labels to obtain a semantic label for each object;
[0022] S303: Generate a descriptive sentence related to the object in the image data based on the semantic tag using an image description generation algorithm, and acquire a semantic relationship in the image data based on the descriptive sentence.
[0023] As an optional embodiment of the present invention, optionally, the expression of the multimodal entity obtained in step S4 is:
[0024]
[0025] Among them, F represents the fused multi-modal entity features, H() represents the feature fusion function, and F t The feature vector set representing the text entity, F g represents the feature vector set of image entities, ψ represents the parameter set of feature fusion function, σ() represents the activation function, n represents the number of text features, i represents the i-th text feature, α i represents the weight of the i-th text feature, f text represents the text feature extraction function, e i represents the i-th text feature, represents the parameters of the text feature extraction function, m represents the number of image features, j represents the jth image feature, β j represents the weight of the jth image feature, f image () represents the image feature extraction function, b j and c j Represents two different image features in the image, Represents the parameters of the image feature extraction function.
[0026] As an optional embodiment of the present invention, optionally, constructing a multimodal deep learning model in step S5, and introducing an attention mechanism into the multimodal deep learning model includes:
[0027] S501, extracting historical text data and historical image data, and preprocessing the historical text data and historical image data respectively;
[0028] S502: construct an initial multimodal deep learning model, input the preprocessed historical text data and historical image data into the initial multimodal deep learning model for training, and obtain a trained multimodal deep learning model;
[0029] S503, introducing an attention mechanism into the trained multimodal deep learning model, and optimizing the multimodal deep learning model;
[0030] S504: Classify the multimodal entities using the trained multimodal deep learning model.
[0031] As an optional embodiment of the present invention, optionally, the expression for classifying the multimodal entity is:
[0032]
[0033] Where C represents the set of multimodal entity labels after classification, J() represents the entity classification function, F represents the fused multimodal entity features, Ω represents the parameter set of the multimodal deep learning model, and c krepresents the category label of the kth output, l represents the total number of labels, represents the maximum probability, L represents the category set, c represents the category label, ω k Represents the classification parameters of a multimodal deep learning model.
[0034] As an optional embodiment of the present invention, optionally, obtaining the parallel relationship between the entity in the text data and the entity in the image data in step S6 includes:
[0035] S601, establishing a mapping relationship between entities in the text data and entities in the image data based on entity similarity calculation;
[0036] S602: Based on the mapping relationship, identify a parallel relationship between the entity in the text data and the entity in the image data, wherein the parallel relationship indicates that the entity in the text data and the entity in the image data are associated and independent of each other;
[0037] S603: verify the parallel relationship.
[0038] As an optional embodiment of the present invention, optionally, the expression for constructing the triple in step S6 is:
[0039] T triplet =K triplet (E text ,E image ,R text ,R image ,Ξ triplet )
[0040] ={(u x ,r,u y )∣(u x ,u y )∈U,r∈R(u x ,u y ,x r )}
[0041] Among them, T triplet represents the constructed triple set, K triplet represents the triple construction function, E text Represents a text entity set, E image represents the image entity set, R text Represents the set of semantic relations between text entities, R image represents the set of semantic relations between image entities, triplet Represents the parameter set of the triple construction function, T triplet Represents a tuple of entity x, relationship, and entity y. R() represents a set of semantic relationships. rRepresents a semantic relationship parameter.
[0042] As an optional embodiment of the present invention, optionally, in step S7, constructing a knowledge base based on the classified multimodal entities and the triples includes:
[0043] S701, constructing an initial database;
[0044] S702: Standardize the multimodal entities, uniquely identify each of the multimodal entities, and establish an index of the multimodal entities;
[0045] S703, mapping the relations and entities in the triples to the initial knowledge base;
[0046] S704, representing the relationships and entities in the triples in the form of a graph through a knowledge representation model to obtain a knowledge graph;
[0047] S705. Add attribute information to the entities in the knowledge graph to obtain a knowledge base.
[0048] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned knowledge base construction methods based on automatic extraction, classification and association of multimodal entities.
[0049] The beneficial effects brought by the technical solution provided by the embodiment of the present invention include at least: the present invention firstly pre-processes the text data and the image data, so as to improve the accuracy and efficiency of subsequent entity extraction and semantic relationship acquisition. Then, the entities in the text data are extracted by using the bidirectional encoder and the random field model, and the semantic relationship in the text data is obtained by the sentence analysis algorithm. Compared with the traditional text processing technology, this method can more accurately identify the entities and semantic relationships. At the same time, the entities in the image data are obtained by the target detection model, and the semantic relationship in the image data is obtained by using the image description generation algorithm, so that the knowledge in the image data can also be effectively extracted and utilized. Further, by performing feature fusion on the entities in the text data and the entities in the image data, multi-modal entities are obtained, which helps to associate and integrate data of different modalities. Then, a multi-modal deep learning model is constructed, and the attention mechanism is introduced into the model to classify the multi-modal entities. This classification method can more accurately identify entities of different categories. In addition, the present invention also realizes the association and integration of knowledge by obtaining the parallel relationship between the entities in the text data and the entities in the image data, and constructing triples based on these parallel relationships, the semantic relationship in the text data and the semantic relationship in the image data. Finally, a knowledge base is constructed based on the classified multi-mode entities and triples, and the relationships and entities in the triples are represented in the form of a graph through a knowledge representation model to obtain a knowledge graph, which makes knowledge easier to understand and use and improves the accuracy of the knowledge base. Therefore, the technical solution provided by the embodiment of the present invention can solve the technical problem of low accuracy of the knowledge base in the prior art, realize the effective extraction, association and integration of knowledge, and improve the knowledge management level of the enterprise. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 It is a flow chart of a knowledge base construction method based on automatic extraction, classification and association of multi-modal entities provided by an embodiment of the present invention;
[0052] Figure 2 It is a multimodal deep learning model training flow chart provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0054] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0055] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0056] In the embodiments of the present invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0057] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0058] like Figure 1 As shown, an embodiment of the present invention provides a knowledge base construction method based on automatic extraction, classification and association of multi-modal entities, the method comprising:
[0059] S1, preprocessing text data and image data respectively;
[0060] It should be noted that the preprocessing of text data includes the removal of stop words, word segmentation, part-of-speech tagging, named entity recognition and other steps to extract key information from the text. The preprocessing of image data includes image enhancement, target detection, image segmentation and other steps to identify objects and scenes in the image. Through preprocessing, noise and redundant information in the data can be removed, laying a good foundation for subsequent entity extraction and semantic relationship acquisition.
[0061] S2, extracting entities in the text data based on a bidirectional encoder and a random field model, and obtaining semantic relationships in the text data through a sentence analysis algorithm;
[0062] It should be noted that in step S2, the bidirectional encoder can encode the text data to obtain the context information of the text data, and the random field model can perform sequence annotation on the text data based on the context information, thereby extracting entities in the text data. Through the sentence analysis algorithm, the sentence structure in the text data can be further analyzed to obtain the semantic relationship in the text data. Compared with the traditional rule-based or template-based method, this method can more accurately identify entities and semantic relationships, and improve the accuracy and efficiency of entity extraction and semantic relationship acquisition.
[0063] S3, acquiring entities in the image data based on a target detection model, and acquiring semantic relationships in the image data through an image description generation algorithm;
[0064] It should be noted that in step S3, the target detection model can process the image data to identify objects and scenes in the image, which are regarded as entities in the image data. The image description generation algorithm can generate a description of the image content based on the identified entities, thereby obtaining the semantic relationship in the image data. This method can effectively extract and utilize the knowledge in the image data, providing strong support for subsequent knowledge association and integration.
[0065] S4, performing feature fusion on the entities in the text data and the entities in the image data to obtain a multimodal entity;
[0066] It should be noted that in step S4, by extracting and fusing the features of entities in text data and entities in image data, a multimodal entity containing richer information can be obtained. This feature fusion method can be implemented based on a deep learning model, for example, using a convolutional neural network (CNN) to extract the features of image entities, using a recurrent neural network (RNN) or Transformer to extract the features of text entities, and then splicing or fusing these features to obtain a feature representation of the multimodal entity. This feature fusion helps to associate and integrate data of different modalities, providing strong support for subsequent multimodal entity classification and relationship extraction.
[0067] S5. Construct a multimodal deep learning model and introduce an attention mechanism into the multimodal deep learning model; classify the multimodal entities based on the multimodal deep learning model to obtain classified multimodal entities;
[0068] It should be noted that when constructing a multimodal deep learning model in step S5, a suitable deep learning framework, such as TensorFlow or PyTorch, can be selected, and a neural network structure containing multiple levels can be built based on these frameworks. Among them, the attention mechanism can be implemented by introducing a self-attention mechanism or a multi-head attention mechanism to enhance the model's ability to capture important information in the input data. By training the multimodal deep learning model, it can learn the feature representation of multimodal entities and classify multimodal entities based on these feature representations. The classified multimodal entities can be organized into different knowledge bases according to their category labels to facilitate subsequent knowledge retrieval and utilization.
[0069] S6, obtaining a parallel relationship between entities in the text data and entities in the image data, and constructing a triple based on the parallel relationship, a semantic relationship in the text data, and a semantic relationship in the image data;
[0070] It should be noted that in step S6, the parallel relationship can reflect the complex association between entities in the text data and the image data, and this association is manifested as similarity, complementarity or a certain specific association pattern between entities. By obtaining this parallel relationship, the knowledge representation in the knowledge base can be further enriched, and the accuracy and completeness of the knowledge base can be improved. When constructing triples, the parallel relationship, the semantic relationship in the text data and the semantic relationship in the image data can be used to represent these relationships in a structured form. A triple is usually composed of an entity, a relationship and another entity. By constructing a triple, the knowledge in data of different modalities can be represented and stored in a unified form, which is convenient for subsequent knowledge retrieval and utilization.
[0071] S7, constructing a knowledge base based on the classified multimodal entities and the triples;
[0072] It should be noted that the process of constructing the knowledge base in step S7 is to organize, store and manage the classified multi-modal entities and the constructed triples to form a structured knowledge system. In this process, the classified multi-modal entities can be first indexed and standardized to facilitate subsequent retrieval and utilization. Then, the constructed triples are represented in the form of a graph to form a knowledge graph, which can intuitively display the relationship between entities and improve the readability and comprehensibility of knowledge. At the same time, attribute information such as the type, description, source, etc. of the entity can be added to the entity in the knowledge graph to enrich the content of the knowledge base. Finally, the constructed knowledge base is stored and managed to facilitate subsequent knowledge retrieval, analysis and application.
[0073] S8. Perform a quality assessment on the knowledge base, and optimize the knowledge base based on the assessment result.
[0074] It should be noted that in step S8, the quality assessment includes the evaluation of the accuracy, completeness, consistency and availability of the knowledge base. The accuracy assessment can check whether the knowledge in the knowledge base is accurate and whether there are erroneous or contradictory information. The completeness assessment can check whether the knowledge base contains all relevant knowledge and whether there are missing or missing information. The consistency assessment can check whether the knowledge in the knowledge base is consistent with each other and whether there are contradictory or conflicting information. The usability assessment can check whether the knowledge base is easy to use and whether a good user interface and query function are provided. Through the quality assessment, the problems and deficiencies in the knowledge base can be found, and the knowledge base can be optimized and improved based on the evaluation results to improve the quality and effect of the knowledge base. For example, the erroneous information in the knowledge base can be corrected, the missing information can be supplemented, the contradictory information can be coordinated, and the complex queries can be optimized to improve the accuracy and availability of the knowledge base. In addition, the knowledge base can also be continuously updated and upgraded according to actual needs and technological development to maintain the advancement and applicability of the knowledge base.
[0075] In summary, this embodiment firstly realizes the effective extraction of key information by preprocessing text data and image data, laying a solid foundation for subsequent steps. Then, by using the bidirectional encoder and random field model, the target detection model and the image description generation algorithm, the entities and semantic relationships in the text data and image data are accurately extracted respectively. This multimodal data processing method significantly improves the accuracy and efficiency of entity and semantic relationship recognition. Then, by performing feature fusion on text entities and image entities, a richer and more comprehensive multimodal entity is obtained, which provides strong support for subsequent classification and relationship extraction. In the multimodal entity classification stage, the multimodal deep learning model with the introduction of the attention mechanism shows a strong feature learning ability and classification performance, making the classification results more accurate and reliable. In addition, by obtaining the parallel relationship between text entities and image entities, and constructing triples based on these relationships and the semantic relationships in text and images, the effective association and integration of knowledge is realized. Finally, in the process of building a knowledge base, a structured, easy-to-retrieve and easy-to-use knowledge base is constructed by indexing and standardizing the multimodal entities, representing the triples in the form of graphs to form a knowledge graph, and adding entity attribute information. At the same time, by evaluating and optimizing the quality of the knowledge base, the accuracy and availability of the knowledge base are further improved. Therefore, the embodiment of the present invention not only solves the problem of low accuracy of the knowledge base in the prior art, but also realizes the effective extraction, association and integration of knowledge, providing strong support for the knowledge management of enterprises.
[0076] As an optional embodiment of the present invention, optionally, in step S2, extracting entities in the text data based on a bidirectional encoder and a random field model, and obtaining semantic relationships in the text data through a sentence analysis algorithm includes:
[0077] S201, encoding the preprocessed text data using the bidirectional encoder to capture context information;
[0078] It should be noted that in step S201, the bidirectional encoder can deeply understand the contextual information of the text data. Through the encoding process, the text data is converted into a high-dimensional vector representation. These vectors can capture the vocabulary, phrase and sentence-level information in the text, and provide rich features for subsequent entity extraction and semantic relationship analysis. After capturing the contextual information, the random field model is applied to these encoded vectors, and the entity boundaries and types in the text are identified by sequence labeling of the text data. The random field model can use contextual information to more accurately identify and classify entities. For example, in a text describing a product, the random field model can accurately identify entities such as product name, attributes, and price, providing key information for subsequent knowledge association and integration.
[0079] S202, using a random field model to perform sequence annotation on the encoded text data to identify entity boundaries;
[0080] It should be noted that in step S202, the random field model shows a strong performance in the sequence labeling task. It can label the text data word by word based on the context information, so as to accurately identify the boundaries of the entity. This step is the key to entity extraction, which determines the basis of the subsequent semantic relationship analysis. Through the labeling of the random field model, the present embodiment can obtain the accurate position and type of each entity in the text data, and provide reliable data support for the subsequent knowledge association and integration. For example, in a text describing the relationship between characters, the random field model can accurately label entities such as character names and relationship types, providing key information for the subsequent construction of the knowledge graph.
[0081] S203, determining entity types based on the context information and entity boundaries, and extracting entities from the text data;
[0082] It should be noted that, in step S203, after determining the entity boundary and type, the present embodiment can extract the corresponding entity from the text data. This step is an important part in the knowledge base construction process, because the entity is the basic unit in the knowledge base, and they represent the specific things or concepts in the real world. By accurately extracting the entity, the present embodiment can provide an accurate data basis for subsequent knowledge association and integration. For example, in a text describing a certain event, the present embodiment can extract entities such as the event name, time, place, and participants, and these entities will serve as nodes in the knowledge base and form a knowledge network through relationship connection. In the process of extracting the entity, the present embodiment can also use the context information captured by the bidirectional encoder to have a deeper understanding and analysis of the entity, thereby further improving the accuracy and efficiency of entity extraction. For example, by analyzing the context information, the present embodiment can determine whether a certain entity belongs to a specific category, or identify the potential relationship between entities, which is of great significance for subsequent knowledge association and integration.
[0083] S204: parse the sentence structure in the text data using a dependency syntax analysis algorithm to obtain the grammatical relationship between entities in the sentence.
[0084] It should be noted that in step S204, the dependency syntax analysis algorithm can perform in-depth analysis on the sentences in the text data to obtain the grammatical relationship between the entities in the sentence. This grammatical relationship reflects the dependency and dominance relationship of the entity in the sentence, and is an important basis for understanding the sentence structure and semantics. Through dependency syntax analysis, the present embodiment can identify the key components such as the subject, predicate, and object in the sentence, as well as the dependency relationship between them, such as the verb-object relationship, the subject-predicate relationship, etc. This information is crucial for the subsequent extraction of semantic relationships and the construction of a knowledge base. For example, in a text describing the function of a product, dependency syntax analysis can help the present embodiment accurately identify entities such as product names, function descriptions, and the grammatical relationship between them, such as the verb-object relationship of "product A has function B". This information will be used for the subsequent construction of triples in the knowledge base to achieve effective association and integration of knowledge. Therefore, in step S204, the application of the dependency syntax analysis algorithm is one of the key steps in the knowledge base construction process.
[0085] As an optional embodiment of the present invention, optionally, in step S3, entities in the image data are obtained based on the target detection model, and semantic relationships in the image data are obtained through an image description generation algorithm:
[0086] S301, using a target detection model to locate an object in the image data, and identify the object, and taking the object as an entity;
[0087] It should be noted that in step S301, the target detection model can accurately locate objects in the image and accurately identify the categories of these objects by performing deep learning and feature extraction on the image data. These identified objects are regarded as entities in the image data. They represent the key information in the image and are the basis for subsequent knowledge association and integration. For example, in an image containing a variety of fruits, the target detection model can accurately identify objects such as apples, bananas, and oranges, and extract them as entities.
[0088] S302, classifying the objects by labels to obtain a semantic label for each object;
[0089] It should be noted that in step S302, the label classification process is a process of further refining and classifying the objects identified by the target detection model. By assigning specific semantic labels to these objects, the present embodiment can more accurately understand the entities in the image data and the meanings they represent. For example, in an image containing a vehicle, the target detection model identifies objects such as cars, motorcycles and bicycles, and the label classification process can assign specific semantic labels such as "vehicle-car", "vehicle-motorcycle" and "vehicle-bicycle" to these objects. These semantic labels not only help the present embodiment to understand the image data more deeply, but also provide a more accurate data basis for subsequent knowledge association and integration. In the label classification process, the present embodiment can use a deep learning algorithm or a pre-trained classification model to extract and classify features of the image data to obtain the semantic label of each object. At the same time, the present embodiment can also continuously optimize and improve the label classification algorithm according to actual needs and technological development to improve the accuracy and efficiency of label classification.
[0090] S303: Generate a descriptive sentence related to the object in the image data based on the semantic tag using an image description generation algorithm, and acquire a semantic relationship in the image data based on the descriptive sentence.
[0091] It should be noted that in step S303, the image description generation algorithm can generate descriptive sentences related to the objects in the image based on the semantic tags. These sentences not only describe the appearance characteristics of the objects, but also reflect the spatial relationship and semantic connection between the objects. By parsing these descriptive sentences, the present embodiment can extract the semantic relationships in the image data, such as the parallel relationship, subordinate relationship, action relationship, etc. between the objects. These semantic relationships are of great significance for the subsequent construction of triples in the knowledge base and the effective association and integration of knowledge. For example, in an image describing a kitchen scene, the image description generation algorithm generates a descriptive sentence such as "There is an apple and a banana on the table". By parsing this sentence, the present embodiment can extract the parallel relationship between apples and bananas, and the subordinate relationship between them and the table. This information will be used for the subsequent construction of triples in the knowledge base, such as (apple, parallel relationship, banana), (apple, located, on the table), etc. Therefore, in step S303, the application of the image description generation algorithm is another key step in the knowledge base construction process. By combining the target detection model and the image description generation algorithm, this embodiment can effectively extract entities and semantic relationships from image data, providing strong support for subsequent knowledge association and integration.
[0092] As an optional embodiment of the present invention, optionally, the expression of the multimodal entity obtained in step S4 is:
[0093]
[0094] Among them, F represents the fused multi-modal entity features, H() represents the feature fusion function, and F t The feature vector set representing the text entity, F g represents the feature vector set of image entities, ψ represents the parameter set of feature fusion function, σ() represents the activation function, n represents the number of text features, i represents the i-th text feature, α i represents the weight of the i-th text feature, f text represents the text feature extraction function, e i represents the i-th text feature, represents the parameters of the text feature extraction function, m represents the number of image features, j represents the jth image feature, β j represents the weight of the jth image feature, f image () represents the image feature extraction function, b j and c j Represents two different image features in the image, Represents the parameters of the image feature extraction function.
[0095] As an optional embodiment of the present invention, optionally, constructing a multimodal deep learning model in step S5, and introducing an attention mechanism into the multimodal deep learning model includes:
[0096] S501, extracting historical text data and historical image data, and preprocessing the historical text data and historical image data respectively;
[0097] It should be noted that the historical text data and historical image data are extracted in step S501 in order to train and optimize the multimodal deep learning model. The preprocessing process includes steps such as data cleaning, format conversion, and feature extraction, which aim to improve the quality and consistency of the data and provide a reliable data basis for subsequent training and reasoning. Through in-depth analysis of historical data, this embodiment can understand the distribution patterns and semantic relationships of entities in texts and images, thereby providing strong support for the design and optimization of the model. After the preprocessing is completed, these data will be used to train the multimodal deep learning model to learn the association and mapping relationship between text and images, providing strong support for subsequent multimodal entity classification and relationship extraction.
[0098] S502: construct an initial multimodal deep learning model, input the preprocessed historical text data and historical image data into the initial multimodal deep learning model for training, and obtain a trained multimodal deep learning model;
[0099] like Figure 2 As shown, it should be noted that the first step of building a multimodal deep learning model in step S502 is to preprocess the historical text data and historical image data. This step includes data cleaning, formatting, feature extraction, etc. For text data, operations such as word segmentation, stop word removal, and stem extraction may be required; for image data, resizing, normalization, data enhancement, etc. may be required; for audio and video data, preprocessing steps such as sampling rate adjustment and frame extraction are also required.
[0100] Next, you need to choose a suitable deep learning framework to build the model. For example, deep learning frameworks include TensorFlow, PyTorch, etc. These frameworks provide a wealth of neural network layers, optimization algorithms, and tool functions, which can greatly simplify the model building and training process.
[0101] When designing the network structure of a multimodal deep learning model, it is necessary to consider how to effectively fuse features from different modalities. One approach is to use a multimodal fusion layer, which can concatenate, weight, or use attention mechanisms to process features from different modalities to generate a unified feature representation. In addition, specific network structures can be designed based on specific task requirements, such as convolutional neural networks (CNNs) for image feature extraction, recurrent neural networks (RNNs) or Transformers for sequence data processing, etc.
[0102] The loss function is a function that measures the difference between the model's predicted results and the actual results. In multimodal deep learning models, it is necessary to select a suitable loss function according to the specific task, such as cross entropy loss, mean square error loss, etc. At the same time, it is also necessary to select a suitable optimization algorithm to update the model parameters, such as stochastic gradient descent (SGD), Adam, etc.
[0103] In the model training phase, the model needs to be trained using preprocessed multimodal data. During the training process, it is necessary to monitor the changes in the loss function and the performance indicators (such as accuracy, recall, etc.) on the validation set. The model performance can be optimized by adjusting hyperparameters such as learning rate and batch size. After training, the model needs to be evaluated on the test set to verify its generalization ability.
[0104] Finally, the trained multimodal deep learning model is deployed to actual application scenarios. This includes exporting the model into a deployable format (such as ONNX, PMML, etc.) and integrating it into existing systems or applications. In practical applications, factors such as the real-time performance, robustness, and scalability of the model also need to be considered.
[0105] S503, introducing an attention mechanism into the trained multimodal deep learning model, and optimizing the multimodal deep learning model;
[0106] It should be noted that the introduction of the attention mechanism in step S503 is to further improve the performance and accuracy of the multimodal deep learning model when processing complex data. The attention mechanism can simulate the attention allocation process of humans when processing information. By calculating the importance weights of different parts of data, the model can pay more attention to key information, thereby improving the accuracy of entity extraction and relationship extraction. After the introduction of the attention mechanism, the present embodiment optimizes the multimodal deep learning model, including adjusting the network structure, updating the weight parameters and other steps to ensure that the model can fully utilize the advantages of the attention mechanism and improve the overall performance. Through the optimized model, the present embodiment can more accurately extract entities in texts and images, and effectively associate the semantic relationships between them, providing more reliable data support for the subsequent knowledge base construction. The realization of this step marks that the present embodiment has made important progress in multimodal data processing and knowledge base construction.
[0107] S504: Classify the multimodal entities using the trained multimodal deep learning model.
[0108] It should be noted that the multimodal deep learning model trained in step S504 can classify the extracted multimodal entities. This step is to automatically assign entities to corresponding categories based on the association and mapping relationship between text and image features learned by the model. The classification process involves multiple levels of judgment and analysis to ensure that the entities are accurately classified. Through classification, this embodiment can further understand and organize the entities in the knowledge base, and provide a clearer and more orderly data basis for subsequent knowledge association and integration. At the same time, the classification results can also be used to evaluate and optimize the performance of the model, helping this embodiment to continuously improve and optimize the method of knowledge base construction.
[0109] As an optional embodiment of the present invention, optionally, the expression for classifying the multimodal entity is:
[0110]
[0111] Among them, C represents the set of multimodal entity labels after classification, J() represents the entity classification function, F represents the fused multimodal entity features, Ω represents the parameter set of the multimodal deep learning model, ck represents the category label of the kth output, l represents the total number of labels, represents the maximum probability, L represents the category set, c represents the category label, ω k Represents the classification parameters of a multimodal deep learning model.
[0112] As an optional embodiment of the present invention, optionally, obtaining the parallel relationship between the entity in the text data and the entity in the image data in step S6 includes:
[0113] S601, establishing a mapping relationship between entities in the text data and entities in the image data based on entity similarity calculation;
[0114] It should be noted that the calculation of entity similarity in step S601 is a key step in determining the degree of association between entities in text data and entities in image data. By calculating the similarity between entities, the present embodiment can establish a mapping relationship between them, thereby identifying which entities are corresponding in text and image. The implementation of this step relies on a variety of technologies and algorithms, such as vector space model, cosine similarity calculation, deep learning algorithm, etc. By comprehensively considering the semantic features, appearance features and contextual information of the entities, the present embodiment can calculate the similarity between entities and establish a mapping relationship accordingly. These mapping relationships provide an important basis for subsequent knowledge association and integration, so that the present embodiment can more accurately understand the entities in text and image data and the relationship between them.
[0115] S602: Based on the mapping relationship, identify a parallel relationship between the entity in the text data and the entity in the image data, wherein the parallel relationship indicates that the entity in the text data and the entity in the image data are associated and independent of each other;
[0116] It should be noted that, based on the mapping relationship in step S602, the present embodiment can further analyze the parallel relationship between the entities in the text data and the entities in the image data. The parallel relationship refers to the existence of a certain association between the entities in the text and the image, but this association does not constitute a subordinate or dependent relationship, but is independent and coexistent. For example, in a news report, the text describes a scene of someone walking in the park, while the image shows the specific location and environment of the person in the park. In this case, the "person" in the text and the "person" in the image are mutually related parallel entities. By identifying the parallel relationship, the present embodiment can more comprehensively understand the entities in the text and image data and the complex connections between them, and provide more abundant information for subsequent knowledge association and integration. In order to achieve this goal, the present embodiment adopts a variety of technologies and algorithms, such as rule-based matching methods, deep learning algorithms, etc., to accurately identify parallel relationships and build a corresponding knowledge base.
[0117] S603: verify the parallel relationship.
[0118] It should be noted that the purpose of verifying the parallel relationship in step S603 is to ensure that the identified parallel relationship is accurate and to provide reliable data support for the subsequent knowledge base construction. The verification process involves judgment and analysis in many aspects, such as semantic consistency, context relevance, and the degree of matching between image and text features. By comprehensively considering these factors, the present embodiment can verify the parallel relationship one by one to ensure its accuracy and reliability. During the verification process, the present embodiment also needs to use means such as manual review or expert evaluation to improve the accuracy and efficiency of the verification. The verified parallel relationship will be used in the subsequent knowledge base construction and association integration process, providing strong support for the effective organization and utilization of knowledge.
[0119] As an optional embodiment of the present invention, optionally, the expression for constructing the triple in step S6 is:
[0120] T triplet =K triplet (E text ,E image ,R text ,R image ,Ξ triplet )
[0121] ={(u x ,r,u y )∣(u x ,u y )∈U,r∈R(u x ,u y ,x r )}
[0122] Among them, T triplet represents the constructed triple set, K triplet represents the triple construction function, E text Represents a text entity set, E image represents the image entity set, R text Represents the set of semantic relations between text entities, R image represents the set of semantic relations between image entities, triplet Represents the parameter set of the triple construction function, T triplet Represents a tuple of entity x, relationship, and entity y. R() represents a set of semantic relationships. r Represents semantic relationship parameters.
[0123] As an optional embodiment of the present invention, optionally, in step S7, constructing a knowledge base based on the classified multimodal entities and the triples includes:
[0124] S701, constructing an initial database;
[0125] It should be noted that constructing the initial database in step S701 is a basic link in the knowledge base construction process. This database is intended to store the classified multi-mode entities and constructed triples, providing powerful data support for subsequent knowledge association, integration and utilization. During the construction process, the present embodiment fully considers the characteristics and storage requirements of the data, and designs a reasonable database structure and storage strategy. The initial database contains multiple tables or collections to store information such as text entities, image entities, semantic relationships and triples respectively. At the same time, in order to improve the query efficiency and reliability of the database, the present embodiment also adopts strategies such as indexing and optimizing query statements to ensure that the database can quickly respond to various query requirements. After the construction is completed, the present embodiment also tests and verifies the database to ensure that its performance and stability meet the needs of practical applications. The implementation of this step provides a solid foundation for the subsequent knowledge base construction and association integration.
[0126] S702: Standardize the multimodal entities, uniquely identify each of the multimodal entities, and establish an index of the multimodal entities;
[0127] It should be noted that the purpose of standardizing the multi-mode entities in step S702 is to ensure the consistency and accuracy of the entities in the knowledge base. The standardization process includes the unification of entity naming conventions, attribute definitions, classification standards, etc. By standardizing the multi-mode entities, the present embodiment can eliminate ambiguity and duplication between entities and improve the quality and availability of the knowledge base. At the same time, in order to more efficiently manage and query the entities in the knowledge base, the present embodiment also needs to uniquely identify each multi-mode entity and establish a corresponding index. The unique identifier can ensure that each entity has a unique identity in the knowledge base, while the index can speed up the query and retrieval process of the entity. When establishing the index, the present embodiment needs to consider factors such as the attributes, relationships, and positions of the entities in the knowledge base to design a reasonable index structure and query algorithm.
[0128] S703, mapping the relations and entities in the triples to the initial knowledge base;
[0129] It should be noted that mapping the relationships and entities in the triples to the initial knowledge base in step S703 is a key step in the knowledge base construction process. This step is intended to integrate the previously extracted, classified and associated multi-mode entities and the relationships between them into the knowledge base for subsequent knowledge query, reasoning and application. In the mapping process, this embodiment fully considers the structure and storage requirements of the knowledge base and designs a reasonable mapping strategy and algorithm. Specifically, this embodiment maps the entities in the triples to the corresponding entity nodes in the knowledge base and maps the relationships to the connecting edges between the entity nodes. In this way, this embodiment can represent the entities in the knowledge base and the relationships between them through entity nodes and connecting edges to form a complete knowledge network. At the same time, in order to ensure the accuracy and reliability of the mapping, this embodiment also uses a variety of technologies and algorithms for verification and optimization, such as rule-based matching methods, deep learning algorithms, etc. Through mapping, this embodiment can effectively integrate the previously extracted and associated multi-mode entities and the relationships between them into the knowledge base, providing strong support for subsequent knowledge query, reasoning and application.
[0130] S704, representing the relationships and entities in the triples in the form of a graph through a knowledge representation model to obtain a knowledge graph;
[0131] It should be noted that, in step S704, through the knowledge representation model, this embodiment can intuitively represent the relationship and entity in the triple in the form of a graph, thereby constructing a knowledge graph. The knowledge graph is a structured knowledge representation method that can clearly display the relationship and attributes between entities, and provides a more convenient and efficient way for knowledge query, reasoning and application. When constructing the knowledge graph, this embodiment fully considers the entity and relationship characteristics in the knowledge base, and designs a reasonable node and edge representation method. Specifically, this embodiment represents the entity as a node in the graph, represents the relationship as an edge between the nodes, and represents the direction and strength of the relationship by the direction and weight of the edge. At the same time, in order to enhance the readability and comprehensibility of the knowledge graph, this embodiment also uses a variety of visualization technologies and tools, such as the setting of node color, shape, size and other attributes, as well as interactive operations such as the layout, zooming, and translation of the graph. Through the construction of the knowledge graph, this embodiment can more intuitively understand and analyze the entities in the knowledge base and the relationships between them, and provide more abundant and intuitive information support for subsequent knowledge query, reasoning and application.
[0132] S705. Add attribute information to the entities in the knowledge graph to obtain a knowledge base.
[0133] It should be noted that the purpose of adding attribute information to the entities in the knowledge graph in step S705 is to further improve the content of the knowledge base and improve its practicality and value. Attribute information is a detailed description and supplement to the characteristics of the entity, which can help this embodiment to more comprehensively understand the meaning and context of the entity. When adding attribute information, this embodiment needs to consider multiple aspects, such as the physical attributes, functional attributes, and relationship attributes of the entity. Physical attributes include the shape, size, color, etc. of the entity; functional attributes involve the role, purpose, performance, etc. of the entity; and relationship attributes describe the association and interaction between the entity and other entities. By comprehensively considering these attribute information, this embodiment can provide a more comprehensive and accurate description of the entities in the knowledge base, thereby enhancing the practicality and reliability of the knowledge base. At the same time, the addition of attribute information can also provide richer information support for subsequent knowledge query, reasoning, and application, helping this embodiment to better understand and utilize the knowledge in the knowledge base.
[0134] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0135] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A knowledge base construction method based on automatic extraction, classification and association of multimodal entities, characterized in that: The method comprises: S1, preprocessing text data and image data respectively; S2, extracting entities in the text data based on a bidirectional encoder and a random field model, and obtaining semantic relationships in the text data through a sentence analysis algorithm; S3, acquiring entities in the image data based on a target detection model, and acquiring semantic relationships in the image data through an image description generation algorithm; S4, performing feature fusion on the entities in the text data and the entities in the image data to obtain a multimodal entity; S5. Construct a multimodal deep learning model and introduce an attention mechanism into the multimodal deep learning model; classify the multimodal entities based on the multimodal deep learning model to obtain classified multimodal entities; S6, obtaining a parallel relationship between entities in the text data and entities in the image data, and constructing a triple based on the parallel relationship, a semantic relationship in the text data, and a semantic relationship in the image data; S7, constructing a knowledge base based on the classified multimodal entities and the triples; S8. Perform a quality assessment on the knowledge base, and optimize the knowledge base based on the assessment result.
2. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: In step S2, entities in the text data are extracted based on the bidirectional encoder and the random field model, and semantic relations in the text data are obtained through a sentence analysis algorithm, including: S201, encoding the preprocessed text data using the bidirectional encoder to capture context information; S202, using a random field model to perform sequence annotation on the encoded text data to identify entity boundaries; S203, determining entity types based on the context information and entity boundaries, and extracting entities from the text data; S204: parse the sentence structure in the text data using a dependency syntax analysis algorithm to obtain the grammatical relationship between entities in the sentence.
3. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: In step S3, entities in the image data are obtained based on the target detection model, and semantic relationships in the image data are obtained through an image description generation algorithm: S301, using a target detection model to locate an object in the image data, and identify the object, and taking the object as an entity; S302, classifying the objects by labels to obtain a semantic label for each object; S303: Generate a descriptive sentence related to the object in the image data based on the semantic tag using an image description generation algorithm, and acquire a semantic relationship in the image data based on the descriptive sentence.
4. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: The expression for obtaining the multimodal entity in step S4 is: Among them, F represents the fused multi-modal entity features, H() represents the feature fusion function, and F t The feature vector set representing the text entity, F g represents the feature vector set of image entities, ψ represents the parameter set of feature fusion function, σ() represents the activation function, n represents the number of text features, i represents the i-th text feature, a i represents the weight of the i-th text feature, f text represents the text feature extraction function, e i represents the i-th text feature, represents the parameters of the text feature extraction function, m represents the number of image features, j represents the jth image feature, β j represents the weight of the jth image feature, f image () represents the image feature extraction function, b j and c j Represents two different image features in the image, Represents the parameters of the image feature extraction function.
5. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: In step S5, a multimodal deep learning model is constructed, and an attention mechanism is introduced into the multimodal deep learning model, including: S501, extracting historical text data and historical image data, and preprocessing the historical text data and historical image data respectively; S502: construct an initial multimodal deep learning model, input the preprocessed historical text data and historical image data into the initial multimodal deep learning model for training, and obtain a trained multimodal deep learning model; S503, introducing an attention mechanism into the trained multimodal deep learning model, and optimizing the multimodal deep learning model; S504: Classify the multimodal entities using the trained multimodal deep learning model.
6. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1 or 5, characterized in that: The expression for classifying the multimodal entity is: Where C represents the set of multimodal entity labels after classification, J() represents the entity classification function, F represents the fused multimodal entity features, Ω represents the parameter set of the multimodal deep learning model, and c k represents the category label of the kth output, l represents the total number of labels, represents the maximum probability, L represents the category set, c represents the category label, ω k Represents the classification parameters of a multimodal deep learning model.
7. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: Acquiring the parallel relationship between the entity in the text data and the entity in the image data in step S6 includes: S601, establishing a mapping relationship between entities in the text data and entities in the image data based on entity similarity calculation; S602: Based on the mapping relationship, identify a parallel relationship between the entity in the text data and the entity in the image data, wherein the parallel relationship indicates that the entity in the text data and the entity in the image data are associated and independent of each other; S603: verify the parallel relationship.
8. The method for constructing a knowledge base based on automatic extraction, classification and association of multimodal entities according to claim 1, characterized in that: The expression for constructing the triple in step S6 is: T triplet =K triplet (E text ,E image ,R text ,R image ,Ξ triplet ) ={(in x ,r,u y )∣(in x ,in y )∈U,r∈R(u x ,in y ,x r )} Among them, T triplet represents the constructed triple set, K triplet represents the triple construction function, E text Represents a text entity set, E image represents the image entity set, R text Represents the set of semantic relations between text entities, R image represents the set of semantic relations between image entities, triplet Represents the parameter set of the triple construction function, T triplet Represents a tuple of entity x, relationship, and entity y. R() represents a set of semantic relationships. r Represents semantic relationship parameters.
9. The method for constructing a knowledge base based on automatic extraction, classification and association of multi-modal entities according to claim 1, characterized in that: In step S7, constructing a knowledge base based on the classified multimodal entities and the triples includes: S701, constructing an initial database; S702: Standardize the multimodal entities, uniquely identify each of the multimodal entities, and establish an index of the multimodal entities; S703, mapping the relations and entities in the triples to the initial knowledge base; S704, representing the relationships and entities in the triples in the form of a graph through a knowledge representation model to obtain a knowledge graph; S705. Add attribute information to the entities in the knowledge graph to obtain a knowledge base.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 9.
Citation Information
Cited By
Engineering drawing automatic labeling method and device and medium
CN120340032A