Intangible cultural heritage multi-modal knowledge graph construction method and system based on big data
By cleaning and annotating multimodal intangible cultural heritage data, a unified domain knowledge model is built, and the problem of multi-source heterogeneous data integration is solved, efficient construction and application of intangible cultural heritage knowledge graphs is achieved, and data processing capabilities and consistency are improved.
Patent Information
- Application Number
- CN202510811710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing technology is difficult to effectively integrate multi-source heterogeneous intangible cultural heritage data, resulting in low efficiency in large-scale and real-time applications and insufficient ability to process unstructured data, affecting the construction effect and consistency of the knowledge graph.
By collecting multimodal intangible cultural heritage data, conducting preliminary cleaning and labeling, establishing a first-version multimodal data set, using multimodal data cleaning and labeling models for processing, building a unified domain knowledge model, forming a structured multimodal knowledge graph, and resolving conflicts and contradictions between different modal data.
It realizes efficient integration of multi-source heterogeneous data, improves the efficiency of acquisition and utilization of intangible cultural heritage knowledge, optimizes the storage and query mechanism, improves the construction effect and consistency of the knowledge graph, and supports the query, reasoning and visual application of intangible cultural heritage.
Smart Images

Figure CN120354927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph construction, and particularly to a method and system for constructing a multimodal knowledge graph of intangible cultural heritage based on big data. Background Art
[0002] With the rapid development of information technology and digital technology, the protection and inheritance of intangible cultural heritage have received increasing attention. However, the existing intangible cultural heritage protection methods mainly rely on manual recording and collation, which are inefficient and difficult to make full use of multi-source heterogeneous data. Traditional intangible cultural heritage digitization methods often focus on the collection and display of a single data modality, such as text records, picture preservation, or video recording. However, these methods are unable to cope when faced with increasingly complex intangible cultural heritage information. As a comprehensive cultural phenomenon, intangible cultural heritage has rich and diverse connotations, covering various data forms such as text, images, audio, video, and 3D models, and it is difficult to fully express its value and significance through a single data modality.
[0003] Therefore, how to effectively integrate multi-source heterogeneous intangible cultural heritage data, construct a structured knowledge base, and achieve intelligent management and application of intangible cultural heritage has become an urgent problem to be solved. When dealing with large-scale multimodal data, existing technologies often face problems of low storage and query efficiency, which hinder the application of knowledge graphs in large-scale and real-time intangible cultural heritage fields. In addition, existing technologies have weak processing capabilities for unstructured data such as text, images, and videos, and the knowledge contained in a large amount of unstructured data is ignored, resulting in limited construction effects of knowledge graphs. Summary of the Invention
[0004] One of the purposes of the present invention is to provide a method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data to solve the problem that the knowledge contained in unstructured data in the prior art is ignored.
[0005] The present invention is realized through the following technical solutions. A method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data includes the following steps: S100, collecting intangible cultural heritage data of different modalities, preliminarily cleaning and annotating the collected data, and establishing a preliminary association between different modality data to obtain an initial version of the multimodal dataset; S200, processing the initial version of the multimodal dataset through a multimodal data cleaning and annotation model to obtain a final multimodal dataset, an annotation consistency matrix, and a cross-modal association matrix; S300, constructing a unified domain knowledge model, structuring the modeling of intangible cultural heritage-related concepts, entities, and their relationships to form a semantic expression and knowledge organization form that supports multimodal data; S400, constructing a structured multimodal knowledge graph according to the data in the knowledge organization form output by S300.
[0006] Furthermore, the intangible cultural heritage data in different modalities include: text data related to intangible cultural heritage, including: texts, literature, historical documents, research papers, project introductions, and expert interviews describing the background, development, inheritance, etc. of intangible cultural heritage items; image data related to intangible cultural heritage, including: complete photos of intangible cultural heritage projects, decomposition diagrams of technological steps, photos of inheritors and their skill demonstrations, on-site images of folk activities; audio and video data related to intangible cultural heritage, including: recordings and videos of the production process of intangible cultural heritage skills, audio and video of oral intangible cultural heritage inheritance.
[0007] Furthermore, the intangible cultural heritage data in different modalities can also include: structured data from official institutions, such as the UNESCO or national intangible cultural heritage catalogs, including the name, category, region, inheritor information, etc. of the projects. In addition, the inheritor file information, such as personal background, skill inheritance process, social influence, etc., should also be collected.
[0008] Furthermore, the initial cleaning and annotation include: removing duplicate paragraphs from the collected text data and standardizing terms; preliminarily associating different modalities of data based on file names, timestamps, metadata, etc., and annotating the core entities and their mutual relationships in text, image, audio, and video data.
[0009] Furthermore, the multi-modal data cleaning and annotation model ensures that the data cleaning of different modalities can be completed efficiently and accurately through multiple constraints and losses. The multi-modal data cleaning and annotation model includes: S210, denoising processing, removing noise from the data, and avoiding the interference of noise samples through denoising loss to improve the overall data cleaning effect; S220, cross-modal alignment, aligning the features of different modalities so that samples with the same semantics have similar representations between different modalities and optimizing the relationship between modalities; S230, annotation consistency, ensuring that the same class of samples under different modalities maintain consistent labels and avoiding classification errors caused by inconsistent annotations; S240, correlation matrix constraint, promoting the model to establish associations only between high-similarity sample pairs through sparsity constraints, avoiding the model overfitting irrelevant samples, and controlling the relative importance of each objective in the overall optimization by adjusting the weight coefficients of different functions.
[0010] Furthermore, the objective function of the multi-modal data cleaning and annotation model can be expressed by the following formula:
[0011] , where, is the final multi-modal dataset, including, the cleaned text data matrix, the cleaned image data matrix, the cleaned audio data matrix, For the initial version of the multi-modal dataset, including, Initial text data matrix, Initial image data matrix, Initial audio data matrix, is the weight coefficient for controlling the denoising intensity, is the data denoising function, is the weight coefficient for controlling the optimization degree of modal alignment, is the cross-modal correlation optimization function, is the weight coefficient for controlling the annotation consistency, is the annotation consistency constraint function, is the weight coefficient for controlling the sparsity of the correlation matrix, is the correlation matrix constraint function, is the annotation consistency matrix, is the cross-modal correlation matrix.
[0012] Furthermore, the data denoising function is used to ensure that the model can correctly extract effective features from the data. The data denoising function reduces the noise impact by penalizing the inconsistencies or incorrect features in the model, thereby improving the quality of the feature representation; the cross-modal correlation optimization function is used to optimize the alignment of cross-modal features, and by making the features of different modalities as close as possible in a shared embedding space, it improves the performance of the model in cross-modal tasks; the annotation consistency constraint function is used to ensure the annotation consistency of samples during the learning process of the model, ensuring that the same class of samples in different modalities have consistent annotation information; the correlation matrix constraint function sparsifies the correlation matrix so that there are non-zero correlations only between cross-modal features with sufficient similarity. The correlation matrix constraint function uses the L1 norm constraint to encourage most of the elements in the correlation matrix to be close to zero, thereby reducing the connections between irrelevant or low-similarity samples and maintaining the efficiency of the model.
[0013] Furthermore, the data denoising function can be represented by the following formula:
[0014] , where, is the set of initial data matrices, is the set of data matrices after cleaning, is the Frobenius norm, also known as the Euclidean norm of the matrix, used to measure the difference between two matrices. In the above formula represents the sum of the squared differences between the original data and the denoised data, used to measure the effect of the denoising process on the data matrix. If the denoising effect is better, this difference is smaller and the loss value is lower. is the nuclear norm, used to constrain the rank of the matrix and encourage the matrix to have a low-rank structure. In the above formula It is used to force the matrix after denoising to maintain a low rank and remove redundant information, which means that if the rank of the set of data matrices after cleaning is low, it will be more compact and contain fewer redundant terms, helping to clean up invalid or duplicate data. is a hyperparameter for the denoising intensity, used to control the denoising intensity, which balances the weights of the Frobenius norm term and the nuclear norm term in the objective function. When is larger, the denoising effect will be stronger, and the low-rank constraint of the matrix will be more stringent. That is to say, at this time will be closer to a low-rank matrix, thus removing more noise.
[0015] Furthermore, the cross-modal association optimization function can be expressed by the following formula:
[0016] where, is a pair of positive samples (i.e., sample pairs with actual associations), and these sample pairs are usually determined in advance in some way (such as file names, timestamps, labels, etc.), which means that these two samples should have an actual connection across modalities (for example, the text description and video clip of the same event); is a pair of negative samples (i.e., sample pairs without actual associations), and these sample pairs are obtained by random sampling, aiming to help the model learn the differences between different modalities and prevent the model from misidentifying irrelevant samples as relevant, is the feature vector of sample i, is the feature vector of sample j, is the cosine similarity between the feature vector of sample i and the feature vector of sample j, used to measure the similarity between two samples in the feature space, and the value range is [-1, 1]. The closer the value is to 1, the more similar the two samples are, and the closer the value is to -1, the less similar they are. is the similarity threshold,
[0017] Furthermore, the annotation consistency constraint function can be expressed by the following formula:
[0018] , where, is the set of annotated samples, representing the set of samples with known labels, and the labels of these samples are annotated by experts. Usually, the number of these samples is small, but their labels are very accurate, and the model can use the labels of these samples to guide the training. is the annotation consistency probability of sample i, that is, the label prediction probability of the model for sample i. During the optimization process, the model will learn how to adjust to make it closer to the true label , and ensure that similar samples have similar labels by propagating the consistency of the labels; is the true label of sample i, coming from expert annotation; is a hyperparameter for controlling the intensity of label propagation. If is large, the influence of label propagation is strong, and the label consistency constraint between similar samples will be more stringent; if is small, the influence of the label consistency constraint is weak, and the model may rely more on the labeled samples. is the Laplacian matrix, which is constructed based on the similarity between samples and is usually used to represent the relationship between nodes in a graph structure. Specifically, the Laplacian matrix is used to describe the similarity structure between samples and control the label consistency between similar samples. is the trace operation of the matrix, which calculates the sum of the diagonal elements of , representing a global consistency of label propagation. is the transpose matrix of the annotation consistency matrix, is the matrix transpose symbol.
[0019] Furthermore, the correlation matrix constraint function can be expressed by the following formula: , , where is the cross-modal correlation matrix. is the L1 norm of the cross-modal correlation matrix, that is, the sum of the absolute values of all elements in the cross-modal correlation matrix. The role of the L1 norm is to promote the sparsity of the cross-modal correlation matrix, making most elements in the correlation matrix close to zero. By minimizing encourages the model to only retain a few important (high-similarity) correlations, while other low-similarity correlations will be suppressed; s.t. is the abbreviation of subject to, indicating that this function is restricted by the cross-modal feature similarity between sample i and sample j.
[0020] Furthermore, constructing a unified domain knowledge model includes the following sub-steps: S310, defining core ontology classes based on the annotation consistency matrix, counting the categories in the annotation consistency matrix, calculating the total confidence of each category, and if the overall confidence of a certain category is significantly higher than other categories, then define this category as the core ontology class; S320, realizing relationship definition based on the cross-modal correlation and knowledge rules of the cross-modal correlation matrix, inferring the relationship between different entities through the information of cross-modal data pairs in the cross-modal correlation matrix, and achieving alignment between different modalities; S330, extracting text knowledge according to the data in the data matrix set.
[0021] Furthermore, extracting text knowledge includes: entity extraction and relationship extraction, where entity extraction uses the high-confidence labels of the annotation consistency matrix as supervision signals and is realized through the BiLSTM+CRF model; relationship extraction extracts and strengthens the confidence of relationship types from the text through enhanced dependency parsing, combined with the cosine similarity between text and images.
[0022] Furthermore, extracting text knowledge also includes: using images and audio as knowledge supplements, based on the alignment of images and text, using the information in the cross-modal correlation matrix to find the text corresponding to each image, so as to generate annotations for the images; then using an object detection model to identify specific regions in the images and match them through the similarity with text entities, so as to generate relational triples.
[0023] Furthermore, considering that there may be conflicts in the descriptions of the same entity by different modalities, in order to solve the problem of conflicts in the descriptions of the same entity by different modalities, in this embodiment, it is also possible to enhance cross-modal consistency by mapping the embeddings of entities in different modalities into a unified space, thereby solving the conflict problem.
[0024] Furthermore, the conflict loss function can be expressed by the following formula:
[0025] , where is the conflict loss function, is the set of cross-modal data pairs of the same entity, is the data pair of the m-th modality of the same entity, is the data pair of the n-th modality of the same entity, is the input data of the m-th modality, is the input data of the n-th modality, is the embedding mapping function of the m-th modality, is the embedding mapping function of the n-th modality.
[0026] On the other hand, the present invention provides a non-material cultural heritage multi-modal knowledge graph construction system based on big data, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the non-material cultural heritage multi-modal knowledge graph construction method as described in any one of the above.
[0027] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0028] 1. The present invention can effectively integrate multi-source heterogeneous intangible cultural heritage data, make full use of the rich information contained in multi-modal data, improve the acquisition and utilization efficiency of intangible cultural heritage knowledge, and at the same time, through an optimized storage and query mechanism, improve the efficiency of processing large-scale multi-modal data, and support the application of the knowledge graph in the large-scale and real-time intangible cultural heritage field.
[0029] 2. The present invention adopts an efficient processing method for unstructured data such as text, images, videos, etc., fully excavates the knowledge contained in the unstructured data, improves the construction effect of the knowledge graph, and proposes an effective method to solve the conflicts and contradictions existing between different modal data, ensuring the consistency and accuracy of the knowledge graph.
[0030] 3. The present invention realizes the deep fusion and correlation modeling of multi-modal data, fully excavates the internal connections between different modalities, improves the expression ability of the knowledge graph, and finally the constructed multi-modal knowledge graph can support applications such as querying, reasoning, and visualization of intangible cultural heritage, which is beneficial to the effective protection and inheritance of intangible cultural heritage. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0032] Figure 1 It is a flowchart of the method provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0034] Embodiment 1
[0035] The inheritance and protection of intangible cultural heritage (hereinafter referred to as 'ICH') are manifestations of a country's soft power at the current stage. With the development of artificial intelligence technology, exploring the combination of artificial intelligence and ICH data to achieve intelligent management of ICH data has become an urgent research topic. In the prior art, it is difficult to effectively integrate multi-source heterogeneous ICH data, and it is impossible to make full use of the rich information contained in multi-modal data, resulting in low efficiency in obtaining and using ICH knowledge; when dealing with large-scale multi-modal data, the storage and query efficiency is relatively low, which hinders the application of multi-modal knowledge graphs in large-scale and real-time ICH fields; moreover, the existing knowledge graph construction methods have weak processing capabilities when dealing with unstructured data such as text, images, videos, etc., and a large amount of knowledge contained in unstructured data is ignored, affecting the construction effect of the knowledge graph. When constructing an ICH knowledge graph, there is a lack of effective methods to solve the conflicts and contradictions existing between different modal data, affecting the consistency and accuracy of the knowledge graph.
[0036] To solve these problems and promote the integration and application of artificial intelligence in intangible cultural heritage, this embodiment provides a method for constructing a multi-modal knowledge graph of intangible cultural heritage based on big data. Figure 1 The flowchart of the method in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:
[0037] Step 1: First, obtain diverse intangible cultural heritage data, preprocess the data to eliminate noise and establish preliminary associations, laying a foundation for subsequent construction of a multi-modal knowledge graph of intangible cultural heritage.
[0038] The diverse intangible cultural heritage data includes: text data, collecting text information related to intangible cultural heritage, covering literature, historical documents, research papers, project introductions, expert interviews, etc. The focus should be on collecting texts describing the background, development, inheritance, etc. of intangible cultural heritage projects.
[0039] Image data, collecting high-definition images related to intangible cultural heritage, including complete photos of intangible cultural heritage projects, decomposition diagrams of technological steps, photos of inheritors and their skill demonstrations, on-site images of folk activities, etc. These images help to visually display the forms and details of intangible cultural heritage projects.
[0040] Audio / video data, collecting audio and video files related to intangible cultural heritage projects, including recordings of the production process of intangible cultural heritage skills, videos, audio of oral intangible cultural heritage inheritance, etc. These data provide support for capturing the dynamic information and voice interactions of intangible cultural heritage projects.
[0041] Structured data, collecting metadata of intangible cultural heritage projects provided by official agencies (such as UNESCO or the national intangible cultural heritage list), including the name, category, region, inheritor information, etc. of the project. In addition, inheritor profile information, such as personal background, skill inheritance process, social influence, etc., should also be collected.
[0042] Integrate the above data into a multi-modal intangible cultural heritage database for convenient subsequent data processing operations.
[0043] Then, clean and annotate the data in the multi-modal intangible cultural heritage database. Specifically, it includes:
[0044] Clean the collected text data, including removing duplicate paragraphs and standardizing terms (for example, different names for the same skill should be unified into one term) to ensure data consistency and accuracy.
[0045] Based on file names, timestamps, metadata, etc., conduct preliminary associations for data of different modalities. For example, use the file name of the image and the timestamp of the recording / video to align different audio-visual materials of the same event or skill.
[0046] It is also possible to invite inheritors of intangible cultural heritage, experts in the field of intangible cultural heritage, or relevant researchers to participate in the data annotation work, focusing on annotating the core entities (such as techniques, tools, and figures) in the text, images, and audio, as well as their mutual relationships (such as master-apprentice relationships, technological steps, and regional associations, etc.).
[0047] Integrate the multimodal dataset after data cleaning and annotation, including data in forms such as text, images, and audio, into the initial dataset.
[0048] Step 2: Process the initial dataset through the multimodal data cleaning and annotation model to obtain the final dataset and the cross-modal association matrix. The core goal of the multimodal data cleaning and annotation model is to optimize the quality of multimodal data and ensure annotation consistency for the initial dataset through modules such as denoising, modal alignment, annotation consistency, and association matrix constraint. By jointly optimizing these objective functions, the finally cleaned data can better serve the analysis and annotation tasks of intangible cultural heritage data.
[0049] Specifically, the objective function of the multimodal data cleaning and annotation model in this embodiment can be expressed by the following formula:
[0050] ,
[0051] Among them, is the initial text data matrix, is the initial image data matrix, is the initial audio data matrix, is the weight coefficient for controlling the denoising intensity, is the data denoising function, is the weight coefficient for controlling the optimization degree of modal alignment, is the cross-modal association optimization function, is the weight coefficient for controlling annotation consistency, is the annotation consistency constraint function, is the weight coefficient for controlling the sparsity of the association matrix, is the association matrix constraint function, is the cleaned text data matrix, is the cleaned image data matrix, is the cleaned audio data matrix, is the annotation consistency matrix, is the cross-modal association matrix.
[0052] Note that the objective function is a comprehensive loss function used to improve the performance of cross-modal data cleaning through multi-faceted optimization. This objective function ensures that the model can efficiently clean cross-modal data and maintain a balance in all aspects through the synergistic effects of multiple factors such as denoising, modal alignment, annotation consistency, and correlation matrix constraints. Specifically, the purpose of this mathematical expression is to minimize the objective function, where, is the initial text data matrix, represents a matrix, where each row represents the word vector of a document. represents the number of samples, represents the dimension of the text data, usually obtained through feature extraction methods such as TF-IDF or BERT embedding. is the initial image data matrix, represents a matrix, similar to the text data matrix, where each row is the feature vector of an image, represents the dimension of each image feature vector, which can be extracted by a deep learning model. is the initial audio data matrix, , similarly, each row represents the feature of an audio segment, is the dimension of the audio feature. Common features include MFCC (Mel Frequency Cepstral Coefficients) and mean pooling. And the cross-modal correlation matrix , This is a binary matrix used to represent the relationship between different modalities, where indicates that there is an actual association between sample i and sample j in different modalities. The annotation consistency matrix , This is a matrix used to represent the annotation probability of each sample, k is the number of classes, and each row represents the probability that the sample belongs to k classes. Finally, for convenience of representation, the initial data matrices , , can be integrated into an initial data matrix , and at the same time, the cleaned data matrices , , are also integrated into a set of cleaned data matrices used to represent the set of text, image, or audio data matrices.
[0053] The design idea of the objective function in this embodiment is to simultaneously optimize multiple aspects of the model through multiple constraints and losses, ensuring that data cleaning for different modalities can be completed efficiently and accurately. Specifically, the objective function includes the following main optimization objectives: 1. Denoising: First, the model needs to remove noise from the data, which helps improve the quality and robustness of feature representation. Through the denoising loss, the model can avoid interference from noisy samples and improve the overall data cleaning effect. 2. Cross-modal alignment: Second, a key task in multi-modal data cleaning is to align features of different modalities so that samples with the same semantics have similar representations across different modalities. Through the cross-modal alignment loss, the model can optimize the relationship between modalities. 3. Annotation consistency: Annotation consistency ensures that the same class of samples under different modalities maintain consistent labels, thus avoiding classification errors caused by inconsistent annotations. 4. Association matrix constraint: Finally, the sparsity constraint of the association matrix encourages the model to establish associations only between highly similar sample pairs, thus avoiding the model overfitting to irrelevant samples. Then, by adjusting the weight coefficients of different functions, the relative importance of each objective in the overall optimization can be controlled.
[0054] Specifically, in this embodiment, the data denoising function in the objective function aims to remove noise and ensure that the model can correctly extract effective features from the data. The noise here may be caused by irrelevant samples, incorrect annotations, or other factors. Processing through this denoising function helps improve the robustness of the model. The data denoising function reduces the impact of noise by penalizing inconsistencies or incorrect features in the model, thereby improving the quality of feature representation.
[0055] This function can be represented by the following formula:
[0056] ,
[0057] where, is the set of initial data matrices, is the set of data matrices after cleaning, is the Frobenius norm, also known as the Euclidean norm of the matrix, used to measure the difference between two matrices. In the above formula represents the sum of the squares of the differences between the original data and the denoised data, used to measure the effect of the denoising process on the data matrix. If the denoising effect is better, this difference is smaller and the loss value is lower. is the nuclear norm, used to constrain the rank of the matrix and encourage the matrix to have a low-rank structure. In the above formula is used to force the denoised matrix to maintain a low rank and remove redundant information, meaning that if the rank of the set of data matrices after cleaning is lower, it will be more compact and contain fewer redundant terms, helping to clean up invalid or duplicate data. is a hyperparameter for denoising intensity, used to control the denoising intensity, which balances the weights of the Frobenius norm term and the nuclear norm term in the objective function. When is larger, the denoising effect will be stronger, and the low-rank constraint of the matrix will be more stringent. That is to say, at this time will be closer to the low-rank matrix, thus removing more noise. If is smaller, the denoising effect is weaker, and the matrix may retain more original features, thus retaining some redundant information.
[0058] It should be noted that the objective of the above formula is to optimize the set of cleaned data matrices by minimizing the data denoising loss, so as to remove redundant information and noise. The above formula combines the Frobenius norm and the nuclear norm, enabling the model to retain the effective features of the data while removing the irrelevant and redundant parts. The hyperparameter is used to control the denoising intensity. Through such an optimization process, a cleaner and lower-rank multi-modal data set can be obtained, which is helpful for subsequent cross-modal association and annotation consistency optimization.
[0059] The cross-modal association optimization function in the objective function is used to optimize the alignment of cross-modal features. Its goal is to make the features of different modalities (such as images and texts) as close as possible in a shared embedding space, so as to improve the performance of the model in cross-modal tasks. The modal alignment loss function can measure the similarity or alignment degree between different modalities, forcing the modal features with the same semantics to be closer, while the features with different semantics are as far away as possible.
[0060] This function can be represented by the following formula:
[0061]
[0062] where is a positive sample pair (i.e., a sample pair with an actual association), and these sample pairs are usually pre-determined in some way (such as file name, timestamp, label, etc.), meaning that these two samples should be actually related cross-modally (for example, the text description and video clip of the same event); is a negative sample pair (i.e., a sample pair without an actual association), and these sample pairs are obtained by random sampling, aiming to help the model learn the differences between different modalities and prevent the model from misidentifying irrelevant samples as relevant. is the feature vector of sample i, is the feature vector of sample j, is the cosine similarity between the feature vector of sample i and the feature vector of sample j, used to measure the similarity between two samples in the feature space, and the value range is [-1, 1]. The closer the value is to 1, the more similar the two samples are, and the closer the value is to -1, the less similar they are. is the similarity threshold, which is used to control whether two samples are considered similar. It is usually a small positive value, such as 0.5, and is used to distinguish the cosine similarity between positive sample pairs and negative sample pairs. If the cosine similarity is greater than , then these two samples are considered related; if the cosine similarity is less than , then they are considered not associated. is the association matrix value, indicating whether there is an actual association between sample i and sample j. If , it means that sample i and sample j are related (for example, they describe the same event or object). If the two are not related, then .
[0063] It should be noted that the goal of the above formula is to enhance the cross-modal association ability of the model by optimizing the similarity between cross-modal samples. The formula contains two parts: the optimization of positive sample pairs and the optimization of negative sample pairs. Among them, the optimization of positive sample pairs is achieved by for each pair of , that is, each pair of predefined related sample pairs, making the cosine similarity of their feature vectors in the feature space close to 1, that is, they should be as similar as possible. By minimizing the difference between the cosine similarity and the association matrix value, it is expected that the cosine similarity is close to 1; if the similarity is less than 1, the function will increase, prompting the model to adjust the feature vectors to make the similarity greater.
[0064] The optimization of negative sample pairs is achieved by for each pair of , that is, each pair of negative sample pairs without actual association, making the cosine similarity of their feature vectors as close to 0 or lower, indicating that they should be as far away as possible in the feature space. is used to penalize those negative sample pairs whose similarity is greater than the threshold, that is, if the cosine similarity of two unassociated samples is higher than the threshold, a loss will be generated, prompting the model to adjust its feature representation to make the negative sample pairs more separated in the feature space.
[0065] That is to say, for positive sample pairs in the above formula, their cosine similarity is maximized so that they are closer in the feature space. For negative sample pairs, their cosine similarity is minimized to avoid them being misjudged as related. Through the function designed in this way, the model can learn more accurate cross-modal feature representations, so that data from different modalities can be effectively corresponding and aligned in the shared feature space.
[0066] The annotation consistency constraint function in the objective function is used to ensure that the model can maintain the annotation consistency of samples during the learning process. That is to say, the model should ensure that the same class of samples in different modalities have consistent annotation information. The annotation consistency function ensures that the same class of samples (such as images and texts) in different modalities have the same or similar label representations by penalizing inconsistent annotation information, thereby improving the accuracy of annotation between different modalities.
[0067] This function can be represented by the following formula:
[0068] ,
[0069] where, is the set of annotated samples, representing the set of samples with known labels, and the labels of these samples are annotated by experts. Usually, the number of these samples is small, but their labels are very accurate, and the model can use the labels of these samples to guide the training. is the annotation consistency probability of sample i, that is, the label prediction probability of the model for sample i. During the optimization process, the model will learn how to adjust to make it closer to the true label , and ensure that similar samples have similar labels by propagating the consistency of labels; is the true label of sample i, coming from expert annotation; is a hyperparameter that controls the strength of label propagation. If is larger, the influence of label propagation is stronger, and the label consistency constraint between similar samples will be more stringent; if is smaller, the influence of the label consistency constraint is weaker, and the model may rely more on the annotated samples. is the Laplacian matrix, which is constructed based on the similarity between samples and is usually used to represent the relationship between nodes in a graph structure. Specifically, the Laplacian matrix is used to describe the similarity structure between samples and control the label consistency between similar samples. is the trace operation of the matrix, which calculates the sum of the diagonal elements of , representing a global consistency of label propagation, is the transpose matrix of the annotation consistency matrix, is the matrix transpose symbol.
[0070] It should be noted that the objective of the above formula is to optimize the model by introducing label consistency constraints, so that the model can reasonably propagate label information, so that the labels of unannotated samples can be reasonably inferred, while keeping the labels of annotated samples unchanged. In the above formula, the first term represents the label consistency constraint of annotated samples; for each annotated sample i, its annotation consistency probability It should be as close as possible to the true labels provided by experts . By minimizing the difference between the prediction result of the model is made consistent with the actual labels. The second term is the constraint term for label propagation, which is used to ensure the consistency of labels between similar samples. This term uses the Laplacian matrix to represent the similarity structure between samples, and the annotation consistency matrix contains the annotation consistency probability for each sample. By forcing similar samples to have similar labels in the feature space, the optimization process will make the labels of similar samples develop in the same direction. That is to say, this formula guides the model to learn the label relationship across samples through annotation consistency constraints, ensuring that similar samples have similar labels while ensuring the accuracy of the labels of the labeled samples. Through such constraints, the model can not only make full use of the limited labeled samples, but also infer the labels of unlabeled samples through label propagation.
[0071] The correlation matrix constraint function in the objective function sparsifies the correlation matrix, so that there are non-zero correlations only between cross-modal features with sufficient similarity. This correlation matrix constraint function is constrained by the L1 norm, which prompts most of the elements in the correlation matrix to be close to zero, thereby reducing the connections between irrelevant or low-similarity samples and maintaining the efficiency of the model.
[0072] This function can be expressed by the following formula:
[0073] ,
[0074] ,
[0075] where is the cross-modal correlation matrix. is the L1 norm of the cross-modal correlation matrix, that is, the sum of the absolute values of all elements in the cross-modal correlation matrix. The role of the L1 norm is to promote the sparsity of the cross-modal correlation matrix, making most of the elements in the correlation matrix close to zero. By minimizing the model is encouraged to retain only a few important (high-similarity) correlations, while other low-similarity correlations will be suppressed; s.t. is the abbreviation of subject to, indicating that this function is restricted by the cross-modal feature similarity between sample i and sample j.
[0076] It should be noted that the above formula optimizes the cross-modal learning model by promoting the sparsity of the cross-modal correlation matrix $\mathbf{C}$ and ensuring that only the correlations between samples with higher similarity are retained through the threshold . It helps to improve the efficiency of learning, reduces the computational amount by only focusing on relevant sample pairs, and avoids invalid correlations. In addition, by adjusting , it is possible to control the sensitivity of the model to similarity, thereby affecting the performance of the model in cross-modal alignment.
[0077] Step 3: Based on the dataset and the cross-modal association matrix, construct a unified domain knowledge model to structurally model the concepts, entities, and their relationships related to intangible cultural heritage, and form a semantic expression and machine-understandable knowledge organization form that supports multi-modal data.
[0078] The data matrix set obtained after the multi-modal data cleaning and annotation model processes the initial dataset , the annotation consistency matrix and the cross-modal association matrix . It should be noted that the data matrix set contains the multi-modal data after cleaning; the annotation consistency matrix records the class labels of the samples and their corresponding confidence levels; the cross-modal association matrix records the corresponding relationships between different modal samples.
[0079] In the process of constructing the unified domain knowledge model, first, based on the annotation consistency matrix, define the core ontology classes. Statistically analyze the classes in the annotation consistency matrix, calculate the total confidence level of each class. If the overall confidence level of a certain class is significantly higher than that of other classes, then define this class as the core ontology class. Thus, construct the hierarchical structure between classes and clarify the parent-child class relationships between different classes (for example, the class of "Shu embroidery" can be defined as: traditional textile craft → embroidery → Shu embroidery).
[0080] Then, based on the cross-modal associations and knowledge rules of the cross-modal association matrix, implement relationship definition. Through the information of the cross-modal data pairs in the cross-modal association matrix, infer the relationships between different entities and achieve alignment between different modalities. For example, if text i is "Shu embroidery production process" and image j is "Shu embroidery production step diagram", then infer that the relationship is a relevant relationship, match the predefined relationship pattern in the knowledge graph to establish the rule relationship in the knowledge graph, thereby ensuring the reasoning ability of the knowledge graph for intangible cultural heritage data.
[0081] Finally, according to the data in the data matrix set, extract text knowledge. The extraction of text knowledge includes entity extraction and relationship extraction. In this embodiment, entity extraction uses high-confidence labels (from the annotation consistency matrix) as supervision signals and is implemented through the BiLSTM+CRF model. Relationship extraction strengthens the confidence level of relationship types by enhanced dependency parsing and combining the cosine similarity of text and images.
[0082] At the same time, taking images and audio as knowledge supplements, based on the alignment of images and text, using the information in the cross-modal correlation matrix to find the text corresponding to each image, so as to generate annotations for the images (for example, when the image is described as "Sichuan embroidery", the text description "the production process of Sichuan embroidery" will be automatically assigned to the image entity); then using an object detection model such as Mask R-CNN to identify specific regions in the image and match them through the similarity with text entities, so as to generate relationship triples. And perform automatic speech recognition (ASR) on the audio data, extract entities from the generated text, and use voiceprint recognition to determine whether the speaker in the audio matches the identity of the inheritor of intangible cultural heritage, so as to verify the authenticity and relevance of the recording.
[0083] Considering that there may sometimes be conflicts in the descriptions of the same entity in different modalities, in order to solve the conflicts in the descriptions of the same entity in different modalities, in this embodiment, it is also possible to enhance cross-modal consistency by mapping the entities of different modalities into a unified space, thereby solving the conflict problem.
[0084] Specifically, in this embodiment, the conflict loss function can be used to force the representations of the same entity in different modalities to be close. The conflict loss function can be expressed by the following formula:
[0085] ,
[0086] Where, is the conflict loss function, is the set of cross-modal data pairs of the same entity, is the data pair of the m-th modality of the same entity, is the data pair of the n-th modality of the same entity, is the input data of the m-th modality, is the input data of the n-th modality, is the embedding mapping function of the m-th modality, is the embedding mapping function of the n-th modality.
[0087] It should be noted that the embedding mapping function is a mapping function that projects the original data (such as text, images, etc.) into a unified embedding space. The features of text and images can be mapped to a shared vector space through a neural network. In this space, the different modality representations of the same entity should be as close as possible. In this way, it helps with multi-modal understanding, enabling the model to understand and reason about various forms of data (such as understanding the relationship between text and images). is the square of the Euclidean distance, representing the difference between the embeddings of the corresponding entities in modality n and modality n. The smaller this value is, the closer the representations of the same entity in different modalities are. And in contains all the cross-modal data pairs of the same entity; in other words, in It contains all different modality pairs corresponding to the same entity (such as the same "intangible cultural heritage project"), for example, text descriptions, images, audio, etc. To sum up, the core of this conflict loss function is to minimize the Euclidean distance between different modality representations (such as text and image) of the same entity by embedding them in a shared vector space. For the same entity, the representations of different modalities should be clustered in the same embedding space rather than scattered. By minimizing this conflict loss function, the relationship between multimodal data is enhanced, thus helping the model better understand and reason about the complex knowledge in the field of intangible cultural heritage.
[0088] Step 4: In the knowledge organization form output in Step 3, it contains a complete ontology of the intangible cultural heritage field (including classes, relationships, constraints, etc.) and a set of cross-modal knowledge triples (entities and relationships of text, image, audio). Use these contents in the output knowledge organization form to construct a structured multimodal knowledge graph to achieve applications such as querying, reasoning, and visualization of intangible cultural heritage.
[0089] Embodiment 2
[0090] In this embodiment, a multimodal knowledge graph construction system for intangible cultural heritage based on big data is disclosed. The system includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, it implements the method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data as described in Embodiment 1.
[0091] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for constructing a multi-modal knowledge graph of intangible cultural heritage based on big data, characterized in that, The multi-modal knowledge graph construction method includes: S100. Collect intangible cultural heritage data in different modalities, preliminarily clean and annotate the collected data, establish a preliminary association between data in different modalities, and obtain a preliminary multi-modal data set; S200. Process the preliminary multi-modal data set through a multi-modal data cleaning and annotation model to obtain a final multi-modal data set, an annotation consistency matrix, and a cross-modal association matrix; S300. Construct a unified domain knowledge model, perform structured modeling on intangible cultural heritage-related concepts, entities, and their relationships, and form a semantic expression and knowledge organization form that supports multi-modal data; S400. Construct a structured multi-modal knowledge graph based on the data in the knowledge organization form output by S300.
2. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 1, wherein The intangible cultural heritage data in different modalities includes: Text data related to intangible cultural heritage, including: texts, literature materials, historical documents, research papers, project introductions, and expert interviews describing the background, development, and inheritance of intangible cultural heritage projects; Image data related to intangible cultural heritage, including: complete photos of intangible cultural heritage projects, decomposition diagrams of technological steps, photos of inheritors and their skill demonstrations, and on-site images of folk activities; Audio and video data related to intangible cultural heritage, including: recordings and videos of the production process of intangible cultural heritage skills, audio and video of oral intangible cultural heritage inheritance.
3. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 1, wherein The preliminary cleaning and annotation include: Removing duplicate paragraphs from the collected text data and standardizing terms; Based on file names, timestamps, metadata, etc., making a preliminary association between data in different modalities, Annotating the core entities and their mutual relationships in text, image, audio, and video data.
4. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 1, wherein, The multi-modal data cleaning and annotation model ensures that the data cleaning in different modalities can be completed efficiently and accurately through multiple constraints and losses, The multi-modal data cleaning and annotation model includes: S210. Denoising processing, removing noise from the data, and avoiding interference from noise samples through denoising loss to improve the overall data cleaning effect; S220. Cross-modal alignment, aligning the features of different modalities so that samples with the same semantics have similar representations between different modalities and optimizing the relationship between modalities; S230. Annotation consistency, ensuring that the same type of samples in different modalities have consistent labels and avoiding classification errors caused by inconsistent annotations; S240. Association matrix constraint, prompting the model to establish associations only between highly similar sample pairs through sparsity constraints, avoiding the model from overfitting irrelevant samples, and controlling the relative importance of each objective in the overall optimization by adjusting the weight coefficients of different functions.
5. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 4, wherein The objective function of the multi-modal data cleaning and annotation model is represented by the following formula: , Among them, is the final multi-modal dataset, including the cleaned text data matrix, the cleaned image data matrix, the cleaned audio data matrix, is the initial multi-modal dataset, including the initial text data matrix, the initial image data matrix, the initial audio data matrix, is the weight coefficient for controlling the denoising intensity, is the data denoising function, is the weight coefficient for controlling the optimization degree of modal alignment, is the cross-modal association optimization function, is the weight coefficient for controlling the annotation consistency, is the annotation consistency constraint function, is the weight coefficient for controlling the sparsity of the association matrix, is the association matrix constraint function, is the annotation consistency matrix, is the cross-modal association matrix.
6. The method for constructing a multi-modal knowledge graph of intangible cultural heritage based on big data according to claim 5, characterized in that The data denoising function is used to ensure that the model can correctly extract effective features from the data. The data denoising function reduces the noise impact by penalizing the inconsistencies or incorrect features in the model, thereby improving the quality of feature representation; The cross-modal correlation optimization function is used to optimize the alignment of cross-modal features. By making the features of different modalities as close as possible in a shared embedding space, it improves the performance of the model in cross-modal tasks; The annotation consistency constraint function is used to ensure the annotation consistency of samples during the learning process of the model, ensuring that the same class of samples in different modalities have consistent annotation information; The correlation matrix constraint function sparsifies the correlation matrix, so that there are non-zero correlations only between cross-modal features with sufficient similarity. The correlation matrix constraint function uses the L1 norm constraint to encourage most of the elements in the correlation matrix to be close to zero, thereby reducing the connections between irrelevant or low-similarity samples and maintaining the efficiency of the model.
7. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 1, characterized in that, The construction of the unified domain knowledge model includes the following sub-steps: S310. Define the core ontology class based on the annotation consistency matrix, count the categories in the annotation consistency matrix, and calculate the total confidence of each category. If the overall confidence of a certain category is significantly higher than that of other categories, then define this category as the core ontology class; S320. Implement relationship definition based on the cross-modal correlation and knowledge rules of the cross-modal correlation matrix. Through the information of cross-modal data pairs in the cross-modal correlation matrix, infer the relationships between different entities and achieve alignment between different modalities; S330. Extract text knowledge according to the data in the data matrix set.
8. The method for constructing a multimodal knowledge graph of intangible cultural heritage based on big data according to claim 7, wherein The extraction of text knowledge includes: entity extraction and relationship extraction, where Entity extraction uses the high-confidence labels of the annotation consistency matrix as supervision signals and is implemented through the BiLSTM+CRF model; Relationship extraction uses enhanced dependency parsing, combined with the cosine similarity of text and images, to extract and strengthen the confidence of relationship types from the text.
9. The method for constructing a multi-modal knowledge graph of intangible cultural heritage based on big data according to claim 8, wherein The extraction of text knowledge also includes: Using images and audio as knowledge supplements, based on the alignment of images and text, using the information in the cross-modal correlation matrix, find the text corresponding to each image, so as to generate annotations for the images; Then use the object detection model to identify the specific regions in the images and match them through the similarity with text entities, so as to generate relationship triples.
10. A multi-modal knowledge graph construction system for intangible cultural heritage based on big data, characterized in that, The big-data-based intangible cultural heritage multi-modal knowledge graph construction system includes: A processor; A memory storing a computer program, which when executed by the processor, implements the big-data-based intangible cultural heritage multi-modal knowledge graph construction method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent building knowledge extraction model and method based on multiple modes
CN116737897A
Generating hypothesis candidates associated with an incomplete knowledge graph
US20220156599A1
Cited By
Multi-modal data extraction method and system based on knowledge graph
CN120632163A
Historical literature version knowledge ontology dynamic collaborative construction method and system
CN120893547A
A historical document version knowledge ontology dynamic collaborative construction method and system
CN120893547B
Hierarchical data classification processing and deep correlation analysis system and method based on knowledge graph
CN120974343A
Space-time knowledge graph construction method for regional historical and cultural heritage
CN121365716A