Knowledge graph construction method for multi-modal data
Through a feature extraction method based on entropy weight allocation and cross-modal consistency metric, combined with graph convolution network and search environment adaptation factor Gdf, data redundancy and context perception problems in the knowledge graph are solved, and efficient fusion and accurate retrieval of multimodal data are achieved.
Patent Information
- Application Number
- CN202510294050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-11
AI Technical Summary
The existing knowledge graph construction methods have problems such as data redundancy and repetition, lack of context perception, and difficulty in multimodal data processing, especially in multi-source data integration and cross-modal fusion.
A feature extraction method based on entropy weight allocation and a cross-modal consistency measurement method are used, and semantic alignment is combined with graph convolution network to build a non-redundant knowledge graph; by calculating the search environment adaptation factor Gdf and the cross-modal matching degree index Mcf, the knowledge graph retrieval strategy is dynamically adjusted to realize context-awareness and multimodal data fusion.
Effectively reduce data redundancy, improve query efficiency and data integration quality, optimize search adaptability and the fusion ability of multimodal data, and improve the organizational rationality and retrieval accuracy of knowledge graphs.
Smart Images

Figure CN120296652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and specifically to a method for constructing a knowledge graph of multimodal data. Background Art
[0002] The method for constructing a knowledge graph originated from the combination of artificial intelligence and semantic web technologies, and was initially used to process and represent a large amount of structured and unstructured information. In the early days, the construction of knowledge graphs relied on manual annotation and relation extraction, mainly focusing on the professional knowledge system within the field, which had limitations. With the development of big data and machine learning technologies, the construction of knowledge graphs has gradually shifted from rule-based manual construction to automated semantic analysis and data mining.
[0003] Currently, the application of knowledge graphs has been widely involved in fields such as natural language processing, recommendation systems, and intelligent search, and has become one of the key technologies in the field of artificial intelligence; in particular, the method for constructing a knowledge graph applied to intelligent search has evolved from being rule-driven in the initial stage to being driven by deep learning, gradually realizing multimodal fusion, real-time update, and semantic enhancement, providing strong support for the intelligent processing of data and the efficient organization of knowledge.
[0004] However, the method for constructing a knowledge graph applied to intelligent search still has some drawbacks in the prior art, which are mainly reflected in the following aspects:
[0005] 1. Data redundancy and duplication problems: In the process of constructing large-scale knowledge graphs, the phenomena of data redundancy and duplication are relatively common, especially in the integration of multi-source data. For example, the same entity may obtain different descriptions from different data sources, resulting in redundant information in the knowledge graph, affecting query efficiency and data storage.
[0006] 2. Lack of context awareness: Existing knowledge graph construction methods often lack context awareness and cannot dynamically adjust the graph content according to the actual needs of search users. For example, users may have different requirements for the same entity in different search environments (such as academic search and shopping search), and existing graphs are difficult to optimize search results according to the context.
[0007] 3. Difficulty in processing multimodal data: Many existing knowledge graph construction methods mainly process structured data and are difficult to effectively integrate unstructured data such as images, voices, and videos. For example, in a search engine, when processing multimodal data (such as image search), the graph construction lacks an effective cross-modal fusion mechanism, restricting the comprehensiveness of the search. Summary of the Invention
[0008] In view of the deficiencies of the prior art, the present invention provides a method for constructing a knowledge graph of multimodal data, which solves the technical drawbacks mentioned in the background art.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for constructing a knowledge graph of multimodal data, comprising the following steps:
[0010] Step 1: Adopt a feature extraction method based on entropy weight distribution to vectorize the features of text, image, speech, and video data, construct an initial knowledge entity set, and assign unique identification numbers and multimodal space mappings to the knowledge entities;
[0011] Step 2: Based on the cross-modal consistency measurement method, match similar knowledge entities in different data sources, and adopt the information entropy de-duplication strategy to merge redundant data to construct a non-redundant knowledge graph. At the same time, use the graph convolutional network for semantic alignment;
[0012] Step 3: Calculate and evaluate the search environment adaptation factor Gdf, construct a context-aware knowledge graph, and verify whether the search environment has matched enough context information;
[0013] Step 4: Adopt a cross-modal contrast learning method to extract cross-modal matching related data, calculate and evaluate the cross-modal matching degree index Mcf, construct a unified high-dimensional knowledge representation, and establish a cross-modal data fusion mechanism;
[0014] Step 5: Calculate the multimodal adaptation degree index Qms based on the context-aware knowledge graph, and adjust the knowledge graph retrieval strategy through the dynamic index scheduling mechanism; finally, evaluate and analyze the multimodal adaptation degree index Qms, and adjust the knowledge query path in real time.
[0015] Preferably, Step 1 specifically includes:
[0016] First, obtain multimodal related data including text, image, speech, and video, and based on the feature extraction method of entropy weight distribution, decompose the features of multimodal related data of different modalities, extract the key attributes of each modality data in the multimodal related data, and convert the key attributes into high-dimensional feature vectors; Subsequently, construct an initial knowledge entity set, and based on the deep semantic matching mechanism, calculate the semantic similarity of the high-dimensional feature vectors, and then identify the knowledge entities with high similarity; On this basis, construct a hierarchical clustering model based on the embedding vector calculation of the graph neural network; At the same time, use the hierarchical clustering model to adaptively group the matched knowledge entities, calculate the similarity threshold according to the semantic relevance of the knowledge entities, and then dynamically adjust the clustering structure; Then, assign a unique identification number to each knowledge entity, and use the multimodal space mapping method to map and fuse different modality knowledge in a unified space according to the semantic relationship between entities and the association degree between modalities by establishing cross-modal data correspondence relationships.
[0017] Preferably, Step 2 specifically includes:
[0018] First, obtain knowledge entities from different data sources, and based on cross-modal consistency measurement methods, extract features of knowledge entities including text, images, speech, and videos. Calculate the cosine similarity, Euclidean distance measure, and embedding vector similarity calculation method based on deep neural networks to match the same or similar knowledge entities, and evaluate the confidence of the matching results. Then, based on the deduplication strategy of information entropy, calculate the information entropy of the feature distribution of similar knowledge entities, and according to the information gain threshold, perform knowledge entity merging processing. Among them, knowledge entities with low information gain are merged first, and knowledge entities with high information gain are stored and optimized, and then construct the deduplicated knowledge graph data structure. Subsequently, for semantic deviations caused by different sources, adopt a semantic alignment method based on graph convolutional networks to construct graph structure connections between knowledge entities, and based on the node embedding representation method, calculate the semantic offset of different data sources, adjust the representation vectors between entities, and adjust the representation consistency of knowledge entities from different sources to complete semantic alignment processing.
[0019] Preferably, step three specifically includes:
[0020] Collect and obtain user search behavior data in real time, including feature information such as search keywords, click frequency, stay time, and interaction paths. Then, obtain data related to search environment adaptation in real time by constructing a user behavior prediction model. Extract the search intention deviation Sip, knowledge entity access frequency Kaf, context association stability Cas, and historical search matching degree Hsm from the data related to search environment adaptation, and after dimensionless processing, calculate and obtain the environment adaptation factor Gdf through the following formula;
[0021]
[0022] Preferably, step three specifically further includes:
[0023] Adopt a reinforcement learning method to construct an optimization model for the association weights of knowledge entities, and use the user's search habits, click feedback, and search result preferences as state inputs. Construct a knowledge entity association matrix, calculate the jump probability of the user between knowledge entities using the Markov decision process, and based on the reward feedback mechanism of reinforcement learning, dynamically adjust the association weights between knowledge entities. Among them, the state transition probability is determined by the search history path and entity access frequency. During the training process of reinforcement learning, continuously adjust the association weights between knowledge entities based on policy gradient optimization, and combine the optimal policy selection mechanism to optimize the ranking of knowledge entities and improve the search matching degree at the same time;
[0024] Finally, establish a personalized knowledge entity index based on user group characteristics, continuously optimize the knowledge graph structure using an incremental data update mechanism, and adjust the search strategy based on a dynamic index scheduling method.
[0025] Preferably, step three specifically further includes:
[0026] Preset a search environment adaptation threshold H1, compare and evaluate it with the search environment adaptation factor Gdf, and determine whether the current search environment has met the context matching requirements. The specific evaluation content is as follows:
[0027] If the search environment adaptation factor Gdf ≥ the search environment adaptation threshold H1, it is determined that the current search environment has matched sufficient context information. At this time, extract the data related to cross-modal matching.
[0028] If the search environment adaptation factor Gdf < the search environment adaptation threshold H1, it is determined that the current search environment has not matched sufficient context information. At this time, dynamically adjust the context awareness weight parameter and update the user search behavior prediction model.
[0029] Preferably, step four specifically includes:
[0030] Collect multi-modal related data including text, images, voice, and video in real-time, extract features from the related data of each modality, use the cross-modal contrast learning method to construct a cross-modal representation space, and extract the embedding vectors of the multi-modal related data; on this basis, based on the calculation of the feature similarity between modalities, construct a modality contrast loss function, and use the contrast learning mechanism to adjust the representation vectors of different modality data.
[0031] During the process, gradually collect the inter-modal feature similarity Ma, semantic consistency measure Mb, cross-modal contrast loss value Mc, and multi-scale feature alignment error Md respectively, and after dimensionless processing, calculate and obtain the cross-modal matching degree index Mcf through the following formula:
[0032]
[0033] Preferably, step four specifically further includes:
[0034] Preset a cross-modal matching degree threshold M and compare and evaluate it with the cross-modal matching degree index Mcf, and dynamically adjust the weight allocation of different modality data; the specific evaluation content is as follows:
[0035] When the cross-modal matching degree index Mcf ≥ the cross-modal matching degree threshold M, it means that the current cross-modal matching degree meets the requirements, maintain the weight allocation of the current modality data, and enter the subsequent index optimization stage.
[0036] When the cross-modal matching degree index Mcf < the cross-modal matching degree threshold M, the matching degree of the current modal data does not meet the requirements. At this time, the weights of each modal data are dynamically adjusted, including preferentially adjusting the weight with the lowest matching degree of the modal data and iterative adjustment.
[0037] Finally, a unified high-dimensional knowledge representation method is adopted to map different modal data to a shared vector space, and a cross-modal data fusion mechanism is constructed based on the modal feature relationship; during the data fusion process, the knowledge representation between modalities is adjusted based on the semantic alignment model, and according to the correlation characteristics between different modal data, cross-modal feature alignment operations are performed to adjust the feature expression methods of each modal data, and cross-modal association modeling operations are performed based on the fused knowledge representation.
[0038] Preferably, step five includes:
[0039] First, obtain the knowledge entities and their associated information in the context-aware knowledge graph, and collect real-time multi-modal data such as text, images, voices, and videos involved in the user query; subsequently, based on the multi-modal feature extraction method, vectorize different modal data and construct a cross-modal embedding representation; use the knowledge graph structure relationship to calculate the association strength between knowledge entities, and optimize the semantic consistency of different modal data based on the modal alignment method; finally, through the modal feature fusion technology, generate a unified multi-modal knowledge representation to provide data support for subsequent multi-modal fitness calculation and knowledge retrieval optimization.
[0040] Collect the knowledge entities and the associated information of the knowledge entities in the context-aware knowledge graph in real time, and after feature extraction of the modal data including text, images, voices, and videos, obtain the entity association degree Eri, the modal consistency Mcc, and the search preference degree Upp, and after dimensionless processing, calculate and obtain the multi-modal fitness index Qms according to the following formula:
[0041]
[0042] In the formula, i represents the index variable, X i represents the set of lower-level parameters required to calculate the multi-modal fitness index Qms, where the entity association degree Eri is X1, the modal consistency Mcc is X2, and the search preference degree Upp is X3;
[0043] exp represents the natural exponential function, which is used to compress data and convert the logarithmic mean back to the original unit; ln represents the natural logarithmic function, which is used to restore data and maintain a smooth distribution.
[0044] Preferably, step five specifically includes:
[0045] Preset the multi-modal adaptation threshold Q, compare and evaluate it with the multi-modal adaptation degree index Qms, and adjust the knowledge query path in real time; the specific evaluation content is as follows:
[0046] When the multi-modal adaptation degree index Qms ≥ the multi-modal adaptation threshold Q: the current multi-modal matching degree meets the knowledge retrieval requirements, maintain the existing query path, and execute the knowledge retrieval operation.
[0047] When the multi-modal adaptation degree index Qms < the multi-modal adaptation threshold Q: the current multi-modal matching degree does not meet the knowledge retrieval requirements, and the matching degree is insufficient; at this time, adjust the query path and the cross-modal data matching effect.
[0048] The present invention provides a method for constructing a knowledge graph of multi-modal data. It has the following beneficial effects:
[0049] (1) For the problems of data redundancy and repetition in the method for constructing a knowledge graph of multi-modal data, by using a cross-modal consistency measurement method to match similar knowledge entities in different data sources, and adopting an information entropy de-duplication strategy to merge redundant data, it ensures the construction of a non-redundant knowledge graph data structure; among them, a graph convolutional network is used for semantic alignment to construct a graph structure connection between knowledge entities, and based on the node embedding representation method, the semantic offset of different data sources is calculated, and the representation vectors between entities are adjusted, so as to optimize the representation consistency of knowledge entities from different sources; in addition, during the construction of the knowledge graph, a hierarchical clustering model is used to adaptively group the matched knowledge entities, and the similarity threshold is calculated according to the semantic relevance of the knowledge entities, so as to dynamically adjust the clustering structure, improve the organizational rationality of the knowledge graph, reduce the storage resource occupation caused by data redundancy, and improve the query efficiency and data integration quality.
[0050] (2) For the problem that the existing knowledge graph construction methods lack context awareness ability, the present invention calculates the search environment adaptation factor Gdf based on user search behavior data, including feature information such as search keywords, click frequency, stay time, and interaction path; among them, by calculating the search intention deviation degree Sip, the knowledge entity access frequency Kaf, the context association stability Cas, and the historical search matching degree Hsm, a context-aware knowledge graph is constructed, and it is verified whether the search environment has matched enough context information; when the search environment adaptation factor Gdf is lower than the preset search environment adaptation threshold H1, the context awareness weight parameter is dynamically adjusted, and the user search behavior prediction model is updated to optimize the search adaptability; in addition, a reinforcement learning method is used to construct a knowledge entity association weight optimization model, and the jump probability of the user between knowledge entities is calculated in combination with the Markov decision process, so that the ranking of knowledge entities dynamically adapts to the user's search preferences, and the context-based knowledge graph retrieval optimization is realized.
[0051] (3) The method for constructing a knowledge graph of multimodal data. In view of the difficulties in processing multimodal data in the existing knowledge graph construction methods, the present invention adopts a cross-modal contrast learning method to extract multimodal matching-related data and calculate the cross-modal matching degree index Mcf to ensure the fusion ability of different modal data. Among them, by collecting multimodal data including text, images, voice, and video in real time, and vectorizing each modal data based on a feature extraction method, the inter-modal feature similarity Ma, semantic consistency measure Mb, cross-modal contrast loss value Mc, and multi-scale feature alignment error Md are calculated respectively. After dimensionless processing, the cross-modal matching degree index Mcf is calculated using the logarithmic geometric mean. When the cross-modal matching degree index Mcf is lower than the preset cross-modal matching degree threshold M, the weights of each modal data are dynamically adjusted to optimize the matching effect of cross-modal data. Finally, through a unified high-dimensional knowledge representation method, different modal data are mapped to a shared vector space, and a cross-modal data fusion mechanism is constructed based on the modal feature relationship, enabling different modal data to be associated and modeled and knowledge retrieval optimized in the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flow chart of the steps of the method for constructing a knowledge graph of multimodal data according to the present invention;
[0053] Figure 2 It is a simulation experimental result diagram of the search environment adaptation factor Gdf of the method for constructing a knowledge graph of multimodal data according to the present invention;
[0054] Figure 3 It is a simulation experimental result diagram of the cross-modal matching degree index Mcf of the method for constructing a knowledge graph of multimodal data according to the present invention;
[0055] Figure 4 It is a simulation experimental result diagram of the multimodal adaptation degree index Qms of the method for constructing a knowledge graph of multimodal data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] Embodiment 1
[0058] Please refer to Figure 1 , the present invention provides a method for constructing a knowledge graph of multimodal data, including the following steps:
[0059] Step 1: Adopt a feature extraction method based on entropy weight distribution to vectorize the features of text, image, speech, and video data, construct an initial knowledge entity set, and assign unique identification numbers and multi-modal space mappings to the knowledge entities;
[0060] Step 2: Based on the cross-modal consistency measurement method, match similar knowledge entities in different data sources, and adopt the information entropy de-duplication strategy to merge redundant data to construct a non-redundant knowledge graph. At the same time, use the graph convolutional network for semantic alignment;
[0061] Step 3: Calculate and evaluate the search environment adaptation factor Gdf, construct a context-aware knowledge graph, and verify whether the search environment has matched enough context information;
[0062] Step 4: Adopt a cross-modal contrastive learning method to extract cross-modal matching related data, calculate and evaluate the cross-modal matching degree index Mcf, construct a unified high-dimensional knowledge representation, and establish a cross-modal data fusion mechanism;
[0063] Step 5: Calculate the multi-modal adaptation degree index Qms based on the context-aware knowledge graph, and adjust the knowledge graph retrieval strategy through the dynamic index scheduling mechanism; finally, evaluate and analyze the multi-modal adaptation degree index Qms, and adjust the knowledge query path in real time.
[0064] In this embodiment, in step one, through a feature extraction method based on entropy weight distribution, text, image, voice, and video data are vectorized, enabling the standardization of features of different modality data. An initial knowledge entity set is constructed, and combined with unique identification numbers and multimodal space mapping, the standardization of knowledge entity management and the rationality of data storage are improved; in step two, based on a cross-modal consistency measurement method, similar knowledge entities are matched in different data sources, and an information entropy de-duplication strategy is used to merge redundant data, reducing the redundancy problem caused by multiple data sources. At the same time, a graph convolutional network is used for semantic alignment to enhance the fusion consistency and usability of data from different sources in the knowledge graph; in step three, the search environment adaptation factor Gdf is calculated, and the current search environment adaptation degree is evaluated. A context-aware knowledge graph is constructed through the search intention deviation degree Sip, the knowledge entity access frequency Kaf, the context association stability Cas, and the historical search matching degree Hsm. At the same time, it is verified whether the search environment has matched enough context information to optimize the search adaptability of knowledge entities; in step four, a cross-modal contrast learning method is used to extract cross-modal matching related data, and the cross-modal matching degree index Mcf is calculated. Among them, based on the inter-modal feature similarity Ma, the semantic consistency measurement Mb, the cross-modal contrast loss value Mc, and the multi-scale feature alignment error Md, the cross-modal matching degree index Mcf is calculated, and a unified high-dimensional knowledge representation is constructed based on the calculation results to ensure the deep fusion of different modality data and enhance the accuracy and effectiveness of cross-modal knowledge expression; in step five, the multi-modal adaptation degree index Qms is calculated based on the context-aware knowledge graph, the knowledge graph retrieval strategy is optimized through a dynamic index scheduling mechanism, and it is evaluated and analyzed whether the current retrieval path is optimal according to the multi-modal adaptation degree index Qms. If the multi-modal adaptation degree index Qms is lower than the preset multi-modal adaptation threshold Q, the knowledge query path is adjusted to make it more in line with the matching logic between different modality data, ensuring the accuracy and retrieval efficiency of search results.
[0065] Embodiment 2
[0066] Step one specifically includes:
[0067] First, obtain multi-modal related data including text, images, voice, and video, and based on the feature extraction method of entropy weight allocation, perform feature decomposition on the multi-modal related data of different modalities, extract the key attributes of each modality data in the multi-modal related data, and convert the key attributes into high-dimensional feature vectors. Subsequently, construct an initial knowledge entity set, and based on the deep semantic matching mechanism, calculate the semantic similarity of the high-dimensional feature vectors, and then identify high-similarity knowledge entities. On this basis, construct a hierarchical clustering model based on the embedding vector calculation of the graph neural network. At the same time, use the hierarchical clustering model to adaptively group the matched knowledge entities, calculate the similarity threshold according to the semantic relevance of the knowledge entities, and then dynamically adjust the clustering structure. Then, assign a unique identification number to each knowledge entity, and use the multi-modal space mapping method to map and fuse different modality knowledge in a unified space according to the semantic relationship between entities and the association degree between modalities by establishing cross-modal data correspondence relationships.
[0068] Step 2 specifically includes:
[0069] First, obtain knowledge entities from different data sources, and based on the cross-modal consistency measurement method, extract features from the knowledge entities including text, images, voice, and video. Match the same or similar knowledge entities through cosine similarity calculation, Euclidean distance measurement, and embedding vector similarity calculation methods based on deep neural networks, and evaluate the confidence of the matching results. Then, based on the deduplication strategy of information entropy, calculate the information entropy of the feature distribution of similar knowledge entities, and perform knowledge entity merging processing according to the information gain threshold. Among them, knowledge entities with low information gain are merged first, and knowledge entities with high information gain are stored and optimized, and then construct the deduplicated knowledge graph data structure. Subsequently, for the semantic deviation caused by different sources, use the semantic alignment method based on the graph convolutional network to construct the graph structure connection between knowledge entities, calculate the semantic offset of different data sources based on the node embedding representation method, adjust the representation vectors between entities, and adjust the representation consistency of knowledge entities from different sources to complete the semantic alignment processing.
[0070] In this embodiment, the knowledge graph construction method of multimodal data improves the fusion capability of data of different modalities, and optimizes the organizational structure and retrieval efficiency of knowledge entities; Step 1, through the feature extraction method of entropy weight allocation, obtain multimodal related data including text, image, voice and video, and perform feature decomposition on different modal data, convert key attributes into high-dimensional feature vectors, ensure the unified representation of each modal feature, and use the deep semantic matching mechanism to perform high similarity recognition to improve the accuracy of knowledge entity matching, and construct a hierarchical clustering model through the embedding vector calculation of the graph neural network, so that knowledge entities can be adaptively grouped according to semantic relevance, and dynamically adjust the clustering structure in combination with the similarity threshold, optimize the hierarchical organization of the knowledge graph, and further use the multimodal space mapping method to map knowledge entities of different modalities to a unified space, so as to improve the cross-modal compatibility of the knowledge graph;
[0071] Step 2: Acquire knowledge entities from different data sources, and perform feature extraction on text, image, voice and video data based on a cross-modal consistency measurement method, match the same or similar knowledge entities using cosine similarity calculation, Euclidean distance measurement, and an embedding vector similarity calculation method based on a deep neural network, and screen high-confidence matching results based on confidence evaluation. At the same time, use an information entropy deduplication strategy to calculate the feature distribution of knowledge entities, and merge knowledge entities based on an information gain threshold, wherein knowledge entities with low information gain are merged first, and knowledge entities with high information gain are optimized for storage, thereby constructing a non-redundant knowledge graph data structure, and then use a semantic alignment method based on a graph convolutional network to construct a graph structure connection between knowledge entities, and calculate the semantic offsets of different data sources based on a node embedding representation method, thereby adjusting the representation vectors between entities, enhancing the representation consistency of knowledge entities from different sources, and ultimately improving the knowledge fusion effect across data sources and the global consistency of the knowledge graph, so that the method of the present invention can effectively reduce data redundancy, improve the expression accuracy of the knowledge graph, and optimize the search and matching effect.
[0072] Example 3
[0073] Step three specifically includes:
[0074] Collect and obtain user search behavior data in real time, including feature information such as search keywords, click frequency, dwell time, and interaction path, and then obtain search environment adaptation related data in real time by building a user behavior prediction model; extract the search intention deviation degree Sip, knowledge entity access frequency Kaf, context association stability degree Cas, and historical search matching degree Hsm from the search environment adaptation related data, and after dimensionless processing, calculate the environment adaptation factor Gdf through the following formula;
[0075]
[0076] Step 3 specifically further includes:
[0077] Build a knowledge entity association weight optimization model using reinforcement learning methods, and use the user's search habits, click feedback, and search result preferences as state inputs; construct a knowledge entity association matrix, calculate the jump probability of the user between knowledge entities using the Markov decision process, and dynamically adjust the association weights between knowledge entities based on the reward feedback mechanism of reinforcement learning. Among them, the state transition probability is determined by the search history path and the entity access frequency; during the training process of reinforcement learning, based on policy gradient optimization, continuously adjust the association weights between knowledge entities, and combine the optimal policy selection mechanism to optimize the ranking of knowledge entities and improve the search matching degree at the same time;
[0078] Finally, establish a personalized knowledge entity index based on user group characteristics, continuously optimize the knowledge graph structure using an incremental data update mechanism, and adjust the search strategy based on a dynamic index scheduling method.
[0079] Step 3 specifically further includes:
[0080] Preset a search environment adaptation threshold H1, compare and evaluate it with the search environment adaptation factor Gdf, and judge whether the current search environment has met the context matching requirements. The specific evaluation content is as follows:
[0081] If the search environment adaptation factor Gdf ≥ the search environment adaptation threshold H1, it is determined that the current search environment has matched sufficient context information, and at this time, extract relevant cross-modal matching data.
[0082] If the search environment adaptation factor Gdf < the search environment adaptation threshold H1, it is determined that the current search environment has not matched sufficient context information, and at this time, dynamically adjust the context awareness weight parameter and update the user search behavior prediction model.
[0083] In this embodiment, by collecting user search behavior data in real time, including feature information such as search keywords, click frequency, residence time, and interaction path, and obtaining search environment adaptation-related data based on the user behavior prediction model, the search environment adaptation factor Gdf is calculated to ensure that the search environment can dynamically perceive user needs; during this process, the search intention deviation degree Sip in the search environment adaptation-related data is extracted to measure the stability and jumpiness of the user's search target; the knowledge entity access frequency Kaf is extracted to evaluate the user's attention to specific knowledge entities within a certain period of time; the context association stability Cas is extracted to calculate the entity co-occurrence situation in the current search session to ensure that the knowledge graph can adapt to different search environments; the historical search matching degree Hsm is extracted to evaluate the similarity between the current search request and the user's past search behavior, thereby optimizing the search personalization degree;
[0084] The reinforcement learning method is used to construct an optimization model for the association weights of knowledge entities. The user's search habits, click feedback, and search result preferences are used as state inputs. The Markov decision process is used to calculate the jump probability of the user between knowledge entities, and the association weights between knowledge entities are dynamically adjusted based on the reward feedback mechanism of reinforcement learning. Among them, the state transition probability is determined by the search historical path and the knowledge entity access frequency Kaf to ensure that the search results can be adaptively optimized as the user's needs change;
[0085] During the training process of reinforcement learning, based on policy gradient optimization, the association weights between knowledge entities are continuously adjusted, and the sorting of knowledge entities is optimized by combining the optimal policy selection mechanism to improve the search matching degree and make the search results more in line with user needs; in addition, a personalized knowledge entity index is established based on user group characteristics, and the knowledge graph structure is dynamically optimized using the incremental data update mechanism. At the same time, the search strategy is adjusted based on the dynamic index scheduling method to ensure the real-time performance and accuracy of the retrieval results;
[0086] Finally, by setting the search environment adaptation threshold H1, comparing the calculated search environment adaptation factor Gdf, and performing a search environment adaptation degree evaluation. If the environment adaptation factor Gdf is higher than the environment adaptation threshold H1, it is determined that the current search environment has matched sufficient context information. At this time, the extraction of cross-modal matching-related data is performed to further optimize the fusion and retrieval path of cross-modal data; if the environment adaptation factor Gdf is lower than the environment adaptation threshold H1, it is determined that the current search environment has not matched sufficient context information. At this time, the context awareness weight parameter is dynamically adjusted, and the user search behavior prediction model is updated to enhance the adaptability of the search environment to user needs, so that the knowledge graph can maintain efficient, accurate, and personalized knowledge retrieval capabilities in a complex and changing search environment.
[0087] Embodiment 4
[0088] Step 4 specifically includes:
[0089] Collect multi-modal related data including text, images, voice, and video in real-time, extract features from the related data of each modality, construct a cross-modal representation space using the cross-modal contrast learning method, and extract the embedding vectors of the multi-modal related data; on this basis, based on the calculation of the feature similarity between modalities, construct a modality contrast loss function, and use the contrast learning mechanism to adjust the representation vectors of different modality data;
[0090] During the process, gradually collect the inter-modal feature similarity Ma, semantic consistency measure Mb, cross-modal contrast loss value Mc, and multi-scale feature alignment error Md respectively, and after dimensionless processing, calculate and obtain the cross-modal matching degree index Mcf through the following formula:
[0091]
[0092] Step 4 also specifically includes:
[0093] Compare and evaluate the preset cross-modal matching degree threshold M with the cross-modal matching degree index Mcf, and dynamically adjust the weight distribution of different modality data; the specific evaluation content is as follows:
[0094] When the cross-modal matching degree index Mcf ≥ the cross-modal matching degree threshold M, it means that the current cross-modal matching degree meets the requirements, keep the weight distribution of the current modality data, and enter the subsequent index optimization stage;
[0095] When the cross-modal matching degree index Mcf < the cross-modal matching degree threshold M, the current modality data matching degree does not meet the requirements. At this time, dynamically adjust the weights of each modality data, including preferentially adjusting the weight with the lowest modality data matching degree and iterative adjustment.
[0096] Finally, adopt the unified high-dimensional knowledge representation method to map different modality data to the shared vector space, and construct a cross-modal data fusion mechanism based on the modality feature relationship; during the data fusion process, adjust the knowledge representation between modalities based on the semantic alignment model, and perform cross-modal feature alignment operations according to the correlation characteristics between different modality data, adjust the feature expression methods of each modality data, and perform cross-modal association modeling operations based on the fused knowledge representation.
[0097] Step 5 includes:
[0098] First, obtain the knowledge entities and their associated information in the context-aware knowledge graph, and collect real-time multimodal data such as text, images, voice, and videos involved in the user's query. Subsequently, based on the multimodal feature extraction method, vectorize different modal data and construct a cross-modal embedding representation. Utilize the structural relationship of the knowledge graph to calculate the association strength between knowledge entities, and optimize the semantic consistency of different modal data based on the modal alignment method. Finally, through the modal feature fusion technology, generate a unified multimodal knowledge representation to provide data support for subsequent multimodal fitness calculation and knowledge retrieval optimization.
[0099] Collect the knowledge entities and the associated information of the knowledge entities in the context-aware knowledge graph in real time. After extracting the features of the modal data including text, images, voice, and videos, obtain the entity association degree Eri, the modal consistency Mcc, and the search preference degree Upp, and perform dimensionless processing. Then, calculate and obtain the multimodal fitness index Qms by combining the following formula:
[0100]
[0101] In the formula, i represents the index variable, and X i represents the set of subordinate parameters required to calculate the multimodal fitness index Qms. Among them, the entity association degree Eri is X1, the modal consistency Mcc is X2, and the search preference degree Upp is X3;
[0102] exp represents the natural exponential function, which is used to compress data and convert the logarithmic mean back to the original unit; ln represents the natural logarithmic function, which is used to restore data and maintain a smooth distribution.
[0103] Step five specifically includes:
[0104] Preset a multimodal fitness threshold Q, compare and evaluate it with the multimodal fitness index Qms, and adjust the knowledge query path in real time. The specific evaluation content is as follows:
[0105] When the multimodal fitness index Qms ≥ the multimodal fitness threshold Q: The current multimodal matching degree meets the knowledge retrieval requirements, maintain the existing query path, and perform the knowledge retrieval operation.
[0106] When the multimodal fitness index Qms < the multimodal fitness threshold Q: The current multimodal matching degree does not meet the knowledge retrieval requirements, and the matching degree is insufficient. At this time, adjust the query path and the matching effect of the cross-modal data.
[0107] In this embodiment, by collecting multi-modal related data including text, images, voice, and video in real time, and performing vectorization processing on each modal data based on a multi-modal feature extraction method, a cross-modal embedding representation is constructed to improve the fusion ability and consistency of different modal data; during this process, the inter-modal feature similarity Ma is calculated to measure the similarity degree of different modal data in the feature space;
[0108] The semantic consistency metric Mb is calculated to evaluate the alignment degree of different modal data at the semantic level; the cross-modal contrast loss value Mc is calculated to optimize the feature representations of different modal data, making highly correlated modal data closer and lowly correlated modal data farther away; the multi-scale feature alignment error Md is calculated to measure the consistency of different modal data in local and global features to ensure the fusion quality of cross-modal data;
[0109] Based on these parameters, the cross-modal matching degree index Mcf is calculated and compared with a preset cross-modal matching degree threshold M for evaluation. If Mcf is higher than M, the weight allocation of the current modal data is maintained and the index optimization stage is entered. If Mcf is lower than M, the weights of the modal data are dynamically adjusted, and the cross-modal data fusion mechanism is optimized based on the unified high-dimensional knowledge representation method to improve the matching quality and consistency of cross-modal data;
[0110] During the knowledge retrieval process, the knowledge entities and their associated information in the context-aware knowledge graph are obtained, and data such as text, images, voice, and video involved in the user query are collected in real time to calculate the association strength between knowledge entities and optimize the semantic consistency of different modal data to ensure the retrieval accuracy of the knowledge graph; at the same time, the entity association degree Eri is calculated to evaluate the association strength of the knowledge entities involved in the current query in the context knowledge graph; the modal consistency Mcc is calculated to measure the matching degree between the current query modality (text, image, voice, video) and the knowledge entity in different modalities; the search preference degree Upp is calculated to evaluate the matching degree between the current query content and the user's historical search behavior to optimize the personalized search results; based on these parameters, the multi-modal adaptability index Qms is calculated and compared with a preset multi-modal adaptation threshold Q for evaluation. If Qms is higher than Q, the existing query path is maintained and the knowledge retrieval is executed. If Qms is lower than Q, the query path is dynamically adjusted to improve the matching effect of cross-modal data, making the knowledge graph retrieval more accurate and efficient, and ensuring that the search results can adapt to the fusion requirements of different modal data;
[0111] The specific graph of the simulation experiment results of the multi-modal adaptability index Qms is as follows:
[0112] Among them, for the text modality, TF-IDF and BERT word vectors are used for embedding calculation. For the image modality, features are extracted through a convolutional neural network. For the speech modality, Mel-frequency cepstral coefficients are used for feature dimensionality reduction. For the video modality, key frame features are extracted through a temporal convolutional network. On this basis, the entity association strength Eri is calculated, which is comprehensively calculated from the semantic similarity, path connection weight, and co-occurrence frequency between knowledge entities. The modality consistency Mcc is calculated, and based on the embedding vectors of different modality data, the cosine similarity or cross-modal contrast learning method is used to calculate the consistency between modalities. The search preference degree Upp is calculated, and the matching degree between the current search content and the user's preference is evaluated through the user's historical search data, click behavior, and search path similarity. And a multi-modal fitness index Qms is constructed based on the above parameters.
[0113] Through experiments, it aims to show how the threshold adjustment model adjusts the knowledge query path in real time according to the calculated multi-modal fitness index Qms. The reference table is as follows;
[0114] Table 1
[0115] Number Eri Mcc Upp Qms Result 1 0.58 0.88 0.51 0.64 Not satisfied 2 0.8 0.84 0.53 0.71 Satisfied 3 0.81 0.95 0.90 0.90 Satisfied 4 0.98 0.62 0.98 0.84 Satisfied 5 0.69 0.66 0.41 0.57 Satisfied
[0116] Among them, the multi-modal adaptation threshold Q is set to 0.7.
[0117] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a knowledge graph of multimodal data, characterized in that: It includes the following steps: Step 1: Adopt a feature extraction method based on entropy weight distribution to vectorize the features of text, image, voice, and video data, construct an initial knowledge entity set, and assign unique identification numbers and multi-modal space mappings to the knowledge entities; Step 2: Based on the cross-modal consistency measurement method, match similar knowledge entities in different data sources, and adopt an information entropy de-duplication strategy to merge redundant data to construct a non-redundant knowledge graph, and at the same time use a graph convolutional network for semantic alignment; Step 3: Calculate and evaluate the search environment adaptation factor Gdf, construct a context-aware knowledge graph, and verify whether the search environment has matched enough context information; Step 4: Adopt a cross-modal contrast learning method to extract cross-modal matching related data, calculate and evaluate the cross-modal matching degree index Mcf, construct a unified high-dimensional knowledge representation, and establish a cross-modal data fusion mechanism; Step 5: Calculate the multi-modal adaptation degree index Qms based on the context-aware knowledge graph, and adjust the knowledge graph retrieval strategy through a dynamic index scheduling mechanism; finally, evaluate and analyze the multi-modal adaptation degree index Qms, and adjust the knowledge query path in real time.
2. The method for constructing a knowledge graph of multimodal data according to claim 1, wherein: Step 1 specifically includes: First, obtain multi-modal related data including text, image, voice, and video, and based on the feature extraction method of entropy weight distribution, decompose the features of multi-modal related data of different modalities, extract the key attributes of each modality data in the multi-modal related data, and convert the key attributes into high-dimensional feature vectors; subsequently, construct an initial knowledge entity set, and based on the deep semantic matching mechanism, calculate the semantic similarity of the high-dimensional feature vectors, and then identify the knowledge entities with high similarity; on this basis, construct a hierarchical clustering model based on the embedding vector calculation of the graph neural network; at the same time, use the hierarchical clustering model to adaptively group the matched knowledge entities, calculate the similarity threshold according to the semantic relevance of the knowledge entities, and then dynamically adjust the clustering structure; then, assign a unique identification number to each knowledge entity, and use the multi-modal space mapping method to map and fuse different modality knowledge in a unified space according to the semantic relationship between entities and the association degree between modalities by establishing cross-modal data correspondence relationships.
3. The knowledge graph construction method for multimodal data according to claim 1, wherein: Step 2 specifically includes: First, obtain knowledge entities from different data sources, and based on cross-modal consistency measurement methods, extract features of knowledge entities including text, images, speech, and videos. Through cosine similarity calculation, Euclidean distance measurement, and embedding vector similarity calculation methods based on deep neural networks, match the same or similar knowledge entities, and evaluate the confidence of the matching results; then, based on the deduplication strategy of information entropy, calculate the information entropy of the feature distributions of similar knowledge entities, and according to the information gain threshold, perform knowledge entity merging processing, where knowledge entities with low information gain are merged first, and knowledge entities with high information gain are optimized for storage, and then construct the deduplicated knowledge graph data structure; subsequently, for semantic deviations caused by different sources, adopt a semantic alignment method based on graph convolutional networks to construct graph structure connections between knowledge entities, and based on the node embedding representation method, calculate the semantic offset of different data sources, adjust the representation vectors between entities, and adjust the representation consistency of knowledge entities from different sources to complete semantic alignment processing.
4. A method for constructing a knowledge graph of multimodal data according to claim 1, characterized in that: Step three specifically includes: Collect and obtain user search behavior data in real time, including feature information such as search keywords, click frequencies, residence times, and interaction paths, and then obtain relevant data for search environment adaptation in real time by constructing a user behavior prediction model; extract the search intention deviation Sip, knowledge entity access frequency Kaf, context association stability Cas, and historical search matching degree Hsm from the relevant data for search environment adaptation, and after dimensionless processing, calculate and obtain the environment adaptation factor Gdf through the following formula; 5. A method for constructing a knowledge graph of multimodal data according to claim 1, characterized in that: Step three specifically also includes: Adopt a reinforcement learning method to construct an optimization model for the association weights of knowledge entities, and use user search habits, click feedback, and search result preferences as state inputs; construct an association matrix of knowledge entities, use the Markov decision process to calculate the jump probability of users between knowledge entities, and based on the reward feedback mechanism of reinforcement learning, dynamically adjust the association weights between knowledge entities, where the state transition probability is determined by the search historical path and entity access frequency; during the training process of reinforcement learning, continuously adjust the association weights between knowledge entities based on policy gradient optimization, and combine the optimal policy selection mechanism to optimize the ranking of knowledge entities and improve the search matching degree at the same time; Finally, establish a personalized knowledge entity index based on user group characteristics, continuously optimize the knowledge graph structure using an incremental data update mechanism, and adjust the search strategy based on a dynamic index scheduling method.
6. The knowledge graph construction method for multimodal data according to claim 1, characterized in that: Step three specifically also includes: Preset a search environment adaptation threshold H1, compare and evaluate it with the environment adaptation factor Gdf of the search environment, and judge whether the current search environment has met the context matching requirements. The specific evaluation content is as follows: If the environment adaptation factor Gdf of the search environment ≥ the search environment adaptation threshold H1, it is determined that the current search environment has matched sufficient context information, and at this time, extract relevant data for cross-modal matching. If the search environment adaptation factor Gdf < the search environment adaptation threshold H1, it is determined that the current search environment does not match sufficient context information. At this time, the context awareness weight parameter is dynamically adjusted, and the user search behavior prediction model is updated.
7. The knowledge graph construction method for multimodal data according to claim 1, characterized in that: Step four specifically includes: Collect multi-modal related data including text, images, voice, and video in real-time, extract features from each modal related data, use cross-modal contrast learning method to construct a cross-modal representation space, and extract multi-modal related data embedding vectors; on this basis, based on the calculation of feature similarity between modalities, construct a modal contrast loss function, and use the contrast learning mechanism to adjust the representation vectors of different modal data; During the process, gradually collect the feature similarity Ma between modalities, semantic consistency measure Mb, cross-modal contrast loss value Mc, and multi-scale feature alignment error Md respectively, and after dimensionless processing, calculate and obtain the cross-modal matching degree index Mcf through the following formula:
8. A method for constructing a knowledge graph of multimodal data according to claim 1, characterized in that: Step four specifically further includes: Compare and evaluate the preset cross-modal matching degree threshold M with the cross-modal matching degree index Mcf, and dynamically adjust the weight allocation of different modal data; the specific evaluation content is as follows: When the cross-modal matching degree index Mcf ≥ the cross-modal matching degree threshold M, it means that the current cross-modal matching degree meets the requirements, keep the weight allocation of the current modal data, and enter the subsequent index optimization stage; When the cross-modal matching degree index Mcf < the cross-modal matching degree threshold M, the current modal data matching degree does not meet the requirements. At this time, dynamically adjust the weights of each modal data, including preferentially adjusting the weight with the lowest modal data matching degree and iterative adjustment. Finally, use the unified high-dimensional knowledge representation method to map different modal data to the shared vector space, and construct a cross-modal data fusion mechanism based on the modal feature relationship; during the data fusion process, adjust the knowledge representation between modalities based on the semantic alignment model, and perform cross-modal feature alignment operations according to the correlation characteristics between different modal data, adjust the feature expression methods of each modal data, and perform cross-modal association modeling operations based on the fused knowledge representation.
9. The method for constructing a knowledge graph of multimodal data according to claim 1, wherein: Step five includes: First, obtain the knowledge entities and their association information in the context awareness knowledge graph, and collect multi-modal data such as text, images, voice, and video involved in the user query in real-time; then, based on the multi-modal feature extraction method, vectorize different modal data and construct a cross-modal embedding representation; use the structure relationship of the knowledge graph to calculate the association strength between knowledge entities, and optimize the semantic consistency of different modal data based on the modal alignment method; finally, generate a unified multi-modal knowledge representation through the modal feature fusion technology, providing data support for the subsequent multi-modal adaptation degree calculation and knowledge retrieval optimization. Collect the knowledge entities and knowledge entity association information in the context awareness knowledge graph in real-time, and after extracting features from modal data including text, images, voice, and video, obtain the entity association degree Eri, modal consistency Mcc, and search preference degree Upp, and after dimensionless processing, calculate and obtain the multi-modal adaptation degree index Qms through the following formula: In the formula, i represents an index variable, and X i represents a set of subordinate parameters required for calculating the multimodal fitness index Qms, where the entity association degree Eri is X1, the modal consistency Mcc is X2, and the search preference degree Upp is X3; The exp represents the natural exponential function, which is used to compress data and convert the logarithmic mean back to the original unit; the ln represents the natural logarithmic function, which is used to restore data and maintain a smooth distribution.
10. A method for constructing a knowledge graph of multimodal data according to claim 1, characterized in that: Step five specifically includes: Preset a multi-modal adaptation threshold Q, compare and evaluate it with the multi-modal adaptation degree index Qms, and adjust the knowledge query path in real time; the specific evaluation content is as follows: When the multi-modal adaptation degree index Qms ≥ the multi-modal adaptation threshold Q: the current multi-modal matching degree meets the knowledge retrieval requirements, maintain the existing query path, and perform the knowledge retrieval operation. When the multi-modal adaptation degree index Qms < the multi-modal adaptation threshold Q: the current multi-modal matching degree does not meet the knowledge retrieval requirements, and the matching degree is insufficient; at this time, adjust the query path and the cross-modal data matching effect.
Citation Information
Cited By
Project decision optimization control method, device and equipment based on knowledge graph
CN120509686A
Knowledge graph construction method and system based on large language model technology
CN120523966A
Cross-domain Internet of Things equipment intelligent collaboration method and system based on semantic knowledge graph
CN120750992A
Multi-modal data dynamic fusion method and system based on distributed edge cloud collaboration
CN120805066A
Multi-modal data dynamic fusion method and system based on distributed edge-cloud cooperation
CN120805066B