Multi-modal data and dynamic knowledge graph fusion method

By constructing a three-dimensional weight model and conflict detection for multimodal data, the problem of a single evaluation dimension for multimodal data is solved, thereby improving the accuracy of data screening and the quality of knowledge graphs.

CN121503629APending Publication Date: 2026-02-10MENGLANG SUSTAINABLE DIGITAL TECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202511680537.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies lack a collaborative quantification mechanism for data sources, time, and scenarios in the value assessment of multimodal data, resulting in low accuracy in screening high-value data and affecting the quality of knowledge graph construction.

Method used

By constructing a three-dimensional weight calculation model of data source-time-scene, the data source, time and scene dimensions of multimodal data are quantified, the comprehensive weight is calculated, preprocessing and conflict detection are performed, an initial dynamic knowledge graph is constructed, and multi-dimensional evaluation and optimization are carried out.

Benefits of technology

It improves the accuracy of multimodal data screening, ensures that data value assessment aligns with business needs, quickly locates the root cause of conflicts, reduces the error rate of conflict correction, and improves the accuracy, completeness, and consistency of knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503629A_ABST
    Figure CN121503629A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data and dynamic knowledge graph fusion method, relates to the technical field of knowledge graphs, and solves the technical problems that multi-modal data weight evaluation is single in dimension and lacks a data source, time and scene collaborative quantization mechanism. In combination with data source authority and time freshness, it is ensured that data value evaluation better meets service requirements, dimension weight coefficients can be flexibly adjusted according to scenes, different from fixed coefficients in the prior art, requirements of different fields are met, data full-link metadata is stored through a directed acyclic graph, conflict sources can be rapidly traced, and the data value evaluation efficiency is improved. And automatic classification and targeted correction of conflict types are realized based on a predefined rule, and the error rate of conflict correction is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge graph, in particular to a multi-modal data and dynamic knowledge graph fusion method. BACKGROUND

[0002] With the development of Internet of Things, sensors, image devices and other technologies, multi-modal data such as text, images, videos and sensor signals are growing explosively.

[0003] According to the patent application No. CN202510202520.9, a multi-source heterogeneous data fusion knowledge graph method and system are disclosed. The method steps are: cleaning multi-source data through a self-encoder, dynamically weighting and standardizing after low-rank decomposition and PCA denoising, GNN entity alignment and specification approval serial number authentication; combining CNN / GNN to extract text, image and sensor multi-modal features, BiLSTM to fuse time sequence information and weighted aggregation; using BERT-BiLSTM-SelfAttention-CRF to identify entities, GCN to infer relationships, and TransE to embed entity relationships into a low-dimensional space and store in Neo4j for real-time query; incrementally learning and dynamically updating the graph; and visualizing and displaying the constructed knowledge graph.

[0004] However, the value evaluation of multi-modal data in the prior art is mostly limited to a single dimension corresponding to data source authority or time freshness, without considering scene adaptability, resulting in low screening accuracy of high-value data and directly affecting the quality of subsequent knowledge graph construction. SUMMARY

[0005] To overcome the shortcomings of the prior art, the present application provides a multi-modal data and dynamic knowledge graph fusion method, which solves the problem of single dimension of multi-modal data weight evaluation and lack of data source, time and scene collaborative quantization mechanism.

[0006] To achieve the above purpose, the present application realizes the following technical scheme: a multi-modal data and dynamic knowledge graph fusion method, which specifically includes the following steps: Step one, collect multi-modal data, quantify the data source dimension, time dimension and scene dimension of the multi-modal data, obtain the corresponding dimension weight, and construct a data source-time-scene three-dimensional weight calculation model; Step two, calculate the comprehensive weight of the multi-modal data and select the basic data based on the comprehensive weight, and pre-process the multi-modal data to obtain pre-processed multi-modal data, construct an initial dynamic knowledge graph, and perform conflict detection processing; Step three, extract the features of the pre-processed multi-modal data, calculate the similarity between the features and the entities in the knowledge graph, select the features exceeding the preset threshold for fusion, and update the knowledge graph; Step four, a multi-dimensional evaluation index system covering accuracy, integrity, timeliness and consistency is constructed to evaluate the fused knowledge graph, and the knowledge graph is optimized according to the quality evaluation result.

[0007] As a further scheme of the application, the data source dimension quantification manner of the multi-modal data is as follows: The data source dimension includes data source authority, historical accuracy rate and data integrity, the data source dimension is quantified, the data source type corresponding to the multi-modal data is obtained, and scoring quantization processing is performed based on different data source types; The historical accuracy rate is quantified, the data error rate within time t1 is obtained, and the specific value of t1 is set by an operator, and the historical accuracy rate is calculated according to the formula historical accuracy rate = (1-time t1 data error rate) x calibration coefficient, and the calibration coefficient is dynamically adjusted according to the data source type; The data integrity is quantified, the field missing rate within time t is obtained, and the integrity score is obtained based on different missing rates, and then the data source dimension weight is calculated according to the formula data source dimension weight = (data authority x 0.4 + historical accuracy rate x 0.3 + integrity score x 0.3).

[0008] As a further scheme of the application, the time dimension quantification manner of the multi-modal data is as follows: The number of times corresponding to the multi-modal data within time T is obtained, and the specific value of time T is set by an operator, and the corresponding use frequency is calculated, the use frequency is compared with the preset frequency, the multi-modal data greater than the preset frequency is classified as high-frequency update data, the multi-modal data less than the preset frequency is classified as low-frequency update data, and different classification data is selected according to different classification data; For high-frequency update data, exponential decay mode is adopted, and the time dimension weight is calculated according to the formula time dimension weight = , e is a natural constant, k is an attenuation coefficient, , for low-frequency update data, piecewise linear decay is adopted, when ≤24 hours, the weight is 1.0, when 24 hours ≤72 hours, the weight = 1.0-0.01 ( -24), wherein 0.01 represents that every 1 hour, the weight decreases by 0.01, when >72 hours, the weight is 0.7, and the time dimension weight is obtained.

[0009] As a further scheme of the application, the scene dimension quantification manner of the multi-modal data is as follows: According to the application field of the multi-modal data, core scenes are divided, and a basic adaptation coefficient is given to different data types in each core scene, specifically 0-1, and through reinforcement learning analysis, the data adaptation coefficient is dynamically optimized to obtain the scene dimension weight.

[0010] As a further scheme of the application, the way of constructing the initial dynamic knowledge graph is: The comprehensive weight of each multi-modal data is calculated, and the multi-modal data is sorted according to the comprehensive weight value, and the multi-modal data with high comprehensive weight value is preferentially selected as the basic data for subsequent fusion processing, and the selected basic multi-modal data is preprocessed to obtain multi-modal preprocessed data. According to different application scenarios and business requirements, the entity type and relationship type of the knowledge graph are defined, the entity and relationship information are extracted from the data source, and the dynamic change of the relationship is learned by using the time sequence graph neural network, the entity and relationship in the knowledge graph are updated and expanded in real time, and the initial dynamic knowledge graph is constructed.

[0011] As a further scheme of the application, the way of performing conflict detection processing is: The full-link metadata is recorded for each data point, stored in a directed acyclic graph structure, and a gene map of data flow is constructed, and the entity and relationship information extracted from different data sources are compared, and when it is found that the attribute information of the same entity in different data sources is different, or the description of the same relationship in different data sources is inconsistent, it is marked as potential; The conflict type corresponding to the conflict information is obtained, specifically including numerical conflict, semantic conflict and time conflict, starting from the conflict data node, the directed acyclic graph is traversed in reverse, the traceability path is generated, the conflict reason is automatically classified based on the traceability path and the predefined rule, and the dynamic knowledge graph is obtained through the corresponding conflict correction processing based on the obtained conflict reason.

[0012] As a further scheme of the application, the predefined rule includes: If the historical error rate of the data source is greater than 10% or belongs to a low-authority data source, the conflict reason is data source error; if the processing node has algorithm BUG or parameter configuration error, the conflict reason is processing flow error; if the data timestamp and business time difference is greater than 24 hours or the term is not standardized, the conflict reason is time / semantic error; if the historical error rate of the operator is greater than 20% or the operation log shows that it is not checked, the conflict reason is manual operation error.

[0013] As a further scheme of the application, the features of the preprocessed multi-modal data include: The pre-trained language model is used for text data extraction, and semantic vectors + entity / relationship structured data are output; the CNN backbone network is used for image / video data extraction, and visual vectors + target attribute labels + video time sequence vectors are output; audio data extraction outputs acoustic feature vectors + semantic vectors + domain attribute labels; sensor data extraction outputs time sequence feature vectors + statistical features + anomaly labels.

[0014] As a further scheme of the present application, the accuracy evaluation includes: comparing the entity / relationship information in the graph with the authoritative data source, calculating the accuracy rate, and starting the data backtracking mechanism to find the problem source and correct when the accuracy rate is lower than the set threshold; The integrity evaluation includes: checking whether the entity / relationship is complete according to the pre-defined knowledge graph mode, and supplementing from multi-modal data if there is a missing.

[0015] As a further scheme of the present application, the timeliness evaluation includes: judging whether the graph information reflects the latest situation according to the data generation time and update frequency, and updating or marking the outdated information; The consistency evaluation includes: checking the logical consistency of the internal entity / relationship of the graph, the entity attribute consistency, and the cross-modal data description consistency, and adjusting according to the established rules when inconsistencies are found.

[0016] The present application provides a multi-modal data and dynamic knowledge graph fusion method. Compared with the prior art, the following beneficial effects are achieved: The present application incorporates scene adaptability into weight evaluation, combines data source authority and time freshness, ensures that data value evaluation is more suitable for business needs, and the dimension weight coefficient can be flexibly adjusted according to the scene, which is different from the fixed coefficient of the prior art and adapts to different field needs. Secondly, the directed acyclic graph stores data full-link metadata, which can quickly trace the conflict root cause, improve the root location accuracy, realize automatic classification and targeted correction of conflict types based on pre-defined rules, and reduce the conflict correction error rate. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The present application provides a multi-modal data and dynamic knowledge graph fusion method. Compared with the prior art, the following beneficial effects are achieved: DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] Please refer to Figure 1This application provides a method for fusing multimodal data with dynamic knowledge graphs, which specifically includes the following steps: Step 1: Collect multimodal data and construct a three-dimensional weight calculation model of data source-time-scene. The specific construction method is as follows: The data source dimensions of multimodal data are quantified, including data source authority, historical accuracy, and data completeness. By quantifying the data source dimensions, the data source types corresponding to the multimodal data are obtained, and scores are quantified based on different data source types. For official institutions, the scores are 0.9-1.0; for industry platforms, the scores are 0.7-0.8; and for self-media, the scores are 0.4-0.6. The historical accuracy is quantified by obtaining the data error rate within time t1, where the specific value of t1 is set by the operator. The historical accuracy is calculated according to the formula: historical accuracy = (1 - data error rate within time t1) × calibration coefficient. The calibration coefficient is dynamically adjusted according to the data source type: 0.9-1.0 for official institutions, 0.8-0.9 for industry platforms, and 0.6-0.7 for self-media. Data integrity is quantified by obtaining the field missing rate within time t and scoring it based on different missing rates to obtain an integrity score. Referring to the "National Standard for Data Quality Assessment" (GB / T36344-2018) and considering the characteristics of multimodal data application scenarios, the following scores are set: <5% = 1.0 (no missing core fields, no impact on business use), 5%-20% = 0.8 (minor missing non-core fields, which can be repaired by completion algorithms), 20%-40% = 0.5 (partial missing key fields, requiring caution), and >40% = 0.3 (severely missing core fields, extremely low value). Different missing rates correspond to different data repair costs and application risks. The gradient setting is consistent with the non-linear negative correlation between missing rate and data availability in actual business. Then, the data source dimension weight is calculated according to the formula: Data Source Dimension Weight = (Data Authority × 0.4 + Historical Accuracy × 0.3 + Integrity Score × 0.3). The time dimension of multimodal data is quantified. Based on the interval between the data generation time and the current time, differentiated attenuation rules are designed. Valid timestamps are defined for different modal data: publication time for text, shooting time for images / videos, and acquisition time for sensor data, ensuring consistency in time dimension calculation. At the same time, the number of times the multimodal data is used within time T is obtained, and the specific value of time T is set by the operator. The corresponding usage frequency is calculated and compared with the preset frequency. Multimodal data with a usage frequency greater than the preset frequency is classified as high-frequency update data, and multimodal data with a usage frequency less than the preset frequency is classified as low-frequency update data. Attenuation modes are selected according to different data categories. For frequently updated data, an exponential decay model is adopted, with the weight of the time dimension calculated according to the formula: The time dimension weight is calculated, where e is the natural constant, and k is the decay coefficient. The decay coefficient k for frequently updated data is set according to the data type: k=0.02 for text data, k=0.05 for real-time sensor data, and k=0.03 for image / video data. The values ​​are determined based on the different timeliness requirements of different modalities of data. This represents the time difference between the current time and the data timestamp. For low-frequency updated data, a piecewise linear decay is used. ≤24 hours, weight is 1.0; when 24 hours < ≤72 hours, weight = 1.0 - 0.0067 ( -24), and 0.0067 represents a decay of 0.0067 per hour, when >72 hours, weighted at 0.7-0.001 ( -24), decays by 0.001 per hour, and after 172 hours the weight drops to 0.6, and then remains at 0.6, thus obtaining the time dimension weight; The scenario dimension of multimodal data is quantified, and core scenarios are divided according to the application fields of multimodal data. A basic adaptation coefficient is assigned to different data types in each core scenario, specifically from 0 to 1. For example, in the medical scenario, text data such as medical records and medical orders have a basic adaptation coefficient of 0.9, while image data such as CT / equipment appearance images have a basic adaptation coefficient of 1. For sensor data such as vital signs data such as heart rate and temperature, the basic adaptation coefficient is 0.8. In the industrial safety scenario, text data such as equipment manuals have a basic adaptation coefficient of 1, while image data corresponding to inspection images have a basic adaptation coefficient of 0.9. Furthermore, sensor data corresponding to temperature and pressure have a basic adaptation coefficient of 1.0. At the same time, the data adaptation coefficient is dynamically optimized through reinforcement learning or statistical analysis to obtain the scenario dimension weight. Next, the obtained data source dimension weights, time dimension weights, and scenario dimension weights are weighted and summed. The comprehensive weight is calculated according to the formula: Comprehensive Weight = Data Source Dimension Weight × a + Time Dimension Weight × b + Scenario Dimension Weight × c, where a, b, and c are the corresponding dimension weight coefficients, which are dynamically adjusted according to the scenario.

[0020] Step 2: Calculate the comprehensive weight for each multimodal data point and sort the multimodal data according to the comprehensive weight value. Prioritize multimodal data with high comprehensive weight values ​​as the base data for subsequent fusion processing. Preprocess the selected base multimodal data to obtain multimodal preprocessed data. For text data, perform text cleaning to remove noise information such as special symbols and garbled characters, and perform word segmentation to convert the text into semantic units that computers can understand. For image / video data, perform image enhancement processing, including adjusting the brightness, contrast, and color saturation of the image to improve image quality, and extract key feature information from the image. For sensor data, perform data normalization processing to unify data of different dimensions into the same dimension range. Based on different application scenarios and business needs, entity types and relation types are defined for the knowledge graph. For example, in a medical scenario, entity types can include patients, diseases, drugs, and examination items, while relation types can include patients having diseases, patients taking medications, and patients undergoing examinations. Entity and relation information is extracted from authoritative medical databases, medical records, medical literature, and other data sources to construct an initial dynamic knowledge graph. A temporal graph neural network is then used to learn the dynamic changes in relations, updating and expanding the entities and relations in the knowledge graph in real time. Simultaneously, conflict detection and handling are performed, with the specific conflict detection and handling methods as follows: Full-link metadata is recorded for each data point and stored using a directed acyclic graph (DAG) structure to construct a genetic map of data flow. Entity and relationship information extracted from different data sources is compared. When discrepancies are found in the attribute information of the same entity across different data sources, or when the descriptions of the same relationship are inconsistent across different data sources, these are marked as potential conflicts. Simultaneously, feature analysis is performed on these conflicts to determine their conflict types, including numerical conflicts, semantic conflicts, and temporal conflicts. Then, starting from the conflicting data node, a reverse traversal is performed along the DAG to generate a tracing path. Based on the tracing path and predefined rules... Then, the conflict causes are automatically categorized. Specific predefined rules include: if the historical error rate of the data source is >10%, or it belongs to a low-authority data source, the conflict cause is a data source error; if the processing node has an algorithm bug or parameter configuration error, the conflict cause is a processing flow error; if the difference between the data timestamp and the business time is >24 hours, or the terminology is not standardized, the conflict cause is a time / semantic error; if the operator's historical error rate is >20%, or the operation log shows no verification, the conflict cause is a human operation error. Then, based on the obtained conflict causes, corresponding conflict correction processing is performed to obtain a dynamic knowledge graph.

[0021] Step 3: Integrate the preprocessed multimodal data with the dynamic knowledge graph. First, extract features from the preprocessed multimodal data. For text data, a pre-trained language model is used to extract features, outputting semantic vectors + entity / relation structured data. For image / video data, a CNN backbone network is used to extract features, outputting visual vectors + target attribute labels + video temporal vectors. For audio data, acoustic feature vectors + semantic vectors + domain attribute labels are output. For sensor data, temporal feature vectors + statistical features + anomaly labels are output, obtaining key feature information. Then, this key feature information is matched and associated with entities and relationships in the dynamic knowledge graph. Based on the cosine similarity calculation method, the similarity value between multimodal data features and entities and relations in the knowledge graph is calculated, and the similarity value is compared with a preset threshold. The specific value of the preset threshold is set by the operator. When the similarity exceeds the preset threshold, the multimodal data is fused with the corresponding entity or relation to enrich the content of the knowledge graph. Meanwhile, for newly emerging entities and relationships in multimodal data that do not exist in the knowledge graph, their validity and rationality are verified and confirmed before they are added to the dynamic knowledge graph, thereby realizing the dynamic expansion and updating of the knowledge graph.

[0022] Step 4: Conduct quality assessment and optimization of the integrated dynamic knowledge graph, and construct a multi-dimensional evaluation index system covering accuracy, completeness, timeliness, and consistency. In terms of accuracy assessment, compare with authoritative data sources and calculate the accuracy of entity and relation information according to the formula: Accuracy = (Number of entities consistent with authoritative data sources + Number of relations consistent with authoritative data sources) ÷ (Total number of entities in the graph + Total number of relations in the graph) × 100%. The default threshold is 95%, which can be adjusted according to the scenario, such as 98% for medical scenarios and 92% for industrial scenarios. When the accuracy is lower than the threshold, the data backtracking mechanism is activated to trace conflicting data nodes along the directed acyclic graph, recalibrate the data source dimension weight coefficients corresponding to the node, and replace the low-accuracy data. In terms of completeness assessment, based on the predefined knowledge graph pattern, we check whether all types of entities and relationships exist completely. The coverage rate is calculated as follows: Coverage = (Number of entities / relationships actually existing in the graph) ÷ (Number of necessary entities / relationships predefined based on the business scenario) × 100%. The default threshold is 90%, and for core scenarios such as emergency rescue, the threshold is 95%. When the coverage rate is insufficient, it is supplemented by multimodal data mining. For text data, the BERT model is used to extract missing entities, for image data, the CNN model is used to identify unlabeled relationships, and for sensor data, the time series prediction model is used to complete missing fields. The timeliness assessment is based on the data generation time and update frequency. The formula is Freshness = 1 - (Current Time - Last Update Time) ÷ Data Validity Period, where the validity period is set according to the data type: 7 days for high-frequency data, 30 days for medium-frequency data, and 180 days for low-frequency data. The freshness is calculated and a threshold is set to determine whether the information in the knowledge graph reflects the latest situation in a timely manner. For outdated information, i.e., freshness < 0.5, the corresponding data source update request is automatically triggered, or it is replaced with high-frequency data of the same type. Consistency assessment mainly checks the logical consistency between entities and relationships within the knowledge graph, as well as the consistency with external related data. According to the formula conflict rate = (number of entities with logical conflicts + number of relationships with attribute conflicts) ÷ (total number of entities in the graph + total number of relationships in the graph) × 100%, when inconsistencies are found, i.e. conflict rate ≥ 5%, they are resolved according to predefined priority rules: data source authority priority > time freshness priority > scenario adaptability priority. Conflicting data is automatically corrected or marked for manual review. Based on the quality assessment results, the dynamic knowledge graph was optimized in a targeted manner.

[0023] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.

[0024] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A method for fusing multimodal data with dynamic knowledge graphs, characterized in that, The method specifically includes the following steps: Step 1: Collect multimodal data, quantify the data source dimension, time dimension, and scene dimension of the multimodal data, obtain the corresponding dimension weights, and construct a three-dimensional weight calculation model of data source-time-scene. Step 2: Calculate the comprehensive weight of the multimodal data and select basic data based on it. At the same time, perform preprocessing to obtain preprocessed multimodal data, construct an initial dynamic knowledge graph, and perform conflict detection and processing. Step 3: Extract features from the preprocessed multimodal data and calculate their similarity to entities in the knowledge graph. Simultaneously, select features that exceed a preset threshold for fusion and update the knowledge graph. Step 4: Construct a multi-dimensional evaluation index system covering accuracy, completeness, timeliness, and consistency to evaluate the fused knowledge graph, and optimize it based on the quality evaluation results.

2. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, The method for quantifying the data source dimensions of multimodal data is as follows: Data source dimensions include data source authority, historical accuracy, and data integrity. The data source dimensions are quantified to obtain the data source types corresponding to multimodal data, and scoring and quantification are performed based on different data source types. The historical accuracy is quantified by obtaining the data error rate within time t1, where the specific value of t1 is set by the operator. The historical accuracy is calculated according to the formula: historical accuracy = (1 - data error rate within time t1) × calibration coefficient. The calibration coefficient is dynamically adjusted according to the data source type. Data integrity is quantified by obtaining the field missing rate within time t and scoring it based on different missing rates to obtain an integrity score. Then, the data source dimension weight is calculated according to the formula: Data Source Dimension Weight = (Data Authority × 0.4 + Historical Accuracy × 0.3 + Integrity Score × 0.3).

3. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, The time dimension quantization method for multimodal data is as follows: The system acquires the number of times multimodal data is used within time T, where the specific value of time T is set by the operator. It also calculates the corresponding usage frequency, compares the usage frequency with the preset frequency, classifies multimodal data with a frequency greater than the preset frequency as high-frequency update data, classifies multimodal data with a frequency less than the preset frequency as low-frequency update data, and selects the attenuation mode according to the different data categories. For frequently updated data, an exponential decay model is adopted, with the weight of the time dimension calculated according to the formula: Calculate the time dimension weights, where e is the natural constant and k is the decay coefficient. This represents the time difference between the current time and the data timestamp. For low-frequency updated data, a piecewise linear decay is used. ≤24 hours, weight is 1.0; when 24 hours < ≤72 hours, weight = 1.0 - 0.01 ( -24), where 0.01 represents a weight decrease of 0.01 for every hour exceeding 1 hour. >72 hours, with a weight of 0.7, yields the weight of the time dimension.

4. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, The method for quantifying the scene dimension of multimodal data is as follows: Core scenarios are divided according to the application areas of multimodal data, and basic adaptation coefficients are assigned to different data types in each core scenario. Through reinforcement learning analysis, the data adaptation coefficients are dynamically optimized to obtain the scenario dimension weights.

5. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, The initial dynamic knowledge graph is constructed as follows: The comprehensive weight of each multimodal data is calculated, and the multimodal data are sorted according to the comprehensive weight value. The multimodal data with high comprehensive weight value is selected first as the basic data for subsequent fusion processing. The selected basic multimodal data is preprocessed to obtain multimodal preprocessed data. Based on different application scenarios and business needs, the entity types and relationship types of the knowledge graph are defined, entity and relationship information is extracted from the data source, and the dynamic changes of the relationship are learned using a time-series graph neural network to update and expand the entities and relationships in the knowledge graph in real time, thus constructing an initial dynamic knowledge graph.

6. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, The method for conflict detection and handling is as follows: Full-link metadata is recorded for each data point and stored using a directed acyclic graph structure to construct a gene map of data flow. Entity and relationship information extracted from different data sources is compared. When the attribute information of the same entity differs in different data sources, or the description of the same relationship is inconsistent in different data sources, it is marked as potential. The conflict information is obtained to identify the conflict type, including numerical conflict, semantic conflict, and temporal conflict. Starting from the conflict data node, the system traverses the directed acyclic graph in reverse to generate a source path. Based on the source path and predefined rules, the conflict causes are automatically classified. Based on the obtained conflict causes, the system performs corresponding conflict correction processing to obtain a dynamic knowledge graph.

7. The method for fusing multimodal data and dynamic knowledge graphs according to claim 6, characterized in that, Predefined rules include: If the historical error rate of the data source is >10% or it belongs to a low-authority data source, the conflict is caused by a data source error; if the processing node has an algorithm bug or parameter configuration error, the conflict is caused by a processing flow error; if the difference between the data timestamp and the business time is >24 hours or the terminology is not standardized, the conflict is caused by a time / semantic error; if the historical error rate of the operator is >20% or the operation log shows no verification, the conflict is caused by a human operation error.

8. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, Features extracted from preprocessed multimodal data include: For text data, a pre-trained language model is used to extract semantic vectors and entity / relation structured data; for image / video data, a CNN backbone network is used to extract visual vectors, target attribute labels, and video temporal vectors; for audio data, acoustic feature vectors, semantic vectors, and domain attribute labels are extracted; and for sensor data, temporal feature vectors, statistical features, and anomaly labels are extracted.

9. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, Accuracy assessment includes: comparing entity / relationship information in the graph with authoritative data sources, calculating the accuracy rate, and initiating a data backtracking mechanism to find the source of the problem and correct it when the accuracy rate is lower than a set threshold; Completeness assessment includes checking whether entities / relationships are complete based on predefined knowledge graph patterns, and supplementing them from multimodal data using data mining techniques if any are missing.

10. The method for fusing multimodal data and dynamic knowledge graphs according to claim 1, characterized in that, Timeliness assessment includes: determining whether the map information reflects the latest situation based on the data generation time and update frequency, and updating or marking outdated information; Consistency assessment includes checking the logical consistency of entities / relationships within the graph, the consistency of entity attributes, and the consistency of cross-modal data descriptions. When inconsistencies are found, adjustments are made according to established rules.

Citation Information

Patent Citations

  • Multi-source heterogeneous data fusion knowledge graph method and system

    CN120181198A

  • Multi-sensor data fusion method and device and storage medium

    CN118312926A

  • Multi-modal knowledge graph completion method and system based on embedded synchronization and alignment

    CN118821921A

  • Multi-source heterogeneous data integration method and system of data integration middleware

    CN118964469A

  • Cost consultation data arrangement and analysis system based on cloud platform

    CN119863286A

Cited By

  • Automatic coding agent control system and method for multi-task scene

    CN121742410A

  • Integrated service platform management system and method based on big data

    CN121919520A

  • Integrated Service Platform Management System and Method Based on Big Data

    CN121919520B