Knowledge graph optimization method based on large model and multi-modal data fusion
Through the method of fusion of large models and multimodal data, the semantic alignment and dynamic response problems of knowledge graphs in multimodal data processing are solved, efficient and semantically consistent graph construction is achieved, the graph structure is optimized, and the quality and usability of knowledge graphs are improved.
Patent Information
- Application Number
- CN202510400855.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing knowledge graph construction methods lack a unified semantic alignment mechanism in multimodal data processing, and cannot respond to data changes dynamically, resulting in a long period of graph construction and poor timeliness, and it is difficult to take into account semantic consistency and structural rationality, introducing redundant nodes and error relationships.
The method of fusion of large models and multimodal data is adopted, and multi-level feature extraction and semantic coding is performed through pre-training large models. Combined with dynamic modeling data sampling and graph update of adaptive Poisson distribution model, the graph structure is optimized using adaptive Poisson distribution reinforcement learning to realize cross-modal semantic feature coding and graph optimization.
Deep encoding of cross-modal semantic features is realized, semantic projection errors between modals are reduced, and graph construction is ensured that the semantic driving and time-sensitive. The semantic consistency and structural sparseness of the knowledge graph are optimized, and redundant information is automatically identified and adjusted, which improves the availability and accuracy of the graph.
Smart Images

Figure CN120338067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and in particular to a method for optimizing knowledge graphs based on the fusion of large models and multimodal data. Background Art
[0002] With the rapid development of artificial intelligence, big data, and Internet of Things technologies, knowledge graphs, as an important tool that can structurally represent massive information and reveal semantic relationships between entities, have been widely applied in search engines, intelligent question answering, recommendation systems, and the field of healthcare. To improve the breadth and depth of knowledge graph construction, researchers have begun to attempt to introduce image, audio, and video multimodal data to achieve a comprehensive modeling of complex semantics. However, existing knowledge graph construction methods mainly rely on structured or semi-structured text data, and have limited capabilities in processing non-textual modal data, resulting in obvious deficiencies in the information integrity and semantic richness of the generated graphs.
[0003] Specifically, the existing methods generally have the following defects in the process of multimodal data processing: First, there is a lack of a unified semantic alignment mechanism between multimodal data, and it is difficult to effectively fuse features of different modalities in the same vector space; Second, traditional modeling means cannot dynamically respond to changes in multimodal data in terms of time and semantics, lacking flexible data update and sampling strategies, resulting in a long graph construction cycle and poor timeliness; Third, most of the existing methods for optimizing the structure of knowledge graphs are based on static rules or single strategies, and it is difficult to balance semantic consistency and structural rationality, easily introducing redundant nodes and incorrect relationships, affecting the usability and accuracy of the graphs.
[0004] In summary, the existing technologies have significant deficiencies in multimodal semantic unified modeling, adaptive sampling and update mechanisms, and graph structure optimization. There is an urgent need for a new method for optimizing knowledge graphs that can integrate the semantic encoding capabilities of large models and the perception capabilities of multimodal data to break through the current bottlenecks in heterogeneous information integration and dynamic knowledge modeling. Summary of the Invention
[0005] An object of the present invention is to propose a method for optimizing knowledge graphs based on the fusion of large models and multimodal data. The present invention can adjust the data collection and graph update frequencies according to real-time data semantic changes to ensure that the graph construction process has semantic drivability and time sensitivity.
[0006] A method for optimizing knowledge graphs based on the fusion of large models and multimodal data according to an embodiment of the present invention includes the following steps:
[0007] S1. Automatically collect various multimodal data such as text, images, audio, and video through multiple channels to construct a multimodal data set;
[0008] S2. Preprocess the multi-modal dataset to form a preprocessed multi-modal dataset;
[0009] S3. Use a pre-trained large model to perform multi-level feature extraction and semantic encoding on the preprocessed multi-modal dataset to generate a unified semantic vector representation;
[0010] S4. Input the unified semantic vector representation into an adaptive Poisson distribution model to dynamically model the arrival rate and distribution characteristics of multi-modal data, and determine the adaptive parameters for data sampling and update;
[0011] S5. Based on the unified semantic vector representation and adaptive parameters, use a multi-modal data fusion algorithm to perform semantic alignment and feature integration on each modality data to generate a set of fused semantic representations;
[0012] S6. Use the fused data representation to perform entity extraction and relationship recognition, and construct a preliminary knowledge graph, where the entities and relationships of the constructed knowledge graph are based on the fused data representation;
[0013] S7. Apply a structure optimization algorithm based on adaptive Poisson distribution reinforcement learning to the preliminary knowledge graph to automatically verify the nodes, edges, and their attributes of the knowledge graph, remove redundant information, and perform structural adaptive adjustment to form an optimized knowledge graph.
[0014] Optionally, S1 includes the following steps:
[0015] S11. Set the multi-modal data collection channel set M = {m t , m i , m a , m v}, where m t , m i , m a , m v represent the collection channels for collecting text, image, audio, and video data respectively. Configure the collection rules according to heterogeneous information sources to construct a multi-modal data stream;
[0016] S12. For each modality data s k ∈ M, use a distributed task scheduling mechanism to trigger the collection module in parallel, and perform polling collection of the data stream according to the preset sampling period to form a time-series multi-modal dataset;
[0017] S13. Aggregate all the multi-modal data collection task results into the original dataset to generate a multi-modal dataset D raw :
[0018]
[0019] Among them, d k represents the kth multimodal data sample, N is the total number of collected samples, t k Indicates the data collection timestamp, v k Indicates the original data content, l k A logical identifier that indicates the source of data.
[0020] Optionally, S2 includes the following steps:
[0021] S21. For multimodal dataset D raw Each multimodal data sample d in k , according to the modal data s k , respectively call the corresponding modal processing modules, perform modality-related noise filtering operations, and obtain the denoised multimodal data set;
[0022] S22. The denoised multimodal data set is transformed into the modal data set according to the modal data s k Map to a unified format set to generate a multimodal dataset with standardized format;
[0023] S23. Perform data cleaning operations on the data samples in the multimodal data set to remove data samples with illegal formats, missing fields, abnormal collection, and repeated logical identifiers to form a preprocessed multimodal data set D clean .
[0024] Optionally, S3 includes the following steps:
[0025] S31. The preprocessed multimodal dataset D clean Input the pre-trained large model and define the feature extraction process of the pre-trained large model as a multi-level semantic mapping function F emb , multi-level semantic mapping function F emb Multiple feature encoding layers Sequentially constructed, each feature encoding layer E l The semantic coding unit is implemented based on the pre-trained large model, and the input data feature vector is extracted and encoded to generate semantic feature representations at different levels. The outputs of all feature coding layers are connected to form a unified multi-level feature vector.
[0026] S32. The unified multi-level feature vector is subjected to dimensionality reduction and nonlinear mapping processing through the feature fusion unit of the pre-trained large model to generate a unified semantic vector representation of standard dimensions, forming a unified semantic vector representation set V that can simultaneously represent the semantic information of text, image, audio, and video multimodal data emb :
[0027]
[0028] Among them, represents the semantic vector representation obtained after the k-th multimodal data sample d k passes through the multi-level semantic mapping function F emb .
[0029] Optionally, S4 includes the following steps:
[0030] S41. Input the unified semantic vector representation set V emb into the improved adaptive Poisson distribution model. The improved adaptive Poisson distribution model defines a semantic activity function Λ(·) to characterize the potential impact intensity of the k-th multimodal data sample on the new entities and relationships in the knowledge graph at time t:
[0031]
[0032] where, is the semantic activity of the k-th sample at time t k , Λ0 is the benchmark semantic activity, and β j (t) is the structural weight of the j-th feature at time t, dynamically estimated by combining the change rate of the entity or relationship corresponding to this feature in the current knowledge graph, is the projection function of the j-th semantic feature, and J is the dimension of the unified semantic vector;
[0033] S42. Construct a joint Poisson intensity function that fuses semantic novelty and structural heterogeneity based on the semantic activity function:
[0034]
[0035] where, is the optimal Poisson modeling parameter, θ represents the set of Poisson modeling parameters to be optimized, t k is the timestamp of the data sample, n k is the number of samples within the sampling window, represents the KL divergence between the current sample and the existing nodes in the knowledge graph, reflecting semantic novelty, and γ is a regulation parameter to control the difference weight, represents the semantic activity value of the k-th sample at time t k under the set of Poisson modeling parameters to be optimized;
[0036] S43. Calculate the data sampling period T and the graph update threshold δ t in combination with the current knowledge graph structure state G t =(E t ), R samp , where E upd is the entity set, R t is the relationship set, R tIs a set of relationships:
[0037]
[0038] Among them, Is the average semantic activity at time t, ρ is the structure density adjustment factor, used to adjust the sampling frequency according to the relationship density, η is the update sensitivity factor, Is the average semantic deviation of the current batch of samples;
[0039] S44. Set the data sampling period T samp And the graph update threshold δ upd As the adaptive parameters of the multi-modal data fusion module and the knowledge graph update mechanism.
[0040] Optionally, the S5 includes the following steps:
[0041] S51. Based on the unified semantic vector representation set V emb And the adaptive parameter set {T samp , δ upd} to construct a multi-modal semantic alignment mapping mechanism. For each multi-modal data sample, select the corresponding semantic alignment strategy according to the modality type s k Of the multi-modal data sample, and map the semantic alignment strategy semantic vector To the shared semantic alignment space, so that the semantic features in different modalities are aligned in the same vector space, forming a semantic alignment vector
[0042] S52. Combine the semantic alignment vector With the semantic activity of the corresponding sample And its KL divergence from the existing semantics in the knowledge graph Input into the feature fusion mechanism, and perform feature weighted integration according to the importance and novelty of the multi-modal semantic contribution, dynamically adjust the weights of each modality in the fusion representation, and generate the fused data semantic representation Construct a fused semantic representation set V fused .
[0043] Optionally, the S6 includes the following steps:
[0044] S61. Use the fused semantic representation set V fused As input, use the predefined entity extraction mechanism to analyze each data semantic representation, and identify and extract the candidate entity set E cand ;
[0045] S62. Compare the fused semantic representation V fused With the candidate entity set E candAs input, a relationship recognition mechanism is adopted to detect the semantic associations between candidate entities, identify and extract the candidate relationship set R cand ;
[0046] S63. Based on the candidate entity set E cand and the candidate relationship set R cand , construct a preliminary knowledge graph G init =(E init , R init ), where is the preliminarily determined entity set, is the preliminarily determined relationship set, and all entities and their relationships are constructed based on the fused data representation.
[0047] Optionally, the S7 includes the following steps:
[0048] S71. Input the preliminarily constructed knowledge graph G init into the structure optimization module. The structure optimization module constructs a graph structure optimization policy set Π={π θ} based on the improved adaptive Poisson distribution reinforcement learning algorithm, where π θ is a parameterized optimization policy for dynamically adjusting the node, edge, and their attribute structures in the knowledge graph;
[0049] S72. Define the state space S, action space A, and reward function R for graph optimization. The state space S represents the semantic confidence distribution of each node and relationship in the current graph structure. The action space A includes node retention, edge update, attribute adjustment, and redundant removal structure operations. The reward function R is jointly constructed based on the semantic consistency score and the structure sparsity;
[0050] S73. According to the semantic activity obtained from the adaptive Poisson distribution modeling result and the graph update threshold δ upd , dynamically adjust the action sampling probability and update frequency in the graph structure optimization, so that the reinforcement learning policy has time sensitivity and semantic drivability;
[0051] S74. In each round of structure optimization, use the optimization policy π θ to execute the action a t ∈A on the current graph structure, and obtain the reward r t ∈S based on the current state s t =R(s t , a t ), and update the parameter θ according to the policy gradient method to continuously optimize the structure adjustment process;
[0052] S75. Repeat the structure optimization iteration until the convergence condition is met, and output the finally optimized knowledge graph G opt= (E opt , R opt ).
[0053] Optionally, the final optimized knowledge graph construction satisfies:
[0054] Retain nodes: semantic activity and have structural connectivity with the core sub-graph of the graph;
[0055] Remove nodes: semantic activity and have no high-confidence associated edges;
[0056] Retain edges: Both end nodes connected by the edge satisfy the retention conditions, and the edge semantic confidence ≥ the set edge confidence threshold;
[0057] Merge edges or attributes: Multiple semantically repetitive edges pointing to the same entity, or redundant values of multiple attributes are merged into a unified attribute representation after structural reduction.
[0058] The beneficial effects of the present invention are:
[0059] (1) By inputting heterogeneous modal data such as text, images, audio, and video into a pre-trained large model, designing a multi-level semantic mapping function, the present invention realizes deep encoding of cross-modal semantic features, outputs a unified semantic vector representation, effectively captures context semantics and structural features in different modalities through a multi-layer encoding structure, and significantly reduces the semantic projection error between modalities with the help of a semantic space unified alignment mechanism.
[0060] (2) The present invention designs a semantic activity function based on a unified semantic vector, combines feature dynamic weights and KL divergence to model the potential update value of samples, and dynamically generates a data sampling period and a graph update threshold by optimizing Poisson modeling parameters, which can adjust the data acquisition and graph update frequencies according to real-time data semantic changes to ensure that the graph construction process has semantic drivability and time sensitivity.
[0061] (3) The present invention proposes a graph structure optimization algorithm based on adaptive Poisson distribution reinforcement learning, establishes a dynamic reward function with semantic confidence and structural sparsity as the core, combines the current semantic activity and the update threshold, designs a time-sensitive action sampling mechanism, and the optimization strategy can automatically identify and adjust redundant nodes, weakly related edges, and attribute conflicts in the knowledge graph to form an optimization path with good convergence and strong interpretability. Description of the Drawings
[0062] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0063] Figure 1Flowchart of a knowledge graph optimization method based on the fusion of large models and multimodal data proposed by the present invention. Detailed implementation manners
[0064] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0065] Refer to Figure 1 , a knowledge graph optimization method based on the fusion of large models and multimodal data, includes the following steps:
[0066] S1. Automatically collect various multimodal data such as text, images, audio, and video through multiple channels to construct a multimodal dataset;
[0067] S2. Preprocess the multimodal dataset to form a preprocessed multimodal dataset;
[0068] S3. Use a pre-trained large model to perform multi-level feature extraction and semantic encoding on the preprocessed multimodal dataset to generate a unified semantic vector representation;
[0069] S4. Input the unified semantic vector representation into an adaptive Poisson distribution model to dynamically model the arrival rate and distribution characteristics of multimodal data, and determine the adaptive parameters for data sampling and update;
[0070] S5. Based on the unified semantic vector representation and adaptive parameters, use a multimodal data fusion algorithm to perform semantic alignment and feature integration on each modality of data to generate a set of fused semantic representations;
[0071] S6. Use the fused data representation for entity extraction and relationship recognition to construct a preliminary knowledge graph, where the entities and their relationships in the constructed knowledge graph are based on the fused data representation;
[0072] S7. Apply a structure optimization algorithm based on adaptive Poisson distribution reinforcement learning to the preliminary knowledge graph to automatically verify the nodes, edges, and their attributes of the knowledge graph, eliminate redundant information, and perform adaptive adjustment of the structure to form an optimized knowledge graph.
[0073] In this embodiment, S1 includes the following steps:
[0074] S11. Set the multimodal data collection channel set M = {m t , m i , m a , m v}, where m t , m i , m a , m vThey represent the collection channels for collecting text, image, audio and video data respectively, configure the collection rules according to the heterogeneous information sources, and build a multimodal data stream;
[0075] S12. For each modal data s k ∈M, a distributed task scheduling mechanism is used to trigger the acquisition module in parallel, and the data stream is polled and collected according to the preset sampling period to form a time-series multimodal data set;
[0076] S13. All multimodal data collection task results are aggregated into the original dataset to generate a multimodal dataset D with consistent structure. raw :
[0077]
[0078] Among them, d k represents the kth multimodal data sample, N is the total number of collected samples, t k Indicates the data collection timestamp, v k Indicates the original data content, l k A logical identifier that indicates the source of data.
[0079] In this implementation, S2 includes the following steps:
[0080] S21. For multimodal dataset D raw Each multimodal data sample d in k , according to the modal data s k , respectively call the corresponding modal processing modules, perform modality-related noise filtering operations, and obtain the denoised multimodal data set;
[0081] S22. The denoised multimodal dataset is transformed into the modal data according to the modal data s k Map to a unified format set to generate a multimodal dataset with standardized format;
[0082] S23. Perform data cleaning operations on the data samples in the multimodal dataset to remove data samples with illegal formats, missing fields, abnormal collection, and repeated logical identifiers to form a preprocessed multimodal dataset D clean .
[0083] In this implementation, S3 includes the following steps:
[0084] S31. The preprocessed multimodal dataset D clean Input the pre-trained large model and define the feature extraction process of the pre-trained large model as a multi-level semantic mapping function F emb , multi-level semantic mapping function F emb Multiple feature encoding layers are sequentially formed, where each feature encoding layer E l is implemented based on the semantic encoding unit of the pre-trained large model, and extracts and encodes the input data feature vector to generate semantic feature representations at different levels. The outputs of all feature encoding layers are concatenated to form a unified multi-level feature vector;
[0085] S32. Pass the unified multi-level feature vector through the feature fusion unit of the pre-trained large model for dimensionality reduction and non-linear mapping processing to generate a unified semantic vector representation with a standard dimension, forming a set V of unified semantic vector representations that can simultaneously represent the semantic information of multi-modal data such as text, images, audio, and video emb :
[0086]
[0087] wherein, represents the semantic vector representation obtained after the k-th multi-modal data sample d k passes through the multi-level semantic mapping function F emb ;
[0088] In this embodiment, S4 includes the following steps:
[0089] S41. Input the set V of unified semantic vector representations emb into the improved adaptive Poisson distribution model. The improved adaptive Poisson distribution model defines a semantic activity function Λ(·) to characterize the potential influence intensity of the k-th multi-modal data sample on the new entities and relationships in the knowledge graph at time t:
[0090]
[0091] wherein, is the semantic activity of the k-th sample at time t k , Λ0 is the benchmark semantic activity, and β j (t) is the structural weight of the j-th feature at time t, dynamically combined with the change rate estimation of the entity or relationship corresponding to this feature in the current knowledge graph is the projection function of the j-th semantic feature, and J is the dimension of the unified semantic vector;
[0092] S42. Construct a joint Poisson intensity function that combines semantic novelty and structural heterogeneity based on the semantic activity function:
[0093]
[0094] wherein, is the optimal Poisson modeling parameter, θ represents the set of Poisson modeling parameters to be optimized, t k is the timestamp of the data sample, and n kis the number of samples within the sampling window, represents the KL divergence between the current sample and the semantics of the existing nodes in the knowledge graph, reflecting semantic novelty, and γ is a tuning parameter to control the difference weight, represents the k-th sample at time t k and the semantic activity value under the set of Poisson modeling parameters to be optimized;
[0095] S43. According to the optimal Poisson modeling parameters combined with the current knowledge graph structure state G t =(E t ,R t ) calculate the data sampling period T samp and the graph update threshold δ upd , where E t is the entity set, R t is the relationship set:
[0096]
[0097] Among them, is the average semantic activity at time t, ρ is the structure density adjustment factor, used to adjust the sampling frequency according to the relationship density, η is the update sensitivity factor, is the average semantic deviation of the current batch of samples;
[0098] S44. Use the data sampling period T samp and the graph update threshold δ upd as the adaptive parameters of the multi-modal data fusion module and the knowledge graph update mechanism.
[0099] In this embodiment, S5 includes the following steps:
[0100] S51. Based on the unified semantic vector representation set V emb and the adaptive parameter set {T samp ,δ upd}, construct a multi-modal semantic alignment mapping mechanism. For each multi-modal data sample, select the corresponding semantic alignment strategy according to the modal type s k of the multi-modal data sample, and map the semantic alignment strategy semantic vector to the shared semantic alignment space, so that the semantic features under different modalities are aligned in the same vector space to form a semantic alignment vector
[0101] S52. Combine the semantic alignment vector with the semantic activity of the corresponding sample and the KL divergence between it and the existing semantics of the knowledge graph In the input feature fusion mechanism, feature weighted integration is performed according to the importance and novelty of multimodal semantic contributions, the weights of each modality in the fusion representation are dynamically adjusted, and the semantic representation of the fused data is generated. Construct a set V of fused semantic representations fused .
[0102] In this embodiment, S6 includes the following steps:
[0103] S61. Take the set V of fused semantic representations fused as input, use a predefined entity extraction mechanism to analyze each data semantic representation, and identify and extract a candidate entity set E cand ;
[0104] S62. Take the fused semantic representation V fused and the candidate entity set E cand as input, adopt a relationship recognition mechanism to detect the semantic associations between candidate entities, and identify and extract a candidate relationship set R cand ;
[0105] S63. Based on the candidate entity set E cand and the candidate relationship set R cand , construct a preliminary knowledge graph G init =(E init , R init ), where is the preliminarily determined entity set, is the preliminarily determined relationship set, and all entities and their relationships are constructed based on the fused data representation.
[0106] In this embodiment, S7 includes the following steps:
[0107] S71. Input the preliminarily constructed knowledge graph G init into the structure optimization module. The structure optimization module constructs a set of graph structure optimization strategies Π={π θ} based on the improved adaptive Poisson distribution reinforcement learning algorithm, where π θ is a parameterized optimization strategy used to dynamically adjust the node, edge, and their attribute structures in the knowledge graph;
[0108] S72. Define the state space S, action space A, and reward function R for graph optimization, where the state space S represents the semantic confidence distribution of each node and relationship in the current graph structure, the action space A includes node retention, edge update, attribute adjustment, and redundant removal structure operations, and the reward function R is jointly constructed based on the semantic consistency score and the structural sparsity;
[0109] S73. Based on the semantic activity obtained from the adaptive Poisson distribution modeling results and the graph update threshold δ upd , dynamically adjust the action sampling probability and update frequency in the graph structure optimization, so that the reinforcement learning strategy has time sensitivity and semantic drive;
[0110] S74. In each round of structure optimization, adopt the optimization strategy π θ to execute the action a t ∈A on the current graph structure, and obtain the reward r t ∈S based on the current state s t = R(s t , a t ), and update the parameter θ according to the policy gradient method to continuously optimize the structure adjustment process;
[0111] S75. Repeat the structure optimization iteration until the convergence condition is met, and output the finally optimized knowledge graph G opt =(E opt , R opt ).
[0112] In this embodiment, the construction of the finally optimized knowledge graph satisfies:
[0113] Retain nodes: Semantic activity and have structural connectivity with the core subgraph of the graph;
[0114] Remove nodes: Semantic activity and have no high-confidence associated edges;
[0115] Retain edges: Both end nodes of the connection satisfy the retention conditions, and the edge semantic confidence ≥ the set edge confidence threshold;
[0116] Merge edges or attributes: Multiple semantically repeated edges pointing to the same entity, or redundant multi-attribute values are merged into a unified attribute representation after structure reduction.
[0117] Example 1:
[0118] At 9:35 am on November 13, 2024, in the big data laboratory of the neurology department of a hospital in City A, a data acquisition server numbered "BTH-MOD1" started a new round of multi-modal acquisition tasks for neurology case data within the hospital, aiming to build a dynamic knowledge graph between the "neuropathy - clinical manifestations - treatment methods" triples to serve the hospital's intelligent consultation assistance system.
[0119] Within just two hours, the system collected a total of 13,486 multi-modal sample data, including 7,850 text medical records, 2,331 MRI image scans, 2,700 voice records, and 605 video clips from the surgical live broadcast system. The data was packed according to the timestamp and sent into the "Unified Data Preprocessing Channel D-PIPE" to execute the preprocessing process of denoising, standardization, and modal synchronization marking:
[0120] At 11:46 am, the system automatically marked a data sample with a relatively large semantic conflict:
[0121] Medical record number: TRH20241113-01095;
[0122] Time: 2024-11-13 10:24:31;
[0123] Modality: Text, Image;
[0124] Text content: "The patient suddenly developed left limb paralysis in the past two days, and the preliminary diagnosis is brainstem hemorrhage;"
[0125] Image content: The MRI image shows a high-density lesion in the right basal ganglia area, which is initially determined to be a bleeding lesion.
[0126] The method of the present invention performs fusion on its semantic vectors and finds that the KL divergence between the text semantic vector `[0.13, 0.22, 0.91,..., 0.08]` and the image semantic vector `=[0.10, 0.19, 0.95,..., 0.05]` is 0.18. The system automatically marks it as "inter-modal semantic contradiction", and the semantic activity Λ value is 0.91, which is higher than the update threshold of 0.62. The system decides to retain this sample for subsequent fusion judgment of conflict knowledge.
[0127] At 15:12 that afternoon, the data fusion engine activated the automatic atlas update mechanism according to the set 3 hours. The fusion module mapped the sample semantic vectors to the shared space and then calculated the semantic weights. The new representation `[0.12, 0.21, 0.93,..., 0.06]` generated after semantic fusion of this conflict sample was identified as a "bleeding site conflict sample" with high semantic confidence. The system judged this sample as a false alarm of the real condition through the reinforcement learning strategy, and a new relationship was added to the atlas:
[0128] Entity 1: "Brainstem hemorrhage"
[0129] Entity 2: "Basal ganglia hemorrhage"
[0130] Relationship type: "Exclusion - false alarm conflict"
[0131] Time label: 2024-11-13 10:24.
[0132] Such false positive relationships were completely ignored in traditional knowledge graphs before, and there was no error repair mechanism. The method of the present invention actively models modal conflicts for the first time and annotates the metadata of the knowledge graph, providing an important basis for the subsequent intelligent medical interview system.
[0133] At 12:04 on November 15th, the project team simulated the clinical medical interview process in the deployment test environment. The doctor entered the question: "The patient shows sudden left-sided paralysis. Which parts are bleeding?" The traditional knowledge graph only returned two candidate entities, "brainstem" and "thalamus", while the knowledge graph generated by the present invention returned three candidates, "brainstem", "basal ganglia", and "thalamus", and automatically attached confidence ranking and conflict prompt explanations. The confidence results are as follows:
[0134] Brainstem: 0.91;
[0135] Basal ganglia: 0.87 (semantic conflict marked);
[0136] Thalamus: 0.72;
[0137] Based on this, doctors can more specifically judge the imaging results and avoid misjudgment.
[0138] To further verify the system performance, the research team conducted a one-week AB control experiment, using the traditional knowledge graph construction process (Text-GCN + EarlyFusion) and the method of the present invention respectively, to build graphs and conduct question-answering tests on 200,000 multi-modal neurology case data collected from 5 hospitals. The test data covers three modal combinations of text, image, and voice. The experimental scenarios include three indicators: accuracy evaluation of the question-answering system, graph construction efficiency evaluation, and graph consistency evaluation.
[0139] The following is a summary of the comparative experiment data:
[0140]
[0141]
[0142] In one experiment, in the record numbered "TRH20241115-05573", the knowledge triple extracted by the traditional method is:
[0143] Entity 1: "Thalamic hemorrhage";
[0144] Entity 2: "Speech disorder";
[0145] Relationship: "Causal";
[0146] While the method of the present invention identifies a more complete graph structure:
[0147] Entity 1: "Thalamic hemorrhage";
[0148] Entity 2: "Insufficient blood supply to the frontal lobe";
[0149] Entity 3: "Speech disorder";
[0150] Relationship 1: "Trigger";
[0151] Relationship 2: "Secondary to";
[0152] Time link: 2024-11-12 → 2024-11-14;
[0153] More entities and relationships are identified in the knowledge graph, and the time evolution path is accurately expressed, which is of great significance for clinical causal analysis.
[0154] As of November 30, 2024, the knowledge graph construction team of a certain hospital has successfully generated 42 optimized knowledge graphs using the method of the present invention, covering 8 professional directions in neurology, respiratory medicine, and cardiology. Deployed in the hospital's intelligent diagnosis and decision-making system, after going online, the system automatically identifies the clinical input statement "The patient suddenly becomes confused and has cold hands and feet" and matches the entity "Insufficient brainstem perfusion" in the knowledge graph with a confidence level of 0.88, correctly prompting the doctor to prioritize the investigation of cerebral blood perfusion and avoid misdiagnosis as "diabetic hypoglycemia".
[0155] This embodiment truly demonstrates the actual ability of the present invention in the processing of complex multi-modal medical data. From semantic encoding, modal fusion to knowledge graph optimization, each step provides quantitative data and clear scenarios, comprehensively verifying the core advantages of this method being superior to the existing technical solutions in terms of semantic consistency, construction efficiency, and knowledge graph quality.
[0156] The present invention inputs heterogeneous modal data such as text, images, audio, and video into a pre-trained large model, designs a multi-level semantic mapping function, realizes the deep encoding of cross-modal semantic features, outputs a unified semantic vector representation, effectively captures the context semantics and structural features in different modalities through a multi-layer encoding structure, and significantly reduces the semantic projection error between modalities with the help of a semantic space unified alignment mechanism.
[0157] The present invention designs a semantic activity function based on a unified semantic vector, combines feature dynamic weights and KL divergence to model the potential update value of samples, and dynamically generates a data sampling period and a knowledge graph update threshold by optimizing Poisson modeling parameters, which can adjust the data collection and knowledge graph update frequencies according to real-time data semantic changes, ensuring that the knowledge graph construction process has semantic drive and time sensitivity.
[0158] The present invention proposes a graph structure optimization algorithm based on adaptive Poisson distribution reinforcement learning, establishes a dynamic reward function with semantic confidence and structural sparsity as the core, combines the current semantic activity and update threshold, designs a time-sensitive action sampling mechanism, and the optimization strategy can automatically identify and adjust redundant nodes, weakly related edges and attribute conflicts in the knowledge graph, forming an optimization path with good convergence and strong interpretability.
[0159] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A knowledge graph optimization method based on the fusion of large models and multi-modal data, characterized in that It includes the following steps: S1. Automatically collect various multimodal data such as text, images, audio, and video through multiple channels to construct a multimodal dataset; S2. Preprocess the multimodal dataset to form a preprocessed multimodal dataset; S3. Use a pre-trained large model to perform multi-level feature extraction and semantic encoding on the preprocessed multimodal dataset to generate a unified semantic vector representation; S4. Input the unified semantic vector representation into an adaptive Poisson distribution model to dynamically model the arrival rate and distribution characteristics of multimodal data, and determine the adaptive parameters for data sampling and updating; S5. Based on the unified semantic vector representation and adaptive parameters, use a multimodal data fusion algorithm to perform semantic alignment and feature integration on each modal data to generate a set of fused semantic representations; S6. Use the fused data representation to perform entity extraction and relationship recognition to construct a preliminary knowledge graph; S7. Implement a structure optimization algorithm based on adaptive Poisson distribution reinforcement learning on the preliminary knowledge graph to automatically verify the nodes, edges, and their attributes of the knowledge graph, remove redundant information, and perform structure adaptive adjustment to form an optimized knowledge graph.
2. The knowledge graph optimization method based on the fusion of large models and multimodal data according to claim 1, wherein The S1 includes the following steps: S11. Set the multi-modal data acquisition channel set M = {m t , m i , m a , m v}, where m t , m i , m a , m v respectively represent the acquisition channels for collecting text, image, audio, and video data. According to the heterogeneous information sources, configure the acquisition rules to construct a multi-modal data stream; S12. For each piece of modal data s k ∈ M, a distributed task scheduling mechanism is used to trigger the acquisition module in parallel, and the data stream is polled and acquired according to the preset sampling period to form a time - serialized multi - modal data set; S13. Consolidate all the results of the multi-modal data acquisition tasks into the original dataset uniformly to generate a multi-modal dataset D with a consistent structure raw : Among them, d k represents the k-th multimodal data sample, N is the total number of collected samples, t k represents the data collection timestamp, v k represents the original data content, l k represents the logical identifier of the data source.
3. A knowledge graph optimization method based on the fusion of large models and multi-modal data according to claim 1, characterized in that The S2 includes the following steps: S21. For each multimodal data sample d raw in the multimodal dataset D k , according to the modality data s k to which it belongs, respectively call the corresponding modality processing module to perform modality-related noise filtering operations to obtain a denoised multimodal dataset; S22. Map the denoised multi-modal data set according to the modal data s k to a unified format set to generate a multi-modal data set with standardized format; S23. Perform a data cleaning operation on the data samples in the multi-modal dataset, removing data samples with illegal formats, missing fields, abnormal acquisitions, and duplicate logical identifiers, to form a pre-processed multi-modal dataset D clean .
4. A method for optimizing a knowledge graph based on the fusion of a large model and multimodal data according to claim 1, characterized in that, The S3 includes the following steps: S31. Input the preprocessed multi-modal dataset D clean into a pre-trained large model, and define the feature extraction process of the pre-trained large model as a multi-level semantic mapping function F emb , where the multi-level semantic mapping function F emb consists of multiple feature encoding layers in sequence. Each feature encoding layer E l is implemented based on the semantic encoding unit of the pre-trained large model, and performs feature extraction and encoding on the input data feature vector to generate semantic feature representations at different levels. The outputs of all feature encoding layers are concatenated to form a unified multi-level feature vector; S32. Process the unified multi-level feature vector through the feature fusion unit of the pre-trained large model for dimensionality reduction and non-linear mapping to generate a unified semantic vector representation with a standard dimension, forming a set V of unified semantic vector representations that can simultaneously represent the semantic information of multi-modal data such as text, images, audio, and video emb : Among them, represents the semantic vector representation obtained after the k-th multi-modal data sample d k passes through the multi-level semantic mapping function F emb 5. The knowledge graph optimization method based on large model and multi-modal data fusion according to claim 4, wherein The S4 includes the following steps: S41. Input the unified semantic vector representation set V emb into the improved adaptive Poisson distribution model, which defines a semantic activity function Λ(·) to characterize the potential impact intensity of the k-th multi-modal data sample on the new entities and relationships in the knowledge graph at time t: wherein, is the semantic activity of the k-th sample at time t k , Λ0 is the baseline semantic activity, and β j (t) is the structural weight of the j-th feature at time t, estimated by dynamically combining the change rate of the entity or relationship corresponding to the feature in the current knowledge graph is the projection function of the j-th semantic feature, and J is the dimension of the unified semantic vector S42. Construct a joint Poisson intensity function that combines semantic novelty and structural heterogeneity based on a semantic activity function: Among them, is the optimal Poisson modeling parameter, θ represents the set of Poisson modeling parameters to be optimized, t k is the timestamp of the data sample, n k is the number of samples within the sampling window, represents the KL divergence between the current sample and the semantics of the existing nodes in the knowledge graph, reflecting semantic novelty, and γ is a tuning parameter to control the difference weight, represents the k-th sample at time t k and the semantic activity value under the set of Poisson modeling parameters to be optimized; S43. According to the optimal Poisson modeling parameters Combine with the current knowledge graph structure state G t =(E t , R t ) Calculate the data sampling period T samp And the graph update threshold δ upd , where E t Is the entity set, R t Is the relationship set: Among them, is the average semantic activity at time t, ρ is the structure density adjustment factor used to adjust the sampling frequency according to the relationship density, and η is the update sensitivity factor. is the average semantic deviation of the current batch of samples. S44. Set the data sampling period T samp and the atlas update threshold δ upd as the adaptive parameters of the multi-modal data fusion module and the knowledge graph update mechanism.
6. The knowledge graph optimization method based on the fusion of large models and multimodal data according to claim 5, wherein The S5 includes the following steps: S51. Based on the unified semantic vector representation set V emb and the adaptive parameter set {T samp , δ upd}, construct a multimodal semantic alignment mapping mechanism. For each multimodal data sample, according to the modality type s of the multimodal data sample k select the corresponding semantic alignment strategy, and map the semantic alignment strategy semantic vector to the shared semantic alignment space, so that the semantic features under different modalities are aligned in the same vector space, forming a semantic alignment vector S52. Combine the semantic alignment vectors with the semantic activity of the corresponding samples and the KL divergence between it and the existing semantics in the knowledge graph In the input feature fusion mechanism, feature weighted integration is performed according to the importance and novelty of multimodal semantic contributions, and the weights of each modality in the fusion representation are dynamically adjusted to generate the semantic representation of the fused data Construct the set V of the fused semantic representations fused .
7. An optimization method for a knowledge graph based on the fusion of large models and multi-modal data according to claim 6, characterized in that The S6 includes the following steps: S61. Take the set V of the fused semantic representations fused as the input, use a predefined entity extraction mechanism to analyze the semantic representation of each piece of data, and identify and extract a candidate entity set E based on the semantic features in the fused data representation cand ; S62. Take the fused semantic representation V fused and the candidate entity set E cand as inputs, and use a relationship recognition mechanism to detect the semantic associations between candidate entities, identify and extract the candidate relationship set R cand ; S63. Based on the candidate entity set E cand and the candidate relationship set R cand , construct a preliminary knowledge graph G init =(E init , R init ), where is the preliminarily determined entity set, is the preliminarily determined relationship set, and all entities and their relationships are constructed based on the representation of the fused data.
8. A method for optimizing a knowledge graph based on the fusion of a large model and multimodal data according to claim 7, characterized in that, The S7 includes the following steps: S71. Input the preliminarily constructed knowledge graph G init into the structure optimization module, which constructs a set of graph structure optimization strategies Π = {π θ} based on the improved adaptive Poisson distribution reinforcement learning algorithm, where π θ is a parameterized optimization strategy for dynamically adjusting the node, edge, and their attribute structures in the knowledge graph; S72. Define the state space S, action space A, and reward function R for graph optimization, where the state space S represents the semantic confidence distribution of each node and relationship in the current graph structure, the action space A includes node retention, edge update, attribute adjustment, and redundant removal structure operations, and the reward function R is jointly constructed based on the semantic consistency score and structural sparsity; According to the semantic activity obtained from the adaptive Poisson distribution modeling results and the atlas update threshold δ upd , dynamically adjust the action sampling probability and update frequency in the atlas structure optimization, so that the reinforcement learning strategy has time sensitivity and semantic drive; S74. In each round of structure optimization, the optimization strategy π is adopted θ to perform the action a on the current graph structure t ∈ A, and obtain the reward r based on the current state s t ∈ S, where r t = R(s t , a t ). Update the parameter θ according to the policy gradient method to continuously optimize the structure adjustment process; S75. Repeat the structure optimization iteration until the convergence condition is met, and output the finally optimized knowledge graph G opt =(E opt ,R opt ).
9. A method for optimizing a knowledge graph based on the fusion of a large model and multimodal data according to claim 8, characterized in that, The construction of the finally optimized knowledge graph satisfies: Reserved Nodes: Semantic Activity and have structural connectivity with the core subgraph of the graph spectrum; Eliminated node: semantic activity and there are no high-confidence associated edges; Retained edge: Both end nodes of the connected edge satisfy the retention condition, and the edge semantic confidence ≥ the set edge confidence threshold; Merged edge or attribute: Multiple semantically repeated edges pointing to the same entity, or redundant values of multiple attributes are merged into a unified attribute representation after structural reduction.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Domain term automatic extraction method based on abnormal sub-graph detection
CN112528640A
Course recommendation system based on knowledge graph and graph attention network
CN115840853A
Multi-modal knowledge graph completion method and system based on embedded synchronization and alignment
CN118821921A
Multi-modal entity recognition method based on large language model
CN119167937A
Cited By
Knowledge graph construction method and system based on large language model technology
CN120523966A
Fragmented semantic understanding order building system and method based on knowledge graph
CN120725026A
Multi-modal data dynamic fusion method and system based on distributed edge cloud collaboration
CN120805066A
Data knowledge base management method and device for egg industry and medium
CN120852089A
Data label generation and quality control method and system based on pre-training large model
CN121117627A