Multi-source heterogeneous data fusion analysis method and system
By using a multi-source heterogeneous data fusion and analysis method, knowledge graphs are acquired and updated, solving the problem of difficulty in unifying and associating multi-source heterogeneous knowledge. A time-series knowledge graph with the ability to represent time-series information is constructed, realizing the orderly, unified, and associative expression and management of data.
Patent Information
- Application Number
- CN202511149393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-08-18
AI Technical Summary
How to organize and express massive and diverse event knowledge in an orderly, unified, and related manner, and solve the problems of difficulty in unifying, relating, and utilizing multi-source heterogeneous knowledge.
By using a multi-source heterogeneous data fusion analysis method, we can obtain multi-source heterogeneous event knowledge, extract entity information and attribute information, construct a multi-label classification model, add time-series labels to form an initial knowledge graph, extract entity time-series state sequences, fine-tune the knowledge graph using an open base code model, update the knowledge graph, and construct a time-series knowledge graph.
It enables the orderly, unified, and interconnected expression of massive and diverse knowledge, solves the problem of the lack of spatiotemporal dimension in traditional knowledge graphs, and improves the efficiency of data management and utilization.
Smart Images

Figure CN120654077B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for multi-source heterogeneous data fusion and analysis. Background Technology
[0002] Multi-source data intelligent fusion technology is a technology that integrates, processes, and analyzes data from different sources, formats, and semantics to achieve comprehensive, accurate, and in-depth information acquisition and knowledge discovery.
[0003] Multi-source data intelligent fusion technology is a technical system that uses artificial intelligence algorithms to systematically integrate and analyze heterogeneous data from multiple sources. Its core lies in overcoming the limitations of traditional single-modal data processing and achieving synergistic effects across domains. Current mainstream frameworks exhibit a three-layer fusion characteristic: data layer fusion employs distributed storage and federated learning techniques to solve the physical integration challenge of multi-source data; feature layer fusion utilizes graph convolution of deep neural networks to achieve cross-modal feature embedding; and decision layer fusion constructs a dynamic reasoning mechanism through knowledge graphs and reinforcement learning. Current technology can achieve intelligent fusion of multi-dimensional information such as text, images, and time-series signals.
[0004] With the development of internet technology, knowledge graph technology, a technique for organizing knowledge resources in information form, facilitates the rapid analysis of relationships between various types of data. Current research on knowledge graphs in professional fields mainly focuses on objective entities such as personnel, institutions, facilities, and equipment. Event knowledge graphs, however, represent knowledge at a larger granularity than objective entities, possessing characteristics such as temporality, spatiality, and dynamism, and can carry more valuable battlefield information. Research using events or events as the basic unit of knowledge representation to construct event-centric knowledge graphs can effectively organize and manage massive amounts of heterogeneous and multimodal data, deeply correlate and mine key information, and provide auxiliary support for understanding and judging behavioral patterns of events.
[0005] Event information is characterized by its fragmentation, diversity, and disorganization, making it difficult to unify, correlate, and utilize multi-source heterogeneous knowledge. Therefore, how to organize and express massive amounts of diverse event knowledge in an orderly, unified, and correlated manner has become a key issue. Summary of the Invention
[0006] In view of the above problems, the present invention provides a method and system for multi-source heterogeneous data fusion analysis.
[0007] This invention provides a method for multi-source heterogeneous data fusion and analysis, comprising: acquiring multi-source heterogeneous event knowledge; extracting entity information and attribute information from the multi-source heterogeneous data; constructing a multi-label classification model after fusing the attribute information; extracting temporal relationships from the multi-source heterogeneous data through the multi-label classification model; constructing an initial knowledge graph of the multi-source heterogeneous data after adding temporal labels to the temporal relationships; obtaining the entity temporal state sequence of the multi-source heterogeneous data through the initial knowledge graph and entity information; extracting features from the entity temporal state sequence; and processing the features of the entity temporal state sequence to update the initial knowledge graph, thereby obtaining a temporal knowledge graph of the multi-source heterogeneous data.
[0008] According to an embodiment of the present invention, entity elements of multi-source heterogeneous data are extracted, and event information of multi-source heterogeneous data is obtained through entity elements; entity attributes are obtained by detecting and parsing the event information; attribute information of each node of multi-source heterogeneous data is obtained through entity attributes; entity relationships are obtained by classifying and analyzing the attribute information.
[0009] According to an embodiment of the present invention, multi-source heterogeneous data includes multiple entities, and a multi-label classification model can obtain the event relationships of multiple entities.
[0010] According to an embodiment of the present invention, features of multi-source heterogeneous data are acquired and learned through a convolutional layer to obtain a learning result vector of the convolutional layer; the learning result vector is segmented and pooled through a pooling layer to obtain a final output vector; the similarity between the final output vector and the event relationship is calculated to obtain the event association relationship between multiple entities; and time series labels are added to the event association relationship to obtain the time series relationship of multi-source heterogeneous data.
[0011] According to an embodiment of the present invention, an event ontology is constructed using knowledge graphs and entity information; an entity set of multi-source heterogeneous data is constructed using the event ontology; and a current state feature vector of the multi-source heterogeneous data is constructed using the entity set to form an entity temporal state sequence.
[0012] According to an embodiment of the present invention, the entity temporal state sequence is transformed into a labeled sequence; the labeled sequence is fine-tuned using an Open Foundation Code model; the fine-tuned labeled sequence is trained and evaluated using an Open Foundation Code model to obtain the features of the entity temporal state sequence.
[0013] According to embodiments of the present invention, the evaluation metrics include: code accuracy, syntactic regularity, and semantic consistency.
[0014] According to an embodiment of the present invention, multiple proxy modules are obtained based on a language model programming framework; a translation model is constructed through the multiple proxy modules; the features of the entity temporal state sequence are processed through the translation model to obtain the missing entity information in the initial knowledge graph; the missing entity information is trained and then input into the initial knowledge graph to obtain a temporal knowledge graph of multi-source heterogeneous data.
[0015] According to an embodiment of the present invention, the update includes: non-linear relationships between entities, relations and timestamps in the initial knowledge graph, reasoning paths in the initial knowledge graph, and time information and event associations in the initial knowledge graph.
[0016] Another aspect of the present invention provides a system for multi-source heterogeneous data fusion analysis, comprising: a data input module for inputting multi-source heterogeneous data; a data extraction module for extracting entity information and attribute information from the multi-source heterogeneous data; a data processing module for processing the entity information and attribute information to obtain an initial knowledge graph and an entity temporal state sequence; and a data update module for updating the initial knowledge graph to obtain a temporal knowledge graph of the multi-source heterogeneous data.
[0017] The multi-source heterogeneous data fusion and analysis method and system provided by this invention can achieve the following beneficial effects:
[0018] (1) By using the entity and attribute extraction technology of multi-source heterogeneous information, the temporal multi-label relationship extraction technology and the spatiotemporal big data standardization representation technology, the massive and diverse knowledge is organized and expressed in an orderly, unified and related manner, and the scattered, diverse and messy data is unified, related and utilized, thus solving the problem that data is difficult to organize and manage in a unified manner.
[0019] (2) By using the spatiotemporal big data normalization representation technology and expert knowledge extraction and rule construction technology, based on the traditional big data information representation knowledge graph, time sequence labels are introduced into the traditional event association relationship, thereby constructing a time sequence knowledge graph framework structure with time sequence information representation capability, which solves the problem of the lack of spatiotemporal dimension in traditional knowledge graphs. Attached Figure Description
[0020] The above and other objects, features and advantages of the present invention will become more apparent from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0021] Figure 1 A flowchart illustrating a multi-source heterogeneous data fusion and analysis method according to an embodiment of the present invention is shown schematically.
[0022] Figure 2 The schematic diagram illustrates the principle of a multi-source heterogeneous data fusion and analysis method according to an embodiment of the present invention;
[0023] Figure 3 This illustration schematically shows a flowchart of extracting multi-source heterogeneous data information according to an embodiment of the present invention;
[0024] Figure 4 A flowchart illustrating the extraction of temporal relationships of multi-source heterogeneous data according to an embodiment of the present invention is shown.
[0025] Figure 5 The schematic diagram illustrates the principle of extracting temporal state sequence features of entities according to an embodiment of the present invention;
[0026] Figure 6 A block diagram illustrating language model programming according to an embodiment of the present invention is shown schematically;
[0027] Figure 7 A block diagram of a multi-source heterogeneous data fusion and analysis system according to an embodiment of the present invention is shown schematically. Detailed Implementation
[0028] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0031] Before describing specific embodiments of the present invention in detail, technical terms will first be explained to facilitate a better understanding of the present invention.
[0032] Multi-source heterogeneous data refers to a collection of data obtained from different sources that differ in structure, format, type, or storage method. This data is usually difficult to integrate or analyze directly and requires preprocessing before it can be effectively utilized.
[0033] Temporal relationship: refers to the correlation of data, events or phenomena in the time dimension, emphasizing characteristics such as time sequence, duration, periodicity and causality.
[0034] Translation model: Captures the relationship between two objects through translation in vector space (such as addition or subtraction).
[0035] In view of this, the present invention provides a method and system for multi-source heterogeneous data fusion analysis.
[0036] Figure 1 A flowchart illustrating a multi-source heterogeneous data fusion and analysis method according to an embodiment of the present invention is shown schematically. Figure 2 The schematic diagram illustrates the principle of a multi-source heterogeneous data fusion and analysis method according to an embodiment of the present invention.
[0037] like Figure 1 and Figure 2 As shown, the multi-source heterogeneous data fusion and analysis method according to an embodiment of the present invention includes steps S1 to S6.
[0038] In step S1, acquire multi-source heterogeneous event knowledge and extract entity information and attribute information from the multi-source heterogeneous data.
[0039] It extracts entity attributes, entity relationships, and entity elements from multi-source heterogeneous data, outputs information points in a fixed format according to data type, identifies a large number of information points from various documents, and integrates them in a unified form for easy inspection and comparison. Through multi-source data, it completes the automated extraction of multi-source entity relationships and supports custom relationships, providing data support for relationship graph construction.
[0040] This study analyzes the attribute information of entities at each node in multi-source heterogeneous data. Based on the data content, various attributes in the entity attributes are classified, and in-depth analysis is performed using the attribute information of each category. By employing machine learning, deep learning, and statistical methods, an attribute mining model is constructed to learn the attribute information of each node in the biological knowledge network, analyze the correlation between attributes between nodes, obtain the similarity of attributes between different entities, and then discover the correlation between entities under the condition of similar attributes, thereby mining potential entity relationships.
[0041] In step S2, the attribute information is fused to construct a multi-label classification model, and the temporal relationship of multi-source heterogeneous data is extracted through the multi-label classification model.
[0042] For the fused multi-source heterogeneous data, the multi-source heterogeneous data system is classified. A Softmax regression model is applied, with inputs mapped to real numbers between 0 and 1, and normalization ensures the sum is 1; therefore, the sum of probabilities for multiple classifications is l. Using event attribute-based temporal multi-label relation extraction technology, the initially extracted entities and pre-defined simple associations in the ontology are compared. Event entities and associations are equated to words, and a sentence classification model is used to classify indirectly related entities, thus obtaining the temporal relationships of the multi-source heterogeneous data.
[0043] For example, in handling multi-label classification (number of categories) >2) When dealing with this problem, the final output unit of the classifier requires numerical processing using the Soflmax function. The definition of the Soflmax function is as follows:
[0044]
[0045] In the formula, This represents the proportion of the current element's output exponent in the sum of all element output exponents. This indicates the mapped output of the preceding output unit. Indicates category index, This indicates the total number of categories.
[0046] For example, give a specific example to illustrate, when When the value is 4, the linear classifier model produces 4 output values. The result is expressed as:
[0047]
[0048] The four output values obtained As a result, after processing with Soflmax, the relative probabilities after numerical transformation were obtained. :
[0049]
[0050] As can be seen from the above formula... ( =2) has the highest probability, and is more likely to be classified as Class II.
[0051] For example, when using Soflmax for multi-label classification, it's important to be aware of potential numerical overflow. If the null terminology (or scalar values) is very large, the resulting value after exponential operations may overflow. This can be addressed by processing the null terminology by subtracting the maximum value from each element in the null terminology. Specifically, it is expressed as:
[0052]
[0053]
[0054] In the formula, This represents the proportion of the current element's output exponent in the sum of all element output exponents. This indicates the mapped output of the preceding output unit. Indicates category index, This indicates the total number of categories.
[0055] In step S3, after adding time-series labels to the time-series relationships, an initial knowledge graph of multi-source heterogeneous data is constructed.
[0056] By adding time-series labels to temporal relationships, multi-labeled associations with temporal information are obtained. The spatiotemporal structure and spatiotemporal relationships of events are complex, and the associations between event entities in task scenarios even have multiple labels and categories. Combining actual event information extracted from electronic texts, a temporal knowledge graph of the spatiotemporal evolution of events based on an ontology is constructed.
[0057] For example, remote supervision is used to classify and extract multi-label relationships. Multiple relationships between events are modeled using an undirected graph of relationship labels. Relationship label classification based on the undirected graph is implemented using a graph structure traversal algorithm. By finding the connected components in the relationship connected graph, all possible relationship labels belonging to the current statement expression are collected. Finally, K-means clustering is used to complete the screening and division of relationship labels. This can solve problems such as incomplete instances, high spatiotemporal discreteness of information, diverse relationship types, and single association patterns in knowledge bases.
[0058] For example, specific methods for collecting event information include: remote sensing image event extraction, identification and retrieval, information generation, information analysis combined with historical information, image retrieval of events from images obtained from the Internet, and keyword extraction and retrieval of text data.
[0059] In step S4, the entity time-series state sequence of multi-source heterogeneous data is obtained through the initial knowledge graph and entity information.
[0060] By using clustering analysis and deep neural network feature classification, event entities are extracted from heterogeneous data sources such as networks, images, and signals based on the unique identifier attribute information in the initial knowledge graph ontology. An event entity set is constructed, and the event state feature vectors of each temporal state of the event entity set and the state vectors of adjacent event entities are transformed and constructed into event current state feature vectors with stronger spatiotemporal contextual relationships by combining the LTR (Learning to Rank) feature space infiltration optimization method. This forms the temporal state sequence of event entities.
[0061] In step S5, the features of the entity's temporal state sequence are extracted.
[0062] We construct a large-scale language model called CodeLlama (open foundation code model) (with long-term context memory and attention mechanisms, capable of capturing and understanding long-term dependencies in sequences). The generated labeled sequences are input into the CodeLlama model for fine-tuning, clarifying the expectations of event entities, including code structure, syntax, function logic, etc. We control the diversity and quality of the generated code by adjusting the sampling temperature, and encode and extract features from the generated labeled sequences.
[0063] In step S6, the features of the entity temporal state sequence are processed to update the initial knowledge graph and obtain a temporal knowledge graph of multi-source heterogeneous data.
[0064] Design a multi-agent architecture for LangChain (a language model programming framework). Based on the LangChain architecture, develop individual agent modules and define the data flow path from input to output. This can include code input, data transfer and processing between agents, and the final generated code output. Map the tuples of events in the initial knowledge graph to a low-dimensional vector space. Use the temporal order of relations to model knowledge evolution in the time dimension. Apply the Language Model (LM) to the initial knowledge graph to obtain its implicit semantic information for knowledge reasoning. It can mine the association information between multiple quadruplets in the graph, capture semantic knowledge, quickly adapt to new entities and relations, and obtain an updated temporal knowledge graph of multi-source heterogeneous data. This solves the problems of insufficient extraction of time information from timestamps and insufficient information mining of associations in temporal knowledge graphs in existing completion and update methods.
[0065] Figure 3 The flowchart illustrating the extraction of multi-source heterogeneous data information according to an embodiment of the present invention is shown in the illustration.
[0066] like Figure 3 As shown, the extraction of multi-source heterogeneous data information according to an embodiment of the present invention includes steps S11 to S14.
[0067] Step S11: Extract entity elements from multi-source heterogeneous data, and obtain event information from multi-source heterogeneous data through entity elements.
[0068] For example, entity element recognition and extraction includes identifying and automatically labeling concepts such as personnel, military unit names, weapons and equipment, countries, and organizations in the text. The labeling results can be used to distinguish concept categories.
[0069] For example, event information includes the time and place of occurrence, the roles involved, and changes in related actions or states.
[0070] By utilizing deep learning models, pre-training is performed on multi-source heterogeneous input data to obtain labeled sequences corresponding to the input multi-source heterogeneous data sequences. The labeled results are then post-processed (e.g., merging labels) to obtain the final entity elements. Event information is extracted from a large amount of entity element data and presented in a structured form, generating structured event knowledge in batches. Based on the output data hierarchy, event knowledge construction is divided into event detection and event extraction. An event knowledge base is built to provide knowledge support for timeline discovery or event trend analysis.
[0071] Step S12: Detect and parse the event information to obtain entity attributes.
[0072] Extracting event element information from loose, unstructured information and generating refined structured event data, the main tasks of event extraction include three aspects: (1) text understanding, the event description text can be segmented into text units with independent semantics through syntactic component analysis, and the semantic role of the text unit is understood; (2) event parsing, identifying an event data including element units, such as entities, relationships, time, geographical information, and attribute information such as the number of people involved and event type, which can be set manually or automatically generated based on the text understanding results; (3) element filling, according to the filling requirements of element units, converting text units into attribute values that conform to the specifications to obtain entity attributes.
[0073] Step S13: Obtain the attribute information of each node of the multi-source heterogeneous data through entity attributes.
[0074] The attribute information of a specific entity is obtained from the entity attribute information. Mining and analysis are performed according to the attribute characteristics. By organizing the same attribute content, the attribute information is improved, thereby obtaining the attribute information of each node of multi-source heterogeneous data.
[0075] For example, the attribute information of each node includes biological attributes (name, release date, country, etc.), association attributes (name of associated entity, associated content, nature of association, etc.), event attributes (event time, event location, event type, etc.), and organization attributes (organization name, number of people in the organization, nature of the organization, content of the organization's activities), etc.
[0076] Step S14: After classifying the attribute information, analyze it to obtain entity relationships.
[0077] Based on the attribute information of each node, an attribute mining model is constructed using machine learning, deep learning, and statistical methods to learn the attribute information of each node in the biological knowledge network. By analyzing the entity attributes between similar entities, potential associations between entity attributes are discovered. Through a large amount of data attribute information, the similarity of attributes between different entities is obtained, and then the association relationship between entities under the condition of similar attributes is discovered, thus mining potential entity relationships.
[0078] Figure 4 A flowchart illustrating the extraction of time-series relationships of multi-source heterogeneous data according to an embodiment of the present invention is shown.
[0079] like Figure 4 As shown, the extraction of time-series relationships of multi-source heterogeneous data according to an embodiment of the present invention includes steps S21 to S24.
[0080] Step S21: The features of multi-source heterogeneous data are obtained through convolutional layers and learned to obtain the learning result vector of the convolutional layers.
[0081] A convolutional neural network, adept at feature extraction, is used to learn all the features obtained from the vector representation layer. It learns from sentences where two entities co-occur to predict the relationship between them. Feature extraction is achieved by convolving vectors within a sliding window. Typically, multiple convolutional kernels are used in the model to learn various features. Assume that multiple convolutional kernels are used in the model.
[0082] For example, assuming the model uses n convolution kernels, the convolution matrix is: The convolution operation is represented as follows:
[0083]
[0084] In the formula, This represents the resulting vector of the convolution kernel. This indicates the length of the sliding window of the convolution kernel. Represents a sequence.
[0085] Step S22: After segmenting and pooling the learning result vector through a pooling layer, the final output vector is obtained.
[0086] The pooling layer further extracts the features learned by the convolutional layer. It adopts the max pooling strategy, which selects the maximum value among a series of features learned by each convolutional kernel of the convolutional layer to extract the most valuable features. The entire sentence is divided into three segments with two event entities as the dividing point, and max pooling is performed on each segment.
[0087] For example, the result vector of each convolution kernel It will be divided into three sections The three pooling vectors are combined together to obtain the vector. Piecewise max pooling operation It can be represented as:
[0088]
[0089] For example, connecting all vectors Obtain the total vector The pooling layer is then subjected to nonlinear function operations to obtain its final output vector. for:
[0090]
[0091] Step S23: Calculate the similarity between the final output vector and the event relationship to obtain the event association relationship between multiple entities.
[0092] Sentences that correctly express the relationship between events will receive higher weights, while sentences that are incorrectly labeled will receive very low weights. The weight of a sentence is obtained by calculating the similarity between the sentence feature representation vector and the event relationship, thereby obtaining the event association relationship between multiple entities.
[0093] For example, for a set of entity pairs All the sentences in which they appear together Form a set The attention mechanism layer will calculate the corresponding weight vector for this set. Therefore, the characteristics of set T can be calculated as follows:
[0094]
[0095] Step S24: Add time-series labels to the event relationships to obtain the time-series relationships of multi-source heterogeneous data.
[0096] Since knowledge graphs need to reflect the dynamic characteristics of events and the complex relationships between events, some entities have multiple relationships, and the relationship categories are also multi-labeled and multi-attributed. By using the event attribute-based temporal multi-label relationship extraction technology, the initially extracted entities and the pre-set simple relationships in the ontology are compared. Event entities and relationships are equated to words. A sentence classification model is used to classify the relationships of entities with indirect relationships and add temporal labels to obtain the temporal relationships of multi-source heterogeneous data.
[0097] Figure 5 The schematic diagram illustrates the principle of extracting temporal state sequence features of entities according to an embodiment of the present invention; Figure 6 A block diagram illustrating language model programming according to an embodiment of the present invention is shown schematically.
[0098] like Figure 5 and Figure 6 As shown, the relationship update of the knowledge graph based on time sequence mainly includes: (1) relationship update and completion based on logical rules, link prediction of missing elements in the knowledge graph according to a series of reasoning rules formulated by experts or rules obtained by mining algorithms; (2) update and completion based on tensor decomposition, modeling the fact quadruples as fourth-order tensors, and then decomposing the fourth-order tensors so that the quadruples can be embedded in Euclidean space in the form of low-dimensional matrices, which is convenient for link prediction.
[0099] Features extracted from entity time-series state sequences include:
[0100] Based on the characteristics and structure of event expert knowledge in the CodeLlama model, the entity time sequence is transformed into a labeled sequence suitable for input into the CodeLlama model. The rules and representation methods of labeling are defined according to domain knowledge. The key information of each expert knowledge is extracted as a label, or the entire knowledge is used as part of the labeled sequence.
[0101] The generated labeled sequences are encoded and features are extracted for representation and learning in the model. Word vectors, embedding layers, or other encoding techniques can be used to convert the labeled sequences into numerical representations so that the model can understand and process them.
[0102] The generated labeled sequences are fed into the CodeLlama model for fine-tuning. During fine-tuning, the model learns how to effectively handle long-range dependencies to better understand and translate event expert knowledge. Supervised learning tasks with labeled sequences, such as sequence classification, sequence generation, or sequence labeling, can be used to guide the model's training.
[0103] Define evaluation metrics for generative models: Evaluation metrics are used to measure the degree of matching between the generated code and event expert knowledge. For each generated code sample, the evaluation is performed based on the expectations of event expert knowledge to check whether the generated code meets the expectations, including code structure, syntax, function logic, etc.
[0104] The features of entity temporal state sequences are processed to update the initial knowledge graph, including:
[0105] Based on the language model programming framework, multiple proxy modules are obtained. Each proxy module should implement its specific functions and tasks, such as syntax analysis, semantic analysis, and program logic. Appropriate programming languages and tools are used to implement the proxy modules. The developed proxy modules are integrated into the LangChain architecture to ensure smooth collaboration and information exchange between the proxy modules.
[0106] Define the path of data flow from input to output. This can include code input, data transfer and processing between agents, and the final generated code output, ensuring that data flows correctly between agents and performing necessary data processing and transformation. This may involve data structure design, data format conversion, error handling, etc., managing the control logic and flow in the data flow, such as determining the order of agent execution, handling exceptions, decision logic, etc., to ensure the correctness and reliability of the data flow.
[0107] This approach maps tuples of events in a temporal knowledge graph to a low-dimensional vector space, models knowledge evolution in the temporal dimension using the temporal order of relations, and enables translation of entity vectors, relation vectors, and timestamp vectors in space. It also models knowledge evolution in the temporal dimension using the temporal order of relations and regularizes the traditional embedding score function using the observed relation ordering of head entities.
[0108] Adaptively learning the non-linear relationships between entities, relations, and timestamps in a knowledge graph, as well as their dynamic transformations, fully captures its semantic features, thereby enabling link prediction in the knowledge graph. Neural network-based completion methods can extract knowledge sequence features and mine implicit semantic information.
[0109] A multi-layered network structure is used to learn features in the knowledge graph, and the learned important features are used to fill in the missing parts of the graph. A multi-hop reasoning model is used, and a policy network is used to learn multi-hop reasoning paths and continuously adjust the reasoning paths in the temporal knowledge graph to update the initial knowledge graph and obtain the temporal knowledge graph.
[0110] In summary, the embodiments of the present invention provide a method for multi-source heterogeneous data fusion and analysis, which has the following beneficial effects:
[0111] (1) By using the entity and attribute extraction technology of multi-source heterogeneous information, the temporal multi-label relationship extraction technology and the spatiotemporal big data standardization representation technology, the massive and diverse knowledge is organized and expressed in an orderly, unified and related manner, and the scattered, diverse and messy data is unified, related and utilized, thus solving the problem that data is difficult to organize and manage in a unified manner.
[0112] (2) By using the spatiotemporal big data normalization representation technology and expert knowledge extraction and rule construction technology, based on the traditional big data information representation knowledge graph, time sequence labels are introduced into the traditional event association relationship, thereby constructing a time sequence knowledge graph framework structure with time sequence information representation capability, which solves the problem of the lack of spatiotemporal dimension in traditional knowledge graphs.
[0113] Based on the methods disclosed in the above embodiments, the present invention also provides a multi-source heterogeneous data fusion and analysis system, which will be described below in conjunction with... Figure 7 The system is described in detail.
[0114] Figure 7 A block diagram of a multi-source heterogeneous data fusion and analysis system according to an embodiment of the present invention is shown schematically.
[0115] like Figure 7 As shown, the multi-source heterogeneous data fusion and analysis system 700 according to an embodiment of the present invention includes a data input module 710, a data extraction module 720, a data processing module 730, and a data update module 740.
[0116] The data input module 710 is used to input multi-source heterogeneous data.
[0117] The data extraction module 720 is used to extract entity information and attribute information from multi-source heterogeneous data.
[0118] The data processing module 730 is used to process entity information and attribute information to obtain an initial knowledge graph and entity temporal state sequence.
[0119] The data update module 740 is used to update the initial knowledge graph to obtain a time-series knowledge graph of multi-source heterogeneous data.
[0120] It should be noted that the embodiments of the device section are similar to those of the method section, and the technical effects achieved are also similar. For specific details, please refer to the above-mentioned method embodiment section, which will not be repeated here.
[0121] According to embodiments of the present invention, any plurality of the data input module 710, data extraction module 720, data processing module 730, and data update module 740 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the data input module 710, data extraction module 720, data processing module 730, and data update module 740 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any one of the three implementation methods, or in a suitable combination of any of them. Alternatively, at least one of the data input module 710, data extraction module 720, data processing module 730, and data update module 740 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0123] The features described in the various embodiments of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations or combinations fall within the scope of the present invention.
[0124] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. A method for fusion and analysis of multi-source heterogeneous data, characterized in that, include: Acquire multi-source heterogeneous event knowledge, and extract entity information and attribute information from the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes multiple entities; After fusing the attribute information, a multi-label classification model is constructed. This model is then used to extract the temporal relationships from the multi-source heterogeneous data, enabling the model to obtain the event relationships among the multiple entities. The multi-label classification model includes convolutional layers and pooling layers. Extracting the temporal relationships of the multi-source heterogeneous data using the multi-label classification model includes: acquiring and learning features of the multi-source heterogeneous data through the convolutional layer to obtain a learning result vector; segmenting and pooling the learning result vector through the pooling layer to obtain a final output vector; calculating the similarity between the final output vector and the event relationships to obtain the event association relationships between the multiple entities; and adding temporal labels to the event association relationships to obtain the temporal relationships of the multi-source heterogeneous data. After adding time-series tags to the time-series relationships, an initial knowledge graph of the multi-source heterogeneous data is constructed. The entity temporal state sequence of the multi-source heterogeneous data is obtained through the initial knowledge graph and the entity information. Extract features from the entity's temporal state sequence; The features of the entity temporal state sequence are processed to update the initial knowledge graph, resulting in a temporal knowledge graph of the multi-source heterogeneous data. This processing includes: obtaining multiple proxy modules based on a language model programming framework; constructing a translation model using the multiple proxy modules; processing the features of the entity temporal state sequence using the translation model to obtain missing entity information in the initial knowledge graph; training the missing entity information and then inputting it into the initial knowledge graph to obtain the temporal knowledge graph of the multi-source heterogeneous data. The update includes: non-linear relationships between entities, relations, and timestamps in the initial knowledge graph; reasoning paths in the initial knowledge graph; and time information and event associations in the initial knowledge graph.
2. The method according to claim 1, wherein, The entity information includes entity attributes, entity relationships, and entity elements. The entity information and attribute information extracted from the multi-source heterogeneous data include: Extract entity elements from the multi-source heterogeneous data, and obtain event information from the multi-source heterogeneous data through the entity elements; The entity attributes are obtained by detecting and parsing the event information; The attribute information of each node of the multi-source heterogeneous data is obtained through the entity attributes; The entity relationships are obtained by classifying and analyzing the attribute information.
3. The method according to claim 1, wherein, The step of obtaining the entity time-series state sequence of the multi-source heterogeneous data through the knowledge graph and the entity information includes: An event ontology is constructed using the knowledge graph and the entity information. The entity set of the multi-source heterogeneous data is constructed using the event ontology library; The current state feature vector of the multi-source heterogeneous data is constructed by the entity set to form the entity time-series state sequence.
4. The method according to claim 1, wherein, The features extracted from the entity's temporal state sequence include: The entity time-series state sequence is converted into a tag sequence; The marked sequence was fine-tuned using an open base code model; The fine-tuned labeled sequence is trained and evaluated using the open base code model to obtain the features of the entity's temporal state sequence.
5. The method according to claim 4, characterized in that, The evaluation metrics include: code accuracy, syntax correctness, and semantic consistency.
Citation Information
Patent Citations
Graph representation learning method and device based on multi-source heterogeneous medical knowledge graph
CN114741527A
Data processing method and system based on time sequence knowledge graph
CN120031113A