A multi-source heterogeneous data fusion method and system
By building a unified data model and related models, the automation, intelligent fusion and analysis of multi-source heterogeneous data are solved, and the complex and low efficiency of multi-source heterogeneous data fusion methods in the existing technology is solved, and data interoperability and fusion accuracy are improved.
Patent Information
- Application Number
- CN202410785209.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-06-18
AI Technical Summary
In the prior art, the multi-source heterogeneous data fusion method is complex, has low efficiency, poor fusion effect, and is low in practicality, which cannot meet the data fusion needs of multiple business systems.
By building a unified data model and building a strategy generation model, isomorphic data fusion model and fusion data analysis model based on this model, the automation, intelligent fusion and analysis of multi-source heterogeneous data are realized.
It eliminates the heterogeneity of different data sources, improves data interoperability, solves the data island problem, improves the accuracy and efficiency of data fusion, and is suitable for a variety of data sources and meets the needs of multiple business systems.
Smart Images

Figure CN118503915B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data fusion, and in particular relates to a multi-source heterogeneous data fusion method and system. Background Art
[0002] With the rapid development of enterprise informatization, various business systems have emerged, such as OA (office automation), CRM (customer relationship management), ERP (enterprise resource planning), WMS (warehouse management system), PLM (product lifecycle management), SaaS (software as a service), etc. While these systems provide enterprises with convenient business processing functions, they also bring about the problem of data islands. The heterogeneity of data between systems makes it difficult for data to communicate with each other, affecting the overall collaborative work efficiency of the enterprise. Therefore, it is of great practical significance to study a multi-source heterogeneous data fusion method. In the existing technology, the method of data fusion for multi-source heterogeneous data is complex, inefficient, and has poor fusion effect. In addition, the applicable data structure types are limited, the practicality is low, and it cannot meet the data fusion needs of various business systems. Summary of the invention
[0003] In order to solve the problems of complex methods, low efficiency, poor fusion effect and low practicality in the prior art, the present invention aims to provide a multi-source heterogeneous data fusion method and system.
[0004] The technical solution adopted by the present invention is:
[0005] A multi-source heterogeneous data fusion method comprises the following steps:
[0006] According to the differences in data structures of different data sources, a unified data model is constructed, and based on the unified data model, a strategy generation model, a homogeneous data fusion model, and a fusion data analysis model are constructed;
[0007] Collect a number of real-time heterogeneous data from different data sources, and use the strategy generation model to generate strategies based on the real-time heterogeneous data to obtain a number of corresponding real-time data mapping and conversion strategies;
[0008] According to the real-time data mapping and conversion strategy, the corresponding real-time heterogeneous data is mapped and converted to obtain a number of real-time homogeneous data that conform to the unified data model;
[0009] According to a number of real-time homogeneous data, a homogeneous data fusion model is used to fuse the homogeneous data to obtain corresponding real-time fused data;
[0010] According to the real-time fusion data, the fusion data analysis model is used to perform fusion data analysis to obtain the corresponding real-time fusion data analysis results;
[0011] If the real-time fusion data analysis results meet the requirements, the real-time fusion data is stored in the homogeneous database corresponding to the unified data model, otherwise, return to the strategy generation step.
[0012] Furthermore, according to the data structure differences of different data sources, a unified data model is constructed, and based on the unified data model, a strategy generation model, a homogeneous data fusion model and a fusion data analysis model are constructed, including the following steps:
[0013] Collecting a number of historical heterogeneous data from different data sources, and preprocessing the a number of historical heterogeneous data to obtain a number of corresponding preprocessed historical heterogeneous data;
[0014] According to a number of pre-processed historical heterogeneous data, a corresponding unified data model is constructed, and according to the unified data model, a corresponding historical data mapping and conversion strategy is set for each pre-processed historical heterogeneous data;
[0015] According to the historical data mapping and conversion strategy, the corresponding pre-processed historical heterogeneous data is mapped and converted to obtain a number of historical homogeneous data that conform to the unified data model;
[0016] Based on several pre-processed historical heterogeneous data and corresponding historical data mapping and conversion strategies, a strategy generation model is constructed using a reinforcement learning algorithm;
[0017] According to the data format of several historical homogeneous data and unified data model, a homogeneous data fusion model is constructed using deep learning algorithm, and several corresponding historical fusion data are generated;
[0018] According to several historical fusion data and preset fusion data prediction types, a deep learning algorithm is used to build a fusion data analysis model and generate several corresponding historical fusion data analysis results.
[0019] Further, according to the plurality of pre-processed historical heterogeneous data, a corresponding unified data model is constructed, and according to the unified data model, a corresponding historical data mapping and conversion strategy is set for each pre-processed historical heterogeneous data, including the following steps:
[0020] Parse the pre-processed historical heterogeneous data from different data sources to obtain several core data elements of each data source;
[0021] According to several core data elements of different data sources, set the data structure of the unified data model;
[0022] According to the data structure of the unified data model, set the data relationship between core data elements;
[0023] Set data constraints for each core data element based on the core data elements of different data sources, the data structure of the unified data model, and the data relationships between the core data elements;
[0024] Construct a corresponding unified data model based on the core data elements of different data sources, the data structure of the unified data model, the data relationship between the core data elements, and the data constraints of each core data element;
[0025] According to the unified data model, corresponding historical data mapping and conversion strategies are set for each pre-processed historical heterogeneous data.
[0026] Furthermore, the strategy generation model is built based on the DQN algorithm.
[0027] Furthermore, the homogeneous data fusion model is constructed based on the RF-BiLSTM-Attention algorithm.
[0028] Furthermore, the fusion data analysis model is constructed based on the IPCO-DBN algorithm.
[0029] Furthermore, a plurality of real-time heterogeneous data from different data sources are collected, and a strategy generation model is used to generate strategies based on the plurality of real-time heterogeneous data to obtain a plurality of corresponding real-time data mapping and conversion strategies, including the following steps:
[0030] Collect a number of real-time heterogeneous data from different data sources, and update the action space and state space of the strategy generation model according to the real-time heterogeneous data to obtain an updated action space and an updated state space;
[0031] Input real-time heterogeneous data into the strategy generation model, and generate strategies based on the updated action space and the updated state space to obtain the corresponding real-time data mapping and conversion strategies;
[0032] Traverse all real-time heterogeneous data and obtain corresponding real-time data mapping and conversion strategies.
[0033] Furthermore, based on a plurality of real-time homogeneous data, a homogeneous data fusion model is used to perform homogeneous data fusion to obtain corresponding real-time fused data, including the following steps:
[0034] Input a plurality of real-time homogeneous data into a homogeneous data fusion model, and extract a plurality of real-time homogeneous data key features corresponding to each real-time homogeneous data;
[0035] Extracting real-time isomorphic data deep features corresponding to a number of real-time isomorphic data key features corresponding to the same real-time isomorphic data, and obtaining a number of real-time isomorphic data deep features;
[0036] According to several real-time homogeneous data deep features, preset attention weights and preset fusion functions, homogeneous data fusion is performed to generate corresponding real-time fused data.
[0037] Further, according to the real-time fusion data, a fusion data analysis model is used to perform fusion data analysis to obtain corresponding real-time fusion data analysis results, including the following steps:
[0038] Input the real-time fusion data into the fusion data analysis model to extract the real-time fusion data features of the real-time fusion data;
[0039] According to the real-time fusion data features, fusion data prediction is performed to obtain the corresponding real-time fusion data prediction label;
[0040] According to the real-time fusion data prediction label, the corresponding real-time fusion data analysis result is obtained.
[0041] A multi-source heterogeneous data fusion system is used to implement a multi-source heterogeneous data fusion method. The system includes a model construction unit, a strategy generation unit, a data mapping and conversion unit, a homogeneous data fusion unit, a fusion data analysis unit and a fusion data storage unit which are connected in sequence. The model construction unit outputs a strategy generation model, a homogeneous data fusion model and a fusion data analysis model. The strategy generation unit is provided with a strategy generation model, and the strategy generation unit is respectively connected to several external data sources. The homogeneous data fusion unit is provided with a homogeneous data fusion model, the fusion data analysis unit is provided with a fusion data analysis model, and the fusion data storage unit is provided with a homogeneous database.
[0042] The beneficial effects of the present invention are:
[0043] The present invention discloses a multi-source heterogeneous data fusion method and system, which eliminates the heterogeneity of different data sources, strengthens data intercommunication, and solves the problem of data islands; adopts a unified data model, reduces the difficulty of interoperability between multi-source heterogeneous data, and improves the accuracy of data fusion; uses a policy generation model to automatically and accurately generate policies according to the data source differences of real-time heterogeneous data, realizes adaptive and dynamically adjusted data mapping and conversion policy processes, improves the data fusion efficiency of multi-source heterogeneous data and the applicability to different data sources; uses a homogeneous data fusion model to perform automatic and intelligent data fusion, reduces the complexity of data fusion, is applicable to multiple data sources, and meets the data fusion requirements of multiple business systems; uses a fusion data analysis model to perform real-time analysis on fused data, improves the effect of data fusion, and improves data utilization and system collaborative work efficiency.
[0044] Extract deep features and patterns from data to improve the effect and accuracy of data fusion
[0045] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of the multi-source heterogeneous data fusion method in the present invention.
[0047] Figure 2 It is a structural block diagram of the multi-source heterogeneous data fusion system in the present invention. DETAILED DESCRIPTION
[0048] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.
[0049] Embodiment 1:
[0050] like Figure 1 As shown, this embodiment provides a multi-source heterogeneous data fusion method, including the following steps:
[0051] S1: According to the differences in data structures of different data sources, a unified data model is constructed, and based on the unified data model, a strategy generation model, a homogeneous data fusion model, and a fusion data analysis model are constructed, including the following steps:
[0052] S1-1: Collect a number of historical heterogeneous data from different data sources, and preprocess the historical heterogeneous data to obtain a number of corresponding preprocessed historical heterogeneous data;
[0053] Data sources include business systems such as OA, CRM, ERP, WMS, PLM, and SaaS. Preprocessing includes data cleaning, duplicate elimination, and normalization in sequence to improve data quality and scale the data to a uniform scale to improve the speed and effect of model training.
[0054] S1-2: Based on a number of pre-processed historical heterogeneous data, a corresponding unified data model is constructed, and based on the unified data model, a corresponding historical data mapping and conversion strategy is set for each pre-processed historical heterogeneous data, including the following steps:
[0055] S1-2-1: Analyze the pre-processed historical heterogeneous data from different data sources to obtain several core data elements of each data source; specifically, analyze the data in each business system to identify the core data elements shared across systems;
[0056] S1-2-2: According to several core data elements of different data sources, set the data structure of the unified data model, such as tables, fields, relations, etc.;
[0057] S1-2-3: According to the data structure of the unified data model, set the data relationship between the core data elements, including one-to-one, one-to-many, many-to-many and other relationships;
[0058] S1-2-4: According to the core data elements of different data sources, the data structure of the unified data model and the data relationship between the core data elements, set the data constraints of each core data element, such as data type, value range, default value, etc.;
[0059] S1-2-5: Construct a corresponding unified data model based on the core data elements of different data sources, the data structure of the unified data model, the data relationship between the core data elements, and the data constraints of each core data element;
[0060] The unified data model serves as the basis for data fusion and covers the core data elements in each business system;
[0061] S1-2-6: According to the unified data model, set the corresponding historical data mapping and conversion strategy for each pre-processed historical heterogeneous data;
[0062] Data mapping and conversion strategies include mapping strategies and data conversion strategies. Mapping strategies are used to determine how to map heterogeneous data into a unified data model. Data conversion strategies are used to convert heterogeneous data into data in a unified data model based on a preset conversion algorithm or script. Data mapping and conversion strategies are used to handle data heterogeneity and resolve inconsistencies in data formats, data types, data semantics, etc. in different data sources.
[0063] S1-3: According to the historical data mapping and conversion strategy, the corresponding pre-processed historical heterogeneous data is mapped and converted to obtain a number of historical homogeneous data that conform to the unified data model;
[0064] S1-4: Based on several historical heterogeneous data and corresponding historical data mapping and conversion strategies, a strategy generation model is constructed using the Deep Q-network (DQN) algorithm, including the following steps:
[0065] S1-4-1: Collect the historical data mapping and conversion strategy corresponding to each pre-processed historical heterogeneous data, define the simulation environment of the strategy generation model according to the data mapping and conversion problem corresponding to the historical data mapping and conversion strategy, and construct the deep Q network and the corresponding Q function;
[0066] S1-4-2: Define the state space of the strategy generation model based on the preprocessed historical heterogeneous data and historical data mapping and conversion strategy and action space ,in, For the Status value, is the state indicator, is the state space size, For the Action value, is the action indicator, is the size of the action space;
[0067] S1-4-3: According to the state space and action space , defines the reward function of the policy generation model ,in, For the Status value, is the state and action indicator, For the Action value, For the next moment State value, and build the agency and experience storage pool of the strategy generation model;
[0068] S1-4-4: Based on the state space, action space and reward function, according to the historical data mapping and conversion strategies corresponding to several pre-processed historical heterogeneous data, the agent and deep Q network are optimized and trained, a strategy generation model is constructed, and the data mapping and conversion experience is stored in the experience storage pool;
[0069] The strategy generation model continuously adjusts and optimizes the strategy by interacting with heterogeneous data from different data sources to achieve the best strategy generation effect, realizing an adaptive and dynamically adjusted reinforcement learning process, and providing strategy support for subsequent data mapping and conversion;
[0070] S1-5: Based on the data format of several historical homogeneous data and the unified data model, the Random Forest (RF)-Bidirectional Long Short-Term Memory (BiLSTM)-Attention algorithm is used to build a homogeneous data fusion model and generate several corresponding historical fusion data, including the following steps:
[0071] S1-5-1: Based on a number of historical isomorphic data, use the RF algorithm to build a key feature selection module and generate a number of key features of historical isomorphic data corresponding to each historical isomorphic data;
[0072] S1-5-2: Based on several key features of historical isomorphic data corresponding to different historical isomorphic data, the BiLSTM algorithm is used to build a deep feature extraction module, and generate deep features of historical isomorphic data corresponding to each historical isomorphic data;
[0073] S1-5-3: According to the data format of the unified data model, use the Attention mechanism to set the attention weight of the deep feature extraction module and set the fusion function of the homogeneous data fusion module;
[0074] S1-5-4: Integrate the key feature selection module, the deep feature extraction module and the homogeneous data fusion module to obtain the corresponding homogeneous data fusion model;
[0075] S1-5-5: Based on the deep features, attention weights and fusion functions of several historical homogeneous data, a homogeneous data fusion model is used to fuse homogeneous data and generate several corresponding historical fusion data;
[0076] S1-6: Based on several historical fusion data and preset fusion data prediction types, the Improved Crested Porcupine Optimizer (ICPO)-Deep Belief Network (DBN) algorithm is used to build a fusion data analysis model and generate several corresponding historical fusion data analysis results, including the following steps:
[0077] S1-6-1: According to the preset fusion data prediction type, add the corresponding real fusion data prediction label to each historical fusion data to obtain several corresponding historical fusion data analysis samples;
[0078] S1-6-2: Use the IPCO optimization algorithm to obtain the optimal initial network parameters of the DBN network, including the following steps:
[0079] S1-6-2-1: Encode the candidate initial network parameters of the DBN network as the positions of the IPCO individuals of the IPCO optimization algorithm;
[0080] S1-6-2-2: Set the IPCO population parameters, maximum number of iterations and fitness function of the IPCO optimization algorithm;
[0081] S1-6-2-3: According to the number of individuals in the IPCO population parameters, the Circle chaotic mapping sequence is used for initialization to obtain the initial IPCO population. The formula is:
[0082]
[0083] In the formula, is the initial IPCO individual of the Circle chaos map; is the randomly generated initial IPCO individual; It is the IPCO individual indicator;
[0084] S1-6-2-4: Introduce a cyclic population reduction mechanism to limit the number of individuals in the IPCO population parameters and obtain the updated IPCO population parameters for the next iteration. The formula is:
[0085]
[0086] In the formula, For the The number of individuals in the IPCO population parameter of the iteration; For the The number of individuals in the IPCO population parameter of the iteration; It is the minimum number of individuals in the IPCO population parameter; Evaluate arguments for functions; Evaluate loop parameters for a function; Evaluate loop parameters for the maximum function;
[0087] S1-6-2-5: According to the fitness function, calculate the initial fitness value of the initial IPCO individual in the initial IPCO population. The formula is:
[0088]
[0089] In the formula, For IPCO individuals The fitness value of is the prediction mean square error function; IPCO individuals The corresponding predicted value and true value; is the number of individuals in the IPCO population parameter;
[0090] S1-6-2-6: according to the initial fitness value and the updated IPCO population parameters, the first defense strategy, the second defense strategy, the third defense strategy and the fourth defense strategy are used to update the initial IPCO population to obtain an updated IPCO population;
[0091] The formula for the first defense strategy is:
[0092]
[0093] In the formula, Updated IPCO individuals within the first defense scope; It is the initial IPCO individual within the first defense range; is a random number based on normal distribution; is a random value in the interval [0,1]; It is the optimal solution within the first defense range; is the vector generated between the true optimal solution within the first defense range and the optimal solution randomly selected from the IPCO population; It is the IPCO individual indicator; is the iteration indicator;
[0094] The formula for the second defense strategy is:
[0095]
[0096] In the formula, Updated IPCO individuals within the second defense range; It is the initial IPCO individual within the second defense range; is the search upper limit vector of the second defense range; is a random value in the interval [0,1]; Respectively initial IPCO individuals; All are [1, ]; is the vector generated between the true optimal solution within the second defense range and the optimal solution randomly selected from the IPCO population;
[0097] The formula for the third defense strategy is:
[0098]
[0099] In the formula, Updated IPCO individuals within the third defense scope; It is the initial IPCO individual within the third defense range; is the search upper limit vector of the third defense range; Respectively initial IPCO individuals; is [1, ]; The odor diffusion factor defined for the fitness function; It is a defense factor; Control parameters for search direction;
[0100] The formula for the fourth defense strategy is:
[0101]
[0102] In the formula, Updated IPCO individuals within the fourth defense scope; It is the initial IPCO individual within the fourth defense range; It is the optimal solution within the fourth defense range; All are random values in the interval [0,1]; It is a defense factor; Control parameters for search direction; is the average force affecting the search direction; is the convergence speed factor;
[0103] S1-6-2-7: Use the dynamic reverse learning algorithm to perform dynamic reverse learning on the updated IPCO population to generate a dynamic reverse IPCO population. The formula is:
[0104]
[0105] In the formula, It is a dynamic reverse IPCO individual; γ is the decreasing inertia coefficient; are the maximum and minimum values of the vector space respectively; For the updated IPCO individual;
[0106] S1-6-2-8: According to the fitness function, calculate the fitness values of all IPCO individuals in the updated IPCO population and the dynamically reversed IPCO population, and take the IPCO individual with the minimum fitness value as the optimal individual;
[0107] S1-6-2-9: If the number of iterations reaches the maximum number of iterations or the fitness value of the optimal individual meets the requirements, the optimal solution corresponding to the current optimal individual is output to obtain the optimal initial network parameters of the DBN network;
[0108] The IPCO optimization algorithm can accurately optimize the initial network structure of the DBN network, which is very important for the initialization of network parameters, because a good initialization can improve the convergence speed and final performance of the algorithm. The IPCO optimization algorithm can find the global optimal solution to the integer programming problem, which means that the best initial parameters can be found for the DBN network in all possible parameter configurations. By optimizing the initial network parameters of the DBN network, the generalization ability of the DBN model can be improved, and the risk of overfitting or underfitting can be reduced, thereby obtaining better performance on unseen data.
[0109] S1-6-3: Based on the optimal initial network parameters, use the DBN algorithm to build the initial fusion data analysis model;
[0110] S1-6-4: pre-training the initial fusion data analysis model according to a number of historical fusion data for which real fusion data prediction labels are not set, to obtain a pre-trained fusion data analysis model;
[0111] S1-6-5: Optimize and train the pre-trained fusion data analysis model based on several historical fusion data analysis samples with real fusion data prediction labels to obtain a final fusion data analysis model, and generate several corresponding historical fusion data analysis results;
[0112] S2: Collecting a number of real-time heterogeneous data from different data sources, and using a strategy generation model to generate strategies based on the real-time heterogeneous data to obtain a number of corresponding real-time data mapping and conversion strategies, including the following steps:
[0113] S2-1: Collect some real-time heterogeneous data from different data sources, and update the action space and state space of the strategy generation model based on the real-time heterogeneous data to obtain the updated action space and the updated state space ,in, After the update Status value, is the state indicator, After the update Action value, is the action indicator;
[0114] S2-2: Input real-time heterogeneous data into the policy generation model and update the action space based on it and the updated state space , using the Q function of the deep Q network, we get the predicted Q value of the possible action;
[0115] S2-3: Use the agent to take the possible action corresponding to the highest predicted Q value as the execution action, and obtain the reward value of the execution action according to the reward function ,in, For the updated Status value, is the state and action indicator, For the updated Action value, For the next moment's update Status value;
[0116] S2-4: Update the predicted Q value of the possible action according to the reward value of the executed action, obtain the updated Q value of the possible action, and repeat the previous step until the iteration number threshold is reached;
[0117] The formula is:
[0118]
[0119] In the formula, The updated status value and the updated action value The corresponding updated Q value; Status value and action value The corresponding predicted Q value; is the learning rate; is the highest predicted Q value; is the reward value for executing the action;
[0120] S2-5: Based on the updated Q value, a greedy strategy is used to output the execution action corresponding to the highest updated Q value as a real-time data mapping and conversion strategy;
[0121] S2-6: traverse all real-time heterogeneous data to obtain corresponding real-time data mapping and conversion strategies;
[0122] S3: According to the real-time data mapping and conversion strategy, the corresponding real-time heterogeneous data is mapped and converted to obtain a number of real-time homogeneous data that conform to the unified data model;
[0123] S4: Based on a plurality of real-time homogeneous data, a homogeneous data fusion model is used to fuse the homogeneous data to obtain corresponding real-time fused data, including the following steps:
[0124] S4-1: Input a plurality of real-time homogeneous data into the homogeneous data fusion model, and extract a plurality of real-time homogeneous data key features corresponding to each real-time homogeneous data, including the following steps:
[0125] S4-1-1: Use the RF structure trained in the key feature selection module to extract the feature contribution of several real-time candidate key features in real-time homogeneous data. The formula is:
[0126]
[0127] In the formula, For the Feature contribution of real-time candidate key features; For the Real-time candidate key features are The feature contribution of a random forest classification and regression tree (CART); is the CART tree indicator; It is a real-time alternative key feature indicator; is the total number of CART trees;
[0128]
[0129] In the formula, CART tree node for random forestm ,node and nodes r The Gini index of CART tree node for random forest m Medium Category The proportion of is the total number of categories; m , , r is the node indicator; is the category indicator;
[0130] S4-1-2: normalizing the feature contributions of several real-time candidate key features to obtain several corresponding normalized feature contributions;
[0131] The formula is:
[0132]
[0133] In the formula, It is the contribution of the feature after normalization; J The total number of real-time candidate key features;
[0134] S4-1-3: Generate feature selection standard values for several real-time candidate key features based on the normalized feature contribution;
[0135] The formula is:
[0136]
[0137] In the formula, For the Feature selection criteria values for real-time candidate key features; For the Normalized feature contribution of real-time candidate key features; It is a real-time alternative key feature indicator;
[0138] S4-1-4: sorting a number of real-time candidate key features in descending order according to the feature selection standard value, and selecting the first M real-time candidate key features as the real-time isomorphic data key features of the historical isomorphic data;
[0139] S4-2: Extract the real-time isomorphic data deep features corresponding to several real-time isomorphic data key features corresponding to the same real-time isomorphic data, and obtain several real-time isomorphic data deep features; the BiLSTM network can learn complex features and patterns from the real-time isomorphic data key features. These real-time isomorphic data deep features are crucial for subsequent data analysis and tasks;
[0140] S4-3: Perform homogeneous data fusion according to a number of real-time homogeneous data deep features, preset attention weights and preset fusion functions to generate corresponding real-time fusion data;
[0141] A variety of methods can be used to fuse the deep features of several real-time homogeneous data, such as feature-level fusion or decision-level fusion. This embodiment combines the fusion function of weighted summation and attention mechanism to achieve accurate and efficient feature-level fusion, and the obtained real-time fused data can truly represent the data information.
[0142] S5: Perform fusion data analysis based on the real-time fusion data using the fusion data analysis model to obtain corresponding real-time fusion data analysis results, including the following steps:
[0143] S5-1: input the real-time fusion data into the fusion data analysis model to extract the real-time fusion data features of the real-time fusion data;
[0144] S5-2: According to the real-time fusion data features, fusion data prediction is performed to obtain corresponding real-time fusion data prediction labels;
[0145] S5-3: Predict labels based on real-time fusion data and obtain corresponding real-time fusion data analysis results;
[0146] S6: If the real-time fusion data analysis result meets the requirements, the real-time fusion data is stored in a homogeneous database corresponding to the unified data model; otherwise, return to the strategy generation step.
[0147] Embodiment 2:
[0148] like Figure 2 As shown, this embodiment provides a multi-source heterogeneous data fusion system for implementing a multi-source heterogeneous data fusion method. The system includes a model construction unit, a strategy generation unit, a data mapping and conversion unit, a homogeneous data fusion unit, a fusion data analysis unit, and a fusion data storage unit connected in sequence. The model construction unit outputs a strategy generation model, a homogeneous data fusion model, and a fusion data analysis model. The strategy generation unit is provided with a strategy generation model, and the strategy generation unit is respectively connected to several external data sources. The homogeneous data fusion unit is provided with a homogeneous data fusion model, the fusion data analysis unit is provided with a fusion data analysis model, and the fusion data storage unit is provided with a homogeneous database.
[0149] A model building unit is used to build a unified data model according to the data structure differences of different data sources, and to build a strategy generation model, a homogeneous data fusion model and a fusion data analysis model based on the unified data model;
[0150] A strategy generation unit is used to collect a number of real-time heterogeneous data from different data sources, and generate strategies based on the real-time heterogeneous data using a strategy generation model to obtain a number of corresponding real-time data mapping and conversion strategies;
[0151] A data mapping and conversion unit, used to perform data mapping and conversion on corresponding real-time heterogeneous data according to the real-time data mapping and conversion strategy, to obtain a number of real-time homogeneous data that conform to the unified data model;
[0152] A homogeneous data fusion unit is used to perform homogeneous data fusion according to a plurality of real-time homogeneous data using a homogeneous data fusion model to obtain corresponding real-time fused data;
[0153] A fusion data analysis unit is used to perform fusion data analysis based on the real-time fusion data using a fusion data analysis model to obtain corresponding real-time fusion data analysis results;
[0154] The fusion data storage unit is used to store the real-time fusion data in a homogeneous database corresponding to the unified data model when the real-time fusion data analysis result meets the requirements.
[0155] The present invention discloses a multi-source heterogeneous data fusion method and system, which eliminates the heterogeneity of different data sources, strengthens data intercommunication, and solves the problem of data islands; adopts a unified data model, reduces the difficulty of interoperability between multi-source heterogeneous data, and improves the accuracy of data fusion; uses a policy generation model to automatically and accurately generate policies according to the data source differences of real-time heterogeneous data, realizes adaptive and dynamically adjusted data mapping and conversion policy processes, improves the data fusion efficiency of multi-source heterogeneous data and the applicability to different data sources; uses a homogeneous data fusion model to perform automatic and intelligent data fusion, reduces the complexity of data fusion, is applicable to multiple data sources, and meets the data fusion requirements of multiple business systems; uses a fusion data analysis model to perform real-time analysis on fused data, improves the effect of data fusion, and improves data utilization and system collaborative work efficiency.
[0156] The present invention is not limited to the above optional implementations, and anyone can derive other various forms of products under the enlightenment of the present invention. The above specific implementations should not be understood as limiting the scope of protection of the present invention. The scope of protection of the present invention should be based on the definition in the claims, and the description can be used to interpret the claims.
Claims
1. A multi-source heterogeneous data fusion method, characterized by: The steps include: According to the differences in data structures of different data sources, a unified data model is constructed, and based on the unified data model, a strategy generation model, a homogeneous data fusion model, and a fusion data analysis model are constructed; Collect a number of real-time heterogeneous data from different data sources, and use the strategy generation model to generate strategies based on the real-time heterogeneous data to obtain a number of corresponding real-time data mapping and conversion strategies; According to the real-time data mapping and conversion strategy, the corresponding real-time heterogeneous data is mapped and converted to obtain a number of real-time homogeneous data that conform to the unified data model; According to a number of real-time homogeneous data, a homogeneous data fusion model is used to fuse the homogeneous data to obtain corresponding real-time fused data; According to the real-time fusion data, the fusion data analysis model is used to perform fusion data analysis to obtain the corresponding real-time fusion data analysis results, including the following steps: Input the real-time fusion data into the fusion data analysis model to extract the real-time fusion data features of the real-time fusion data; According to the real-time fusion data features, fusion data prediction is performed to obtain the corresponding real-time fusion data prediction label; Predict labels based on real-time fusion data and obtain corresponding real-time fusion data analysis results; If the real-time fusion data analysis results meet the requirements, the real-time fusion data is stored in the homogeneous database corresponding to the unified data model, otherwise, return to the strategy generation step.
2. A multi-source heterogeneous data fusion method according to claim 1, characterized in that: According to the differences in data structures of different data sources, a unified data model is constructed, and based on the unified data model, a strategy generation model, a homogeneous data fusion model, and a fusion data analysis model are constructed, including the following steps: Collecting a number of historical heterogeneous data from different data sources, and preprocessing the a number of historical heterogeneous data to obtain a number of corresponding preprocessed historical heterogeneous data; According to a number of pre-processed historical heterogeneous data, a corresponding unified data model is constructed, and according to the unified data model, a corresponding historical data mapping and conversion strategy is set for each pre-processed historical heterogeneous data; According to the historical data mapping and conversion strategy, the corresponding pre-processed historical heterogeneous data is mapped and converted to obtain a number of historical homogeneous data that conform to the unified data model; Based on several pre-processed historical heterogeneous data and corresponding historical data mapping and conversion strategies, a strategy generation model is constructed using a reinforcement learning algorithm; According to the data format of several historical homogeneous data and unified data model, a homogeneous data fusion model is constructed using deep learning algorithm, and several corresponding historical fusion data are generated; According to several historical fusion data and preset fusion data prediction types, a deep learning algorithm is used to build a fusion data analysis model and generate several corresponding historical fusion data analysis results.
3. A multi-source heterogeneous data fusion method according to claim 2, characterized in that: According to a number of pre-processed historical heterogeneous data, a corresponding unified data model is constructed, and according to the unified data model, a corresponding historical data mapping and conversion strategy is set for each pre-processed historical heterogeneous data, including the following steps: Parse the pre-processed historical heterogeneous data from different data sources to obtain several core data elements of each data source; According to several core data elements of different data sources, set the data structure of the unified data model; According to the data structure of the unified data model, set the data relationship between core data elements; Set data constraints for each core data element based on the core data elements of different data sources, the data structure of the unified data model, and the data relationships between the core data elements; Construct a corresponding unified data model based on the core data elements of different data sources, the data structure of the unified data model, the data relationship between the core data elements, and the data constraints of each core data element; According to the unified data model, corresponding historical data mapping and conversion strategies are set for each pre-processed historical heterogeneous data.
4. The multi-source heterogeneous data fusion method according to claim 2, characterized in that: The strategy generation model is built based on the DQN algorithm.
5. The multi-source heterogeneous data fusion method according to claim 2, characterized in that: The homogeneous data fusion model is built based on the RF-BiLSTM-Attention algorithm.
6. The multi-source heterogeneous data fusion method according to claim 2, characterized in that: The fusion data analysis model is constructed based on the IPCO-DBN algorithm.
7. The multi-source heterogeneous data fusion method according to claim 4, characterized in that: Collect a number of real-time heterogeneous data from different data sources, and use the strategy generation model to generate strategies based on the real-time heterogeneous data to obtain a number of corresponding real-time data mapping and conversion strategies, including the following steps: Collect a number of real-time heterogeneous data from different data sources, and update the action space and state space of the strategy generation model according to the real-time heterogeneous data to obtain an updated action space and an updated state space; Input real-time heterogeneous data into the strategy generation model, and generate strategies based on the updated action space and the updated state space to obtain the corresponding real-time data mapping and conversion strategies; Traverse all real-time heterogeneous data and obtain corresponding real-time data mapping and conversion strategies.
8. The multi-source heterogeneous data fusion method according to claim 5, characterized in that: According to a plurality of real-time homogeneous data, a homogeneous data fusion model is used to fuse the homogeneous data to obtain corresponding real-time fused data, including the following steps: Input a plurality of real-time homogeneous data into a homogeneous data fusion model, and extract a plurality of real-time homogeneous data key features corresponding to each real-time homogeneous data; Extracting real-time isomorphic data deep features corresponding to a number of real-time isomorphic data key features corresponding to the same real-time isomorphic data, and obtaining a number of real-time isomorphic data deep features; According to several real-time homogeneous data deep features, preset attention weights and preset fusion functions, homogeneous data fusion is performed to generate corresponding real-time fused data.
9. A multi-source heterogeneous data fusion system, used to implement the multi-source heterogeneous data fusion method according to any one of claims 1 to 8, characterized in that: The system includes a model building unit, a strategy generating unit, a data mapping and conversion unit, a homogeneous data fusion unit, a fused data analysis unit and a fused data storage unit which are connected in sequence. The model building unit outputs a strategy generating model, a homogeneous data fusion model and a fused data analysis model. The strategy generating unit is provided with a strategy generating model, and the strategy generating unit is respectively connected with several external data sources. The homogeneous data fusion unit is provided with a homogeneous data fusion model. The fused data analysis unit is provided with a fused data analysis model. The fused data storage unit is provided with a homogeneous database.
Citation Information
Patent Citations
Consistency expression method for multi-source heterogeneous big data
CN105893612A
Human-machine interaction method for forestry ecological environment based on multi-source information fusion
CN107992904A