Multi-source data integration method and system based on data mining

By extracting and classifying multi-source data features and combining with Bayesian network model output association relationships, the problem that traditional data integration methods are difficult to adapt to dynamic data environments and mining potential associations is solved, and efficient and intelligent multi-source data integration management is achieved.

CN120234355AActive Publication Date: 2025-07-01QINGDAO CHENGYUN DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510299915.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Traditional data integration methods are difficult to adapt to dynamically changing data environments and cannot fully explore potential relationships between multi-source data, resulting in inefficient data integration.

Method used

Through data feature items extraction and classification, a subset of data is constructed, and the Bayesian network model outputs data association relationships to realize integrated management of multi-source data.

Benefits of technology

It improves the efficiency and intelligence of multi-source data integration, realizes potential correlation mining and management between data, and improves the efficiency of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234355A_ABST
    Figure CN120234355A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data integration method and system based on data mining, and relates to the technical field of data management.Data features are extracted from data uploaded by a front-end data source, and the uploaded data are divided into different data subsets according to the extracted data features; after the data in each data subset is optimized, determining a data association relationship among the data in different data subsets through a Bayesian network model, and constructing a corresponding integrated class database for the data with the data association relationship by utilizing the data association relationship among the data; the index sequences of different levels are utilized to bind and associate the data with the data association relationship, so that classification and arrangement of the data uploaded by different front-end data sources are realized, and data management is more efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and specifically to a multi-source data integration method and system based on data mining. Background Art

[0002] With the advent of the digital age, data has become the core resource in all walks of life. Enterprises, research institutions, and government departments have accumulated a vast amount of data in their daily operations. These data come from a wide range of sources, including sensors, social media, enterprise databases, public data platforms, etc. The heterogeneity, dispersion, and complexity of multi-source data have brought huge challenges to data integration and utilization. Traditional data integration methods often rely on manual rules or fixed patterns, making it difficult to adapt to the dynamically changing data environment and unable to fully explore the potential associations between data. At the same time, the rapid development of data mining technology has provided new possibilities for multi-source data integration. By automatically analyzing the patterns, trends, and correlation relationships in data, data mining can help achieve more efficient and intelligent data integration.

[0003] How to perform corresponding feature classification on multi-source data and conduct integrated management of multi-source data based on the correlations existing between the data after feature classification, so as to achieve more efficient multi-source data integration, is the problem we need to solve. For this reason, a multi-source data integration method and system based on data mining are provided herein. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-source data integration method and system based on data mining.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A multi-source data integration method based on data mining includes the following steps:

[0006] Step S1: Extract data feature items from the data uploaded by the front-end data source, classify the extracted data feature items, construct corresponding data subsets according to the classification results, and import the uploaded data into each data subset.

[0007] Step S2: Optimize the data in each data subset, and import each data subset after data optimization into the Bayesian network model to output the data correlation relationships between the data in different data sets.

[0008] Step S3: Conduct integrated management of the data uploaded by each front-end data source according to the data correlation relationships between the data in each data subset output.

[0009] Further, the process of extracting data feature items from the data uploaded by the front-end data source includes:

[0010] Generate a corresponding data guidance sequence according to the front-end data source of the uploaded data and the data upload time;

[0011] Associate the generated data guidance sequence with the data uploaded by the front-end data source to obtain the original data;

[0012] Construct a corresponding data feature set according to the generated data guidance sequence;

[0013] Traverse the obtained original data and input the obtained original data into the trained decision tree model;

[0014] Extract features from the input original data through the decision tree model to obtain corresponding data feature items.

[0015] Furthermore, the training process of the decision tree model is as follows:

[0016] Collect sample data and summarize the collected sample data to obtain a corresponding sample data set;

[0017] Classify the sample data in the sample data set to obtain corresponding sample types;

[0018] Obtain the proportion of each sample type in the sample data in the sample data set and obtain the information entropy of the sample data set;

[0019] Take the obtained sample types as corresponding sample features, obtain the quantity of each sample feature in the sample data set, and obtain several sample feature subsets corresponding to the sample features;

[0020] Construct a corresponding root node according to the obtained sample type and obtain the information gain of each sample feature corresponding to the sample type;

[0021] Record the sample feature corresponding to the maximum value in the obtained information gain as the root node of the decision tree;

[0022] Set corresponding internal nodes according to each root node and set corresponding node division conditions for each internal node;

[0023] Obtain the sample features in the sample data set that meet the corresponding node division conditions, obtain the information gain of the corresponding sample features, and then generate sub-nodes corresponding to the sample features;

[0024] Then set corresponding internal nodes with this sub-node and set corresponding node division conditions for each internal node, and so on;

[0025] Set a corresponding quantity threshold for each sample feature;

[0026] When the number of sample features that meet the corresponding node division conditions is less than the set quantity threshold, stop the division, mark the corresponding sub-node as a leaf node, and thus complete the construction of the decision tree model;

[0027] Divide the collected sample data into a training set, a test set, and a validation set for training the decision tree model.

[0028] Furthermore, the process of classifying the extracted data feature items, constructing corresponding data subsets according to the classification results, and importing the uploaded data into each data subset includes:

[0029] Classify the obtained data feature items, and the classification results of the data feature items include attribute feature items, parameter feature items, and association feature items;

[0030] Establish an attribute subset, a parameter subset, and an association subset respectively in the data feature set;

[0031] Import the obtained attribute feature items into the attribute subset, import the obtained parameter feature items into the parameter subset, and import the obtained association feature items into the association subset.

[0032] Furthermore, the process of optimizing the data in each data subset includes:

[0033] Mark the data ranges where each data feature item in the obtained data feature set is located in the original data;

[0034] Set a number of redundant comparison fields and set corresponding sequence codes for each redundant comparison field;

[0035] Match the unmarked data ranges in the original data with the redundant comparison fields, and mark the data content that is the same as the redundant comparison fields as redundant data;

[0036] Convert the original data into a data stream, and replace the data stream segment corresponding to the redundant data with the corresponding sequence code;

[0037] After completing the replacement of all data stream segments of the redundant data in the original data, complete the optimization of the original data.

[0038] Furthermore, the process of importing each data subset after data optimization into the Bayesian network model and outputting the data association relationship between the data in different data sets includes:

[0039] Construct a Bayesian network model and train the constructed Bayesian network model;

[0040] Import the optimized original data in the attribute subset, parameter subset, and association subset into the trained Bayesian network model;

[0041] Output the probability of association between the attribute feature items in the attribute subset, the parameter feature items in the parameter subset, and the association feature items in the association subset;

[0042] Set a probability estimation threshold, and compare the probability of association between each output feature item with the set probability estimation threshold;

[0043] If the probability of association between each feature item is higher than the set probability estimation threshold, it indicates that there is a data association relationship between the corresponding feature items, otherwise there is no data association relationship;

[0044] Traverse each attribute feature item in the attribute subset. If there are corresponding feature items in the parameter subset and the association subset that are data-associated with the attribute feature item, mark this attribute feature item as the reference feature item;

[0045] Construct a reference subset corresponding to the reference feature item, and import other feature items that are data-associated with this reference feature item into the reference subset;

[0046] Obtain the association coefficient Gx between the reference feature item and the reference subset according to the probability of association between the feature items in the reference subset and this reference feature item;

[0047] Set the association relationship threshold intervals [G1], [G2], [G3];

[0048] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G1], it indicates that this reference feature item is associated with the corresponding reference subset;

[0049] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G2], it indicates that this reference feature item is partially associated with the corresponding reference subset, and then generate an association ratio coefficient;

[0050] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G3], it indicates that this reference feature item is not associated with the corresponding reference subset, and no operation is performed.

[0051] Furthermore, when the reference feature item is associated with the reference subset, generate an association index sequence of all the feature items in the reference feature item and the reference subset;

[0052] When there is a partial association between the reference feature item and the reference subset, each feature item in the reference subset is sorted according to the probability of association with the reference feature item from high to low, and then the corresponding features are selected according to the association ratio coefficient from high to low, and an association index sequence of the reference feature item and the features in the selected reference subset is generated;

[0053] The data association relationship between the reference feature item and each feature item in the reference subset is determined through the association index sequence.

[0054] Further, the process of integrating and managing the data uploaded by each front-end data source according to the data association relationship between the data in each output data subset includes:

[0055] Read the association index sequences corresponding to each reference feature item, and establish integrated class databases of corresponding types according to the types of the features corresponding to the association index sequences;

[0056] Import the original data corresponding to each feature and the corresponding association index sequence into the integrated class database of the corresponding type;

[0057] Mark the association index sequence as the first index sequence;

[0058] Generate a corresponding second index sequence for the integrated class database, and associate the second index sequence with each first index sequence in the corresponding integrated class database;

[0059] Associate each second index sequence with the corresponding reference feature item, thereby completing the corresponding data integration management.

[0060] Further, a multi-source data integration system applied to a multi-source data integration method based on data mining includes a data management center, and the data management center is communicatively connected to a front-end data transmission module, a multi-source data processing module, an association model output module, and a data integration management module;

[0061] The front-end data transmission module is used to upload data from the front-end data source;

[0062] The multi-source data processing module is used to extract data feature items from the data uploaded by the front-end data source, classify the extracted data feature items, construct corresponding data subsets according to the classification results, and import the uploaded data into each data subset;

[0063] The association model output module is used to optimize the data in each data subset, import each data subset with optimized data into the Bayesian network model, and output the data association relationship between the data in different data sets;

[0064] The data integration management module is used to perform integrated management on the data uploaded by each front-end data source according to the data association relationships among the data within each output data subset.

[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0066] By extracting data features from the data uploaded by the front-end data source, dividing the uploaded data into different data subsets according to the extracted data features, after optimizing the data within each data subset, determining the data association relationships among the data in different data subsets through the Bayesian network model, and then using the data association relationships among the data to construct corresponding integrated databases for the data with data association relationships, and binding and associating the data with data association relationships using different levels of index sequences, the classification and sorting of the data uploaded by different front-end data sources are realized, making the management of the data more efficient. Brief Description of the Drawings

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0068] Figure 1 It is a flowchart of the present invention. Detailed Embodiments

[0069] As Figure 1 shown, a multi-source data integration method based on data mining includes the following steps:

[0070] Step S1: Extract data feature items from the data uploaded by the front-end data source, classify the extracted data feature items, construct corresponding data subsets according to the classification results, and import the uploaded data into each data subset;

[0071] Step S2: Optimize the data within each data subset, and import each data subset after data optimization into the Bayesian network model to output the data association relationships among the data in different data sets;

[0072] Step S3: Perform integrated management on the data uploaded by each front-end data source according to the data association relationships among the data within each output data subset.

[0073] It should be further noted that in the specific implementation process, the process of extracting data feature items from the data uploaded by the front-end data source includes:

[0074] Generate a corresponding data guidance sequence according to the front-end data source of the uploaded data and the data upload time;

[0075] Associate the generated data guidance sequence with the data uploaded by the front-end data source to obtain the original data;

[0076] Construct a corresponding data feature set according to the generated data guidance sequence;

[0077] Traverse the obtained original data and input the obtained original data into the trained decision tree model;

[0078] Extract features from the input original data through the decision tree model to obtain corresponding data feature items.

[0079] It should be further noted that in the specific implementation process, the training process of the decision tree model is specifically as follows:

[0080] Collect sample data, summarize the collected sample data to obtain a corresponding sample data set, and denote the sample data set as D;

[0081] Classify the sample data in the sample data set to obtain corresponding sample types, and label each sample type as i, where i = 1, 2,..., k;

[0082] Obtain the proportion of each sample type in the sample data in the sample data set, denoted as p i ;

[0083] Then obtain the information entropy of the sample data set, denoted as H(D), where:

[0084]

[0085] Take the obtained sample types as corresponding sample features, denoted as A, and obtain the number of each sample feature in the sample data set, denoted as V;

[0086] Divide the obtained sample features to obtain several sample feature subsets corresponding to the sample feature A, and label each sample feature subset as v, where v = 1, 2,..., V;

[0087] Then denote the sample feature subset labeled as v as D v ;

[0088] And construct a corresponding root node according to the obtained sample types, and obtain the information gain of each sample feature corresponding to the sample type. Denote the information gain corresponding to the sample feature A as Gain(D, A), where:

[0089]

[0090] Mark the sample feature corresponding to the maximum value in the obtained information gain, and record the marked sample feature as the root node of the decision tree;

[0091] Set corresponding internal nodes according to each root node, and set corresponding node division conditions for each internal node;

[0092] Obtain the sample features in the sample data set that meet the corresponding node division conditions according to the node division conditions of each internal node, and obtain the information gain of the corresponding sample features, thereby generating child nodes corresponding to the sample features;

[0093] Then set corresponding internal nodes with the child nodes, and set corresponding node division conditions for each internal node, and so on;

[0094] Set corresponding quantity thresholds for each sample feature;

[0095] When the number of sample features that meet the corresponding node division conditions is less than the set quantity threshold, stop the division, and mark the corresponding child nodes as leaf nodes, thereby completing the construction of the decision tree model;

[0096] Divide the collected sample data into a training set, a test set, and a validation set, input the training set into the constructed decision tree model, train the decision tree model, and after completing the training of the decision tree model, test the data recognition accuracy of the decision tree model through the test set and the validation set. When the data recognition accuracy reaches the expectation, the construction and training of the decision tree model are completed.

[0097] It should be further noted that in the specific implementation process, the process of classifying the extracted data feature items, constructing corresponding data subsets according to the classification results, and importing the uploaded data into each data subset includes:

[0098] Classify the obtained data feature items, and the classification results of the data feature items include attribute feature items, parameter feature items, and association feature items;

[0099] It should be further noted that in the specific implementation process, the number of the attribute feature items is the same as that of the parameter feature items, and each attribute feature item corresponds to a parameter feature item; the association feature item is used to associate the attribute feature items with an association relationship, that is, when multiple attribute feature items are within the same factor, an association feature item corresponding to these attribute feature items is generated, and the factor can be a table, a document, a text, etc. in the specific implementation process;

[0100] Establish an attribute subset, a parameter subset, and an association subset in the data feature set respectively;

[0101] Import the obtained attribute feature items into the attribute subset, import the obtained parameter feature items into the parameter subset, and import the obtained association feature items into the association subset.

[0102] It should be further noted that in the specific implementation process, the process of optimizing the data in each data subset includes:

[0103] Mark the data ranges where each data feature item in the obtained data feature set is located in the original data;

[0104] Set several redundant comparison fields and set corresponding sequence codes for each redundant comparison field;

[0105] Match the unmarked data ranges in the original data with the redundant comparison fields, and mark the data content that is the same as the redundant comparison fields as redundant data;

[0106] Convert the original data into a data stream, and replace the data stream segment corresponding to the redundant data with the corresponding sequence code;

[0107] After completing the replacement of all data stream segments of the redundant data in the original data, the optimization of the original data is completed.

[0108] It should be further noted that in the specific implementation process, the process of importing each data subset after data optimization into the Bayesian network model and outputting the data association relationship between the data in different data sets includes:

[0109] Construct a Bayesian network model and train the constructed Bayesian network model;

[0110] Import the optimized original data in the attribute subset, parameter subset, and association subset into the trained Bayesian network model;

[0111] Output the probability of association between the attribute feature items in the attribute subset, the parameter feature items in the parameter subset, and the association feature items in the association subset;

[0112] Set a probability estimation threshold, and compare the probability of association between each output feature item with the set probability estimation threshold;

[0113] If the probability of association between each feature item is higher than the set probability estimation threshold, it means that there is a data association relationship between the corresponding feature items, otherwise there is no data association relationship;

[0114] Traverse each attribute feature item within the subset of attributes. If there are corresponding feature items in the parameter subset and the associated subset that have a data association with the attribute feature item, then mark this attribute feature item as the reference feature item;

[0115] Construct a reference subset corresponding to the reference feature item, and import other feature items that have a data association relationship with this reference feature item into the reference subset;

[0116] Number the feature items in the reference subset, denoted as j, where j = 1, 2, ……, m;

[0117] Then denote the probability of association between the feature item numbered j and this reference feature item as P j ;

[0118] Obtain the association coefficient between this reference feature item and the reference subset, denoted as Gx, where:

[0119]

[0120] where α is a constant, and 1 > α > 0, P max is the maximum value among the probabilities of association between the feature items and the reference feature item, and P min is the minimum value among the probabilities of association between the feature items and the reference feature item;

[0121] Set the association relationship threshold intervals [G1], [G2], [G3];

[0122] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G1], it indicates that this reference feature item has an association with the corresponding reference subset;

[0123] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G2], it indicates that this reference feature item has a partial association with the corresponding reference subset, and then generate an association ratio coefficient K1, where K1 is a linear coefficient and is related to the position of Gx within the range of [G2];

[0124] When the association coefficient Gx between the obtained reference feature item and each reference subset belongs to [G3], it indicates that this reference feature item has no association with the corresponding reference subset, and no operation is performed;

[0125] When the reference feature item has an association with the reference subset, generate an association index sequence of all the feature items within the reference feature item and the reference subset;

[0126] When there is a partial association between the reference feature item and the reference subset, each feature item in the reference subset is sorted according to the probability of association with the reference feature item from high to low, and then the corresponding features are selected according to the corresponding ratio from high to low according to the association ratio coefficient, and an association index sequence of the reference feature item and the features in the selected reference subset is generated;

[0127] The data association relationship between the reference feature item and each feature item in the reference subset is determined through the association index sequence.

[0128] It should be further noted that in the specific implementation process, the process of integrating and managing the data uploaded by each front-end data source according to the data association relationship between the data in each output data subset includes:

[0129] Read the association index sequences corresponding to each reference feature item, and establish integrated class databases of corresponding types respectively according to the types of the features corresponding to the association index sequences;

[0130] Import the original data corresponding to each feature and the corresponding association index sequence into the integrated class database of the corresponding type;

[0131] Mark the association index sequence as the first index sequence;

[0132] Generate a corresponding second index sequence for the integrated class database, and associate the second index sequence with each first index sequence in the corresponding integrated class database;

[0133] Associate each second index sequence with the corresponding reference feature item, thereby completing the corresponding data integration management.

[0134] In another embodiment of the present invention, a multi-source data integration system based on data mining is also disclosed, including a data management center, and the data management center is communicatively connected with a front-end data transmission module, a multi-source data processing module, an association model output module, and a data integration management module;

[0135] The front-end data transmission module is used for uploading data from the front-end data source;

[0136] The multi-source data processing module is used for extracting data feature items from the data uploaded by the front-end data source, classifying the extracted data feature items, constructing corresponding data subsets according to the classification results, and importing the uploaded data into each data subset;

[0137] The association model output module is used for optimizing the data in each data subset, importing each data subset with optimized data into the Bayesian network model, and outputting the data association relationship between the data in different data sets;

[0138] The data integration management module is used to integrally manage the data uploaded by each front-end data source according to the data association relationships among the data within each output data subset.

[0139] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any modification or equivalent replacement made to the above embodiments based on the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A multi-source data integration method based on data mining, characterized in that: The following steps are involved: Step S1: extracting data feature items from the data uploaded by the front-end data source, classifying the extracted data feature items, constructing corresponding data subsets according to the classification results, and importing the uploaded data into each data subset; Step S2: Optimizing the data in each data subset, and importing each optimized data subset into the Bayesian network model, and outputting the data association relationship between the data in different data sets; Step S3: Integrate and manage the data uploaded by each front-end data source according to the data association relationship between the data in each output data subset.

2. The multi-source data integration method based on data mining according to claim 1 is characterized in that: The process of extracting data feature items from the data uploaded by the front-end data source includes: Generate a corresponding data guide sequence according to the front-end data source of the uploaded data and the data upload time; Associating the generated data guide sequence with the data uploaded by the front-end data source to obtain original data; Constructing a corresponding data feature set according to the generated data guide sequence; Traversing the obtained raw data, and inputting the obtained raw data into the trained decision tree model; The decision tree model is used to extract features from the input raw data to obtain corresponding data feature items.

3. The multi-source data integration method based on data mining according to claim 2 is characterized in that: The training process of the decision tree model is: Collect sample data, and aggregate the collected sample data to obtain a corresponding sample data set; Classify the sample data in the sample data set to obtain the corresponding sample type; Obtain the proportion of each sample type in the sample data in the sample data set, and obtain the information entropy of the sample data set; Taking each sample type obtained as a corresponding sample feature, obtaining the number of each sample feature in the sample data set, and obtaining a number of sample feature subsets corresponding to the sample features; And construct the corresponding root node according to the obtained sample type, and obtain the information gain of each sample feature corresponding to the sample type; The sample feature corresponding to the maximum value of the obtained information gain is recorded as the root node of the decision tree; Set a corresponding internal node according to each root node, and set a corresponding node division condition for each internal node; Obtain sample features in the sample data set that meet the corresponding node division conditions, and obtain the information gain of the corresponding sample features, and then generate child nodes corresponding to the sample features; Then, the corresponding internal node is set with the child node, and the corresponding node division condition is set for each internal node, and so on; Set the corresponding quantity threshold for each sample feature; When the number of sample features that meet the corresponding node division conditions is less than the set quantity threshold, the division is stopped and the corresponding child nodes are marked as leaf nodes, thereby completing the construction of the decision tree model; The collected sample data is divided into training set, test set and validation set to train the decision tree model.

4. The multi-source data integration method based on data mining according to claim 3 is characterized in that: The process of classifying the extracted data feature items, constructing corresponding data subsets according to the classification results, and importing the uploaded data into each data subset includes: Classifying the obtained data feature items, wherein the classification results of the data feature items include attribute feature items, parameter feature items, and association feature items; Establishing attribute subsets, parameter subsets and association subsets in the data feature set respectively; The obtained attribute feature items are imported into the attribute subset, the obtained parameter feature items are imported into the parameter subset, and the obtained association feature items are imported into the association subset.

5. The multi-source data integration method based on data mining according to claim 4 is characterized in that: The process of optimizing the data in each data subset includes: Marking the data range where each data feature item in the obtained data feature set is located in the original data; Set a number of redundant comparison fields, and set a corresponding sequence code for each redundant comparison field; Match the unmarked data range in the original data with the redundant control field, and mark the data content identical to the redundant control field as redundant data; Converting the original data into a data stream, and replacing the data stream segment corresponding to the redundant data with a corresponding sequence code; After all the redundant data in the original data are replaced, the optimization of the original data is completed.

6. The multi-source data integration method based on data mining according to claim 5 is characterized in that: The process of importing each data subset that has completed data optimization into the Bayesian network model and outputting the data association relationship between data in different data sets includes: Constructing a Bayesian network model and training the constructed Bayesian network model; Import the optimized original data in the attribute subset, parameter subset and association subset into the trained Bayesian network model; Output the probability of association among the attribute feature items in the attribute subset, the parameter feature items in the parameter subset, and the association feature items in the association subset; Setting a probability estimation threshold, and comparing the probability of association between the outputted feature items with the set probability estimation threshold; If the probability of association between the feature items is higher than the set probability estimation threshold, it means that there is a data association relationship between the corresponding feature items, otherwise there is no data association relationship; Traverse each attribute feature item in the attribute subset, if there is a corresponding feature item in the parameter subset and the associated subset that has a data relationship with the attribute feature item, then record the attribute feature item as the reference feature item; Constructing a benchmark subset corresponding to the benchmark feature item, and importing other feature items that have a data association relationship with the benchmark feature item into the benchmark subset; According to the probability that there is an association between the feature item in the benchmark subset and the benchmark feature item, a correlation coefficient Gx between the benchmark feature item and the benchmark subset is obtained; Set the association relationship threshold interval [G1], [G2], [G3]; When the correlation coefficient between the obtained benchmark feature item and each benchmark subset is Gx∈[G1], it means that the benchmark feature item is correlated with the corresponding benchmark subset; When the obtained correlation coefficient between the benchmark feature item and each benchmark subset is Gx∈[G2], it means that the benchmark feature item is correlated with the corresponding benchmark subset part, and the correlation ratio coefficient is generated; When the obtained correlation coefficient between the benchmark feature item and each benchmark subset is Gx∈[G3], it means that there is no correlation between the benchmark feature item and the corresponding benchmark subset, and no operation is performed.

7. The multi-source data integration method based on data mining according to claim 6 is characterized in that: When the benchmark feature item is associated with the benchmark subset, an associated index sequence of the benchmark feature item and all feature items in the benchmark subset is generated; When the benchmark feature item has a partial correlation with the benchmark subset, the feature items in the benchmark subset are sorted from high to low according to the probability of being associated with the benchmark feature item, and then the corresponding features are selected from high to low according to the corresponding proportion according to the correlation ratio coefficient, and the correlation index sequence of the benchmark feature item and the features in the selected benchmark subset is generated; The data association relationship between the benchmark feature item and each feature item in the benchmark subset is determined by associating the index sequence.

8. The multi-source data integration method based on data mining according to claim 7 is characterized in that: The process of integrating and managing the data uploaded by each front-end data source according to the data association relationship between the data in each output data subset includes: Read the associated index sequence corresponding to each benchmark feature item, and establish an integrated class database of the corresponding type according to the type of feature corresponding to the associated index sequence; Import the original data corresponding to each feature and the corresponding associated index sequence into the integrated database of the corresponding type; The associated index sequence is marked as a first index sequence, a corresponding second index sequence is generated for the integrated database, and the second index sequence is associated with each first index sequence in the corresponding integrated database; each second index sequence is associated with a corresponding benchmark feature item, thereby completing the corresponding data integration management.

9. A multi-source data integration system applied to a multi-source data integration method based on data mining as claimed in any one of claims 1 to 8, comprising a data management center, characterized in that: The data management center is communicatively connected to a front-end data transmission module, a multi-source data processing module, a correlation model output module, and a data integration management module; The front-end data transmission module is used for uploading data from the front-end data source; The multi-source data processing module is used to extract data feature items from the data uploaded by the front-end data source, classify the extracted data feature items, construct corresponding data subsets according to the classification results, and import the uploaded data into each data subset; The association model output module is used to optimize the data in each data subset, import each data subset that has completed data optimization into the Bayesian network model, and output the data association relationship between the data in different data sets; The data integration management module is used to integrate and manage the data uploaded by each front-end data source according to the data association relationship between the data in each output data subset.

Citation Information

Patent Citations

  • General association rule analysis method based on deep learning

    CN115936065A

  • Systems and methods to de-duplicate features for machine learning model

    US20170185911A1

Cited By

  • Multi-source data acquisition processing method and system based on artificial intelligence

    CN121277942A