Multi-source heterogeneous data analysis method and system based on machine learning and electronic equipment
Through a multi-source heterogeneous data analysis method based on machine learning, the data flow relationship is dynamically captured, erroneous fields are repaired, and the visualization and traceability of data flow are achieved. This solves the problem of inaccurate lineage analysis in traditional data governance, improves the intelligence and security of data governance, and supports cross-institutional data sharing and modeling.
Patent Information
- Application Number
- CN202511239952.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Traditional data governance methods are unable to cope with the dual pressures of dynamic changes in data sources and privacy protection in a multi-source heterogeneous data environment, resulting in inaccurate lineage analysis, affecting the accuracy of credit approval and risk assessment, and the existence of data silos and privacy leakage risks in cross-institutional collaboration.
By obtaining data flow logs from multi-source heterogeneous databases, using graph databases to store lineage information, combining node label assignment and graph traversal algorithms to dynamically capture association relationships, judging inconsistent dependencies, repairing erroneous fields based on data cleaning algorithms, and constructing an optimized lineage information set, we can achieve visualization and traceability of data flow, and ensure data security through differential privacy algorithms.
It improves the accuracy of bloodline information, enhances the intelligence and automation level of data governance, ensures the accuracy and security of data association, supports cross-institutional data sharing and federated learning modeling, and improves the accuracy of credit approval and risk assessment.
Smart Images

Figure CN120804597A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data governance, in particular to a multi-source heterogeneous data analysis method and system based on machine learning and an electronic device. BACKGROUND
[0002] At present, with the diversification of data sources and the explosive growth of data volume in financial institutions, data governance plays a crucial role in ensuring data security, improving decision-making efficiency and meeting regulatory requirements. Traditional data governance methods often struggle to cope with the dual pressures of dynamic data sources and privacy protection when faced with complex data environments. They usually rely on static rules or manual intervention, making it difficult to adapt to scenarios where data sources change frequently, and in data sharing, the risk of privacy leakage can lead to compliance issues. In cross-institutional collaboration, data silos further exacerbate the governance difficulty. In addition, in multi-source heterogeneous data governance, dynamic blood analysis and federated learning security sharing are key technical challenges. In cross-system data integration, if blood analysis cannot accurately capture real-time changes in data flow, it may lead to incorrect data association, which in turn affects the accuracy of credit approval or risk assessment. This inaccuracy further exacerbates the difficulty of federated learning security sharing.
[0003] The above information disclosed in the background section is only intended to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0004] Therefore, the present disclosure provides a multi-source heterogeneous data analysis method based on machine learning, which can improve the accuracy of blood analysis while ensuring data security.
[0005] In a first aspect, embodiments of the present application provide a multi-source heterogeneous data analysis method based on machine learning, the method comprising: obtaining a data flow transfer log in a multi-source heterogeneous database, the data flow transfer log comprising a timestamp, an IP address of a data source, an ID of a flow transfer node, and a size of a data packet; determining a blood relationship information set according to the data flow transfer log, the blood relationship information set comprising a data source, a first data flow transfer path, and a flow transfer node dependency relationship; storing the first data flow transfer path and the flow transfer node dependency relationship by using a graph database to generate a graph database structure; assigning a corresponding unique identifier to each flow transfer node in the graph database structure based on a node label assignment algorithm to generate a node identifier set; determining a dynamic association relationship set between each flow transfer node based on the node identifier set according to a graph traversal algorithm; judging whether there is an inconsistent dependency relationship in the dynamic association relationship set; if it is judged that there is the inconsistent dependency relationship, obtaining an inconsistent dependency field; determining an error field set according to the inconsistent dependency field, the error field comprising a missing value, an abnormal value, and a format inconsistency; repairing the error field set based on a data cleaning algorithm to generate a repaired data set; constructing a second data flow transfer path according to the repaired data set; judging whether the second data flow transfer path has path integrity based on a path traversal algorithm; if it is judged that the path integrity is possessed, generating an optimized blood relationship information set according to the second data flow transfer path.
[0006] In the second aspect, an embodiment of the present application provides a multi-source heterogeneous data analysis system based on machine learning, which includes a first acquisition module, a first determination module, a first generation module, a second generation module, a second determination module, a first judgment module, a second acquisition module, a third determination module, a repair module, a construction module, a second judgment module, and a third generation module. Among them, the first acquisition module is used to obtain the data flow log in the multi-source heterogeneous database, and the data flow log includes a timestamp, an IP address of the data source, an ID of the flow node, and the size of the data packet; the first determination module is used to determine the lineage information set based on the data flow log, and the lineage information set includes the data source, the first data flow path, and the flow node dependency; the first generation module is used to use a graph database to store the first data flow path and the flow node dependency to generate a graph database structure; the second generation module is used to assign a corresponding unique identifier to each of the flow nodes in the graph database structure based on a node label allocation algorithm to generate a node identifier set; the second determination module is used to determine the dynamic relationship between each of the flow nodes based on the node identifier set based on a graph traversal algorithm. dynamic association relationship set; a first judgment module, used to judge whether there is an inconsistent dependency in the dynamic association relationship set; a second acquisition module, used to obtain the inconsistent dependency field if it is judged that the inconsistent dependency exists; a third determination module, used to determine the error field set based on the inconsistent dependency field, the error field including missing values, abnormal values and format inconsistency; a repair module, used to repair the error field set based on the data cleaning algorithm to generate a repair data set; a construction module, used to construct a second data flow path based on the repair data set; a second judgment module, used to judge whether the second data flow path has path integrity based on the path traversal algorithm; a third generation module, used to generate an optimized lineage information set according to the second data flow path if it is judged that the path integrity is possessed.
[0007] The embodiment of the present application provides a multi-source heterogeneous data analysis method based on machine learning. Through analyzing and processing the data flow transfer log in the multi-source heterogeneous database, information such as a timestamp, an IP address of a data source, and an ID of a transfer node is accurately extracted, and a first transfer path and a node dependency relationship in a blood relationship information set are stored in combination with a graph database, so that the blood relationship information is visualized and traceable, and the data source and the transfer context are clear and transparent. With the aid of a node label allocation algorithm and a graph traversal algorithm, a dynamic association relationship set between the transfer nodes is determined, the association relationship between the transfer nodes can be dynamically captured, inconsistent dependency relationships can be discovered in time, and inconsistent dependency fields can be located, so that the accuracy of data association is ensured. An error field is determined through the inconsistent dependency field, and the error field is repaired through a data cleaning algorithm to generate a repaired data set, and a second data transfer path is constructed according to the repaired data set, and after the path integrity is verified, an optimized blood relationship information set is generated, so that the accuracy of the blood relationship information is effectively improved, the real-time changes of data transfer can be accurately captured, and the accuracy of credit approval or risk assessment is improved. High-quality and reliable foundations are provided for subsequent cross-institutional data sharing, federal learning modeling and the like, and the intelligentization and automation level of data governance is enhanced, and efficient response to dynamic changes and complex association challenges in a multi-source heterogeneous data environment is facilitated. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, a brief introduction will be given to the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0009] Figure 1 is a flowchart of a multi-source heterogeneous data analysis method based on machine learning provided by an exemplary embodiment of the present application.
[0010] Figure 2 is a flowchart of a multi-source heterogeneous data analysis method based on machine learning provided by another exemplary embodiment of the present application.
[0011] Figure 3 is a flowchart of a multi-source heterogeneous data analysis method based on machine learning provided by another exemplary embodiment of the present application.
[0012] Figure 4 is a flowchart of a multi-source heterogeneous data analysis method based on machine learning provided by another exemplary embodiment of the present application.
[0013] Figure 5is a flowchart of a method for multi-source heterogeneous data analysis based on machine learning provided by another example embodiment of the present application.
[0014] Figure 6 is a flowchart of a method for multi-source heterogeneous data analysis based on machine learning provided by another example embodiment of the present application.
[0015] Figure 7 is a flowchart of a method for multi-source heterogeneous data analysis based on machine learning provided by another example embodiment of the present application.
[0016] Figure 8 is a flowchart of a method for multi-source heterogeneous data analysis based on machine learning provided by another example embodiment of the present application. DETAILED DESCRIPTION
[0017] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout.
[0018] The terms "one", "a", "said" are used to indicate that there is one or more of the elements / components / etc.; the terms "include" and "has" are used to indicate an open-ended inclusion in such a way that additional elements / components / etc. can be present in addition to those listed; the terms "first" and "second" etc. are used as labels for the purpose of distinguishing between elements, and are not meant to limit the number of such elements.
[0019] At present, with the diversification of data sources and the explosive growth of data volume in financial institutions, data governance plays a crucial role in ensuring data security, improving decision-making efficiency, and meeting regulatory requirements. Traditional data governance methods often struggle to cope with the dual pressures of dynamic data sources and privacy protection in complex data environments. They usually rely on static rules or manual intervention, making it difficult to adapt to scenarios where data sources change frequently, and may lead to compliance issues due to privacy leakage risks during data sharing, especially in cross-institutional collaboration, where data silos further exacerbate governance difficulties. In addition, in multi-source heterogeneous data governance, dynamic bloodline analysis and federated learning secure sharing are key technical challenges. In cross-system data integration, if bloodline analysis cannot accurately capture real-time changes in data flow, it may lead to incorrect data association, which in turn affects the accuracy of credit approval or risk assessment, further exacerbating the problem of federated learning secure sharing.
[0020] A machine learning-based multi-source heterogeneous data analysis method is provided in the embodiments of the present disclosure, as shown in Figure 1 The machine learning-based multi-source heterogeneous data analysis method can include the following steps: Step S110: Obtain data flow transfer logs in a multi-source heterogeneous database, the data flow transfer logs including timestamps, IP addresses of data sources, IDs of transfer nodes, and sizes of data packets; Step S120: Determine a blood relationship information set according to the data flow transfer logs, the blood relationship information set including data sources, first data flow transfer paths, and transfer node dependency relationships; Step S130: Store the first data flow transfer paths and the transfer node dependency relationships by using a graph database to generate a graph database structure; Step S140: Assign corresponding unique identifiers to each transfer node in the graph database structure based on a node label assignment algorithm to generate a node identifier set; Step S150: Determine a dynamic association relationship set between each transfer node according to the node identifier set based on a graph traversal algorithm; Step S160: Determine whether there is an inconsistent dependency relationship in the dynamic association relationship set; Step S170: If it is determined that there is an inconsistent dependency relationship, obtain an inconsistent dependency field; Step S180: Determine an error field set according to the inconsistent dependency field, the error field including missing values, abnormal values, and inconsistent formats; Step S190: Repair the error field set based on a data cleaning algorithm to generate a repaired data set; Step S192: Construct a second data flow transfer path according to the repaired data set; Step S194: Determine whether the second data flow transfer path has path integrity based on a path traversal algorithm; Step S196: If it is determined that the second data flow transfer path has path integrity, generate an optimized blood relationship information set according to the second data flow transfer path.
[0021] According to the machine learning-based multi-source heterogeneous data analysis method provided by the present disclosure, the method can obtain data flow transfer logs in a multi-source heterogeneous database, the data flow transfer logs including timestamps, IP addresses of data sources, IDs of flow transfer nodes, and sizes of data packets; determine a blood relationship information set according to the data flow transfer logs, the blood relationship information set including data sources, first data flow transfer paths, and flow transfer node dependency relationships; store the first data flow transfer paths and the flow transfer node dependency relationships by using a graph database to generate a graph database structure; assign corresponding unique identifiers to each flow transfer node in the graph database structure based on a node label assignment algorithm to generate a node identifier set; determine a dynamic association relationship set between each flow transfer node based on a graph traversal algorithm according to the node identifier set; determine whether there is an inconsistent dependency relationship in the dynamic association relationship set; if it is determined that there is an inconsistent dependency relationship, obtain an inconsistent dependency field; determine an error field set according to the inconsistent dependency field, the error field including missing values, abnormal values, and inconsistent formats; repair the error field set based on a data cleaning algorithm to generate a repaired data set; construct a second data flow transfer path according to the repaired data set; determine whether the second data flow transfer path has path integrity based on a path traversal algorithm; and if it is determined that the second data flow transfer path has path integrity, generate an optimized blood relationship information set according to the second data flow transfer path.
[0022] In the above method, by analyzing and processing the data flow transfer logs in the multi-source heterogeneous database, the information including timestamps, IP addresses of data sources, IDs of flow transfer nodes, and the like can be accurately extracted, and the first flow transfer path and the node dependency relationship in the blood relationship information set are stored in the graph database to realize the visualization and traceability of the blood relationship information, so that the data sources and the flow transfer context are clear and transparent. By means of the node label assignment algorithm and the graph traversal algorithm, the dynamic association relationship set between each flow transfer node is determined, which can dynamically capture the association relationship between the flow transfer nodes, discover the inconsistent dependency relationship in time, locate the inconsistent dependency field, and ensure the accuracy of data association. The error field is determined by the inconsistent dependency field, and the error field is repaired by the data cleaning algorithm to generate a repaired data set, and a second data flow transfer path is constructed according to the repaired data set. After verifying that the path integrity is passed, an optimized blood relationship information set is generated, which effectively improves the accuracy of the blood relationship information, accurately captures the real-time changes of data flow transfer, and further improves the accuracy of credit approval or risk assessment. It provides high-quality and reliable basis for subsequent cross-institutional data sharing, federal learning modeling, and the like, while enhancing the intelligentization and automation level of data governance, and helping to efficiently cope with dynamic changes and complex association challenges in the multi-source heterogeneous data environment.
[0023] The steps of the machine learning-based multi-source heterogeneous data analysis method provided by the present embodiment will be described in detail as follows: In an embodiment of the present disclosure, in step S110, the data flow transfer log in the multi-source heterogeneous database is obtained, and the data flow transfer log includes a timestamp, an IP address of a data source, an ID of a flow transfer node, and a size of a data packet. Specifically, the data flow transfer log is obtained from the multi-source heterogeneous data system, and can be implemented by a distributed message queue such as Kafka. A Kafka cluster (3 nodes, throughput of 100,000 logs per second) is deployed, and each data flow transfer log includes a timestamp, an IP of a data source, an ID of a flow transfer node, and a size of a data packet (for example, 1024 bytes).
[0024] For example, in the multi-source heterogeneous data system of a certain financial institution, a flow transfer log of a user credit data is collected through the Kafka cluster, and the content is: “time_stamp: 2025-08-01-14:30, source_ip: 192.168.1.101, node_id: NODE_007, packet_size: 1536 bytes”. Among them, the timestamp “2025-08-01-14:30:22.567” accurately records the time of data flow transfer, the data source IP “192.168.1.101” points to the user information database server, the flow transfer node ID “NODE_007” corresponds to the processing node of the credit approval system, and the data packet size is “1536 bytes”. It can be seen that the data flow transfer log includes user basic information, credit score and other data content. The data flow transfer log is collected in real time through a distributed log collection tool, and provides an original basis for analyzing data sources, flow transfer paths and node dependency relationships.
[0025] In an embodiment of the present disclosure, in step S120, the blood relationship information set is determined according to the data flow transfer log, and the blood relationship information set includes a data source, a first data flow transfer path, and a flow transfer node dependency relationship. Specifically, if the data flow transfer log records “timestamp: 2025-08-01-09:00:00, data source IP address: 192.168.1.102, flow transfer node ID: flow transfer node A, B, C, data packet size: 2048 bytes”, then in the blood relationship information set formed after analysis, the data source (such as “192.168.1.101”), the first data flow transfer path (such as “flow transfer node A -> flow transfer node B -> flow transfer node C”), and the node dependency relationship (such as “flow transfer node B depends on the output of flow transfer node A, and flow transfer node C depends on the output of flow transfer node B”) are extracted.
[0026] In the above method, by extracting and structuring the data source, the first data flow path and the node dependency relationship from the data flow transfer log, a clear blood relationship information set is formed, not only realizing the full-link tracing of data from the source to the terminal, making the "origin and destination" of each piece of data traceable, laying a standardized foundation for subsequent graph database structure storage and dynamic correlation analysis, improving the transparency and controllability of data governance in a multi-source heterogeneous data environment, and providing a reliable blood basis for cross-system data collaboration and model training.
[0027] In an embodiment of the present disclosure, in step S130, the first data flow path and the flow node dependency relationship are stored by using a graph database to generate a graph database structure. Specifically, if the first data flow path is: flow node A -> flow node B -> flow node C, and the flow node dependency relationship is: flow node B depends on the output of flow node A, and flow node C depends on the output of flow node B. In the graph database (such as Neo4j), flow nodes A, B and C are used as nodes, the first data flow path (such as A→B→C) is represented by a directed edge, and weights are added to the edge (such as the weight of 0.8 from flow node A to B representing 80% of the data successfully flowing), forming a "node-edge-weight" graph database structure.
[0028] In the above method, the "node-edge" structure of the graph database naturally adapts to the mesh relationship of data flow, can intuitively store the directionality of the path and the strength of the node dependency (the strength of the node dependency is judged according to the size of the weight value), and compared with the traditional table structure, can more easily trace the full link of data, can significantly improve the visual management efficiency and traceability of multi-source heterogeneous data flow, and support the precision and efficiency of data governance.
[0029] In an embodiment of the present disclosure, in step S140, based on a node label assignment algorithm, each flow node in the graph database structure is given a corresponding unique identifier to generate a node identifier set. Specifically, if the graph database structure includes order generation nodes, payment processing nodes, inventory update nodes and other flow nodes, each flow node is assigned a unique identifier, such as assigning the order generation node the identifier "8f3b2a1c-4d5e-11eb-ae93-0242ac130002", and assigning the payment processing node the identifier "9g4c3b2d-6e7f-22fc-bf04-1353bd241113". These one-to-one identifiers are stored in the node attributes to form a node identifier set containing the unique identifiers of all flow nodes.
[0030] In the above method, the unique identification ensures the uniqueness of each flow transfer node in the graph database structure, accurately distinguishes even if there is repetition in the flow transfer node name or attribute, and provides a reliable basis for the analysis of the dynamic correlation relationship between the flow transfer nodes; through the unique identification, a specific flow transfer node and its associated path can be quickly located, the complete link of the data flow transfer is facilitated to be tracked, the accuracy and efficiency of the node management in the graph database structure are improved, and meanwhile, a clear node identification basis is provided for subsequent operations such as detecting inconsistent dependency relationship and repairing error field, thereby guaranteeing the rigor of the multi-source heterogeneous data flow transfer analysis.
[0031] In an embodiment of the present disclosure, in step S150, the dynamic correlation relationship set between the flow transfer nodes is determined based on the graph traversal algorithm according to the node identification set. Specifically, if the node identification set contains an order node (such as a1b2c3d4), a payment node (such as e5f6g7h8), and an inventory node (such as i9j0k1l2), the graph traversal algorithm is used to traverse the graph database structure, and the path weight is calculated along the directed edge from the order node, such as order→payment (path weight 0.9), which means that 90% of the orders trigger the payment. The dynamic correlation relationship set of “order node→payment node (path weight 0.9), payment node→inventory node (path weight 0.8), and order node→inventory node (path weight 0.3)” is obtained, which reflects the situation that the dependency strength between the nodes changes in real time with the business.
[0032] In the above method, the dynamic correlation relationship between the flow transfer nodes can be efficiently mined through the graph traversal algorithm, including direct correlation and indirect correlation, and the weight value can quantify the dependency strength and accurately capture the dynamic changes; this provides a data basis for subsequent detection of inconsistent dependency relationship, can timely discover abnormal correlation, and makes the correlation logic between the flow transfer nodes clearer, thereby improving the depth and accuracy of the multi-source heterogeneous data flow transfer analysis, and providing a reliable basis for the monitoring and optimization of the dynamic relationship in data governance.
[0033] In an embodiment of the present disclosure, in steps S160-S170, it is judged whether there is an inconsistent dependency relationship in the dynamic correlation relationship set, and if it is judged that there is an inconsistent dependency relationship, the inconsistent dependency field is obtained, and the method further includes the following steps, as shown in Figure 2 The specific content is as follows: Step S210: obtaining the edge weight value between the current flow transfer nodes, the edge weight value being used to represent the dependency strength between the flow transfer nodes; Step S220: obtaining the edge weight value between the current flow transfer nodes, the edge weight value being used to represent the dependency strength between the flow transfer nodes; Step S230: judging whether the difference between the current edge weight value and the initial edge weight value is greater than a preset deviation threshold; Step S240: if the judgment is greater than the preset deviation threshold, it is judged that there is an inconsistent dependency relationship; Step S250: taking any flow conversion node as the starting point, recording all third data flow conversion paths corresponding thereto, and calculating the path weight values corresponding to each third data flow conversion path; Step S260: judging whether the difference between the path weight values corresponding to each third data flow conversion path and the historical data flow conversion path is greater than a second preset deviation threshold; Step S270: if the judgment is greater than the second preset deviation threshold, it is marked as a deviation conversion path; Step S280: determining the inconsistent dependency field according to the deviation conversion path.
[0034] Specifically, if the order node (such as o123) and the payment node (such as p456) in the graph database of an e-commerce platform, the current edge weight value is detected to be 0.3. Based on the depth-first search algorithm to traverse the graph database structure, the initial edge weight value is obtained to be 0.8, and the difference between the current edge weight value and the initial edge weight value is calculated to be 0.5, which is greater than the first preset deviation threshold 0.2, so it is judged that there is an inconsistent dependency relationship. Subsequently, taking the order node as the starting point, two third data flow conversion paths are recorded as “order→payment→inventory” and “order→discount→payment”, and the path weight values of each path are calculated to be 1.5 and 0.9 respectively, and the path weight values corresponding to the historical data flow conversion path are 0.5 and 0.3 respectively, and the difference between the path weight values corresponding to the two third data flow conversion paths and the weight values corresponding to the historical data flow conversion path is 1 and 0.6 respectively. For example, the second preset deviation threshold is 0.8, it can be seen that the first third data flow conversion path (order→payment→inventory) is greater than 0.8 and is marked as a deviation conversion path, and the second third data flow conversion path (order→discount→payment) is less than 0.8, so no operation is needed. According to the above deviation conversion path, it is determined that the “inventory” of the payment node is the inconsistent dependency field.
[0035] In the above method, by judging the deviation of edge weight and path weight, the inconsistent dependency field in the inconsistent dependency relationship is accurately located, the automatic detection of the dependency relationship exception is realized, the subjectivity of manual judgment is avoided, and the consistency and reliability of the correlation in the multi-source heterogeneous data flow conversion are effectively guaranteed.
[0036] In an embodiment of the present disclosure, in step S180, the error field set is determined according to the inconsistent dependency field, and the error field includes missing values, abnormal values, and format inconsistencies. Specifically, if the “payment amount field” is determined as the inconsistent dependency field in the transaction data flow of the e-commerce platform, the inconsistent dependency field has three types of error fields in the data flow from the payment node (such as P789) to the order confirmation node (such as C123). For example, the “payment amount” field is empty (missing value), the display amount is “-500 yuan” (unreasonable abnormal value), the format is “300” (lacks currency unit, and is inconsistent with the standard format “300 yuan”), and these collectively constitute the error field set.
[0037] In the above method, through the error field types such as missing values, abnormal values, and format inconsistencies, the data cleaning is more targeted; such accurate positioning can avoid resource waste caused by indiscriminate cleaning processing, ensure efficient and reliable subsequent repair process, lay a high-quality data foundation for generating optimized blood relationship information set, and ultimately improve the accuracy and consistency of multi-source heterogeneous data flow.
[0038] In an embodiment of the present disclosure, in step S190, based on the data cleaning algorithm, the error field set is repaired to generate a repaired data set, which further includes the following steps, as shown in Figure 3 The specific content is as follows: Step S310: standardizing the error field set to generate a first standardized data set; Step S320: calculating the covariance matrix of the first standardized data set to determine the eigenvalues and eigenvectors of the covariance matrix; Step S330: sorting the eigenvalues in the order from the maximum value to the minimum value, and selecting the eigenvectors corresponding to the first five eigenvalues in the sorting result to generate a feature matrix; Step S340: determining the corresponding repair scheme according to the type of error field in the feature matrix.
[0039] Specifically, the error field set of certain financial credit data includes missing values (such as the "income proof" field being empty), outliers (such as the "credit score" being 999 points, far exceeding the reasonable range), and inconsistent formats (such as "date of birth" existing both "1994 / 10 / 01" and "1994-10-01"). Standardizing these error fields can normalize the numerical fields to the [0, 1] interval and generate a first standardized data set. The covariance matrix of the first standardized data set is calculated to obtain eigenvalues of 5.2, 3.8, 2.1, 1.5, and 0.9 (the remaining eigenvalues are less than 0.5), and the eigenvectors corresponding to the first five eigenvalues are selected to generate a feature matrix. Based on this, the corresponding repair scheme is determined according to the type of the error field in the feature matrix.
[0040] In the above method, the field magnitude difference in the error field set is eliminated through standardization processing, ensuring the consistency of subsequent calculations; through covariance matrix and eigenvector analysis, key features and correlation can be extracted from error fields, making the repair scheme more targeted and avoiding blind processing. This can accurately repair error fields and generate high-quality repair data sets, providing reliable data support for constructing complete second data flow paths and optimized bloodline information sets, and ensuring the accuracy and consistency of data flow.
[0041] In an embodiment of the present disclosure, in step S340, determining the corresponding repair scheme according to the type of the error field in the feature matrix further includes the following steps, as shown in Figure 4 The specific content is as follows: Step S410: When the feature matrix has missing values, a plurality of neighbor values corresponding to the missing values are obtained based on the K-Nearest Neighbor interpolation algorithm; Step S420: Weighted average operation is performed on the plurality of neighbor values, and the operation result is taken as the missing value; Step S430: When the feature matrix has outliers, the standard score of the outlier is calculated; Step S440: Determine whether the standard score of the outlier exceeds a preset standard score threshold; Step S450: If it is determined that the preset standard score threshold is exceeded, the corresponding outlier is removed; Step S460: When the feature matrix has inconsistent formats, the inconsistent formats are adjusted to a unified format using a regular expression.
[0042] Specifically, if there are missing values in the "monthly income" field, abnormal values in the "loan amount" field, and inconsistent formats in the "account opening date" field in the feature matrix of the bank credit data, for the missing values in the "monthly income", the K-Nearest Neighbor interpolation algorithm is used to select 5 neighbor (user) samples with similar credit ratings and occupation types, whose monthly incomes are 8000 yuan, 7500 yuan, 9000 yuan, 8500 yuan, and 7800 yuan, respectively. The weighted average is calculated to be 8160 yuan, which is used to fill in the missing values. For the "10000000 yuan" in the "loan amount", the standard score is calculated to be 3.2, which exceeds the preset standard score threshold 3.0, so it is determined as an abnormal value and is removed. For the formats of "2020.05.12" and "2020 / 05 / 12" in the "account opening date", regular expressions are used to adjust them to "2020-05-12".
[0043] In the above method, the K-Nearest Neighbor interpolation algorithm is used to fill in the missing values based on similar sample features, ensuring data relevance; the standard score is used to determine and remove abnormal values, avoiding extreme data interference; and the regular expression is used to unify the format, ensuring field standardization. The three work together to improve data quality. This method makes the repaired data more suitable for subsequent modeling and analysis requirements, provides a reliable foundation for generating complete second data flow path and optimized blood relationship information set, and effectively supports precise management and efficient flow of multi-source heterogeneous data.
[0044] In an embodiment of the present disclosure, in step S192, a second data flow path is constructed according to the repaired data set. Specifically, the second data flow path is constructed based on the repaired data, eliminating the error fields in the first data flow path, ensuring the continuity and accuracy of data flow, laying the foundation for generating the optimized blood relationship information set, and finally improving the reliability and traceability of data flow in the multi-source heterogeneous data system, supporting data collaboration and predictive model training across nodes and systems.
[0045] In an embodiment of the present disclosure, in steps S194-S196, it is determined whether the second data flow path has path integrity based on a path traversal algorithm. If it is determined that the path has integrity, the optimized blood relationship information set is generated according to the second data flow path. Specifically, if the second data flow path of a certain financial institution is “user account system→risk audit node→credit approval node→lending record system”, the path traversal algorithm (such as depth-first search DFS) is used to start from the source node (user account system), traverse each flow node and associated edge layer by layer, check whether there is a broken or missing node or field, and if each flow is found to be complete after traversal, and the input field of each flow node can match the output field of the previous flow node (such as risk score accurately flowing from the audit node to the approval node), it is determined that the path has integrity. If it is found in the traversal that a certain flow node (such as the credit approval node) is missing the necessary “user credit field”, it is determined that the path is not complete and needs to be returned to the data cleaning stage for repair. When the path is confirmed to be complete, the optimized blood relationship information set is generated according to the second data flow path.
[0046] In the above method, the path integrity verification ensures that the repaired data can be reliably transferred in the whole chain, avoids information loss, and improves the completeness of blood relationship information. The optimized blood relationship information set contains the relationship between the repaired data flow path and the field, making the data blood relationship clearer and traceable, providing high-quality blood relationship basis for subsequent cross-institutional data sharing and federal learning modeling, and enhancing the accuracy and credibility of multi-source heterogeneous data governance, supporting the transparency and efficiency of data flow.
[0047] In an embodiment of the present disclosure, in step S196, if it is determined that the path has integrity, the optimized blood relationship information set is generated according to the second data flow path, and then the following steps are included, as shown in Figure 5 The specific content is as follows: Step S510: standardizing the optimized blood relationship information set to generate a second standardized data set; Step S520: calculating the privacy budget of the second standardized data set based on a differential privacy algorithm; Step S530: determining whether the privacy budget exceeds a preset privacy budget threshold; Step S540: if it is determined that the preset privacy threshold has not been exceeded, adding Gaussian noise to the second standardized data set to generate a perturbed data set; Step S550: based on a federal learning framework, fusing local model parameters of each institution in the perturbed data set to generate a shared model parameter set; Step S560: determining whether the shared model parameter set passes the consistency verification of each institution; Step S570: if it is judged that the consistency verification of each institution is satisfied, a secure shared data set is generated.
[0048] Specifically, assuming that the optimized blood relationship information set includes transaction amount (1000-5000 yuan), node flow duration (5-30 minutes), and node dependency strength (0.1-0.9), a second standardized data set is generated by Z-score standardization, such as transaction amount 4000 yuan standardized to 1.0, node flow duration 20 minutes standardized to 1.0, and node dependency strength 0.7 standardized to 1.0. The privacy budget of the second standardized data set is calculated based on the differential privacy algorithm. In differential privacy, the privacy budget ε is used to measure the strength of privacy protection (the smaller ε is, the stricter the privacy protection is, and the data availability is slightly reduced). For example, ε=1.2, which does not exceed the preset privacy budget threshold 1.5, and a perturbation data set can be generated by adding Gaussian noise with a standard deviation of about 0.417. For example, a certain transaction amount standardized value 1.0 is perturbed to 1.03. Then, based on the federated learning framework, the local model parameters of the three financial institutions are fused, and the parameters w1=[0.3, 0.2], w2=[0.32, 0.18], and w3=[0.28, 0.22] of institutions A, B, and C are obtained. The shared parameter set w=[0.3, 0.2] is obtained, and after the consistency verification of each institution (the deviation is less than or equal to 5%), the secure shared data set is generated by integrating the shared parameters and the statistical characteristics of the perturbation data.
[0049] In the above method, differential privacy and noise addition are used to realize data privacy protection and reduce the risk of data privacy leakage. The federated learning framework can break through the data island problem, enable secure data sharing between multiple institutions, and update in a timely manner. Thus, the final secure shared data set not only supports multi-institution joint modeling (such as credit risk prediction), but also prevents privacy leakage, significantly improving the utilization efficiency and security of multi-source heterogeneous data in cross-institution scenarios.
[0050] In an embodiment of the present disclosure, after step S570, if it is judged that the consistency verification of each institution is satisfied, a secure shared data set is generated, the method further includes the following steps, as shown in Figure 6 The specific content is as follows: Step S610: determining a training feature set according to the secure shared data set; Step S620: determining a first initial model parameter according to the training feature set based on a distributed gradient descent algorithm; Step S630: sending the first initial model parameter to each institution; Step S640: each institution calculates a corresponding local gradient update value based on local data; Step S650: aggregating the local gradient update values to generate a global gradient update value; Step S660: updating the first initial model parameter according to the global gradient update value to generate a final model parameter; Step S670: generating a cross-institutional prediction model according to the final model parameter.
[0051] Specifically, when extracting training features from the secure shared dataset, a principal component analysis (PCA) algorithm can be used to extract features from the cross-institutional shared medical dataset. For example, the secure shared dataset includes 5 indicators (age, blood pressure, blood sugar, BMI, and cholesterol) of 1000 patients, and each indicator is normalized to the range of 0 to 1. The PCA algorithm extracts the first two principal components by calculating the covariance matrix, retains 80% of the variance, and generates a feature vector matrix (1000x2). The scikit-learn library of Python can be used to obtain the reduced training feature set, which can reduce the computational complexity and protect the privacy of the original data.
[0052] For example, three medical institutions each hold part of the data, and the first initial model parameter for training can be determined according to the reduced training feature set, and the model parameter is sent to the three medical institutions respectively, and distributed training is realized through a federated learning framework (such as Tensor Flow Federated). Each institution calculates the gradient of the logistic regression model locally, i.e., each local gradient update value. Assuming that the initial learning rate is 0.01, the batch size is 32, and the iteration is 100 times. The local gradient update values are aggregated to generate a global gradient update value, for example, [0.23, -0.15], which is transmitted to the central server after encryption. The encrypted gradient transmission mechanism can use homomorphic encryption (such as Paillier algorithm), and the central server aggregates the ciphertext gradient and updates the first initial model parameter after decryption, ensuring that the data does not leave the local. Finally, a cross-institutional prediction model is obtained through aggregation training. For example, a cross-institutional logistic regression model can be obtained, which can output the disease risk probability of a patient.
[0053] In the above method, the extraction of the training feature set reduces the redundant information, the encrypted transmission ensures data security, the distributed training improves the model generalization ability, and is suitable for cross-institutional collaboration scenarios. In addition, each institution only calculates the local gradient based on local data without sharing the original data, which strictly protects user privacy and complies with data security regulations, while integrating the data features of multiple institutions makes the prediction model more versatile. Distributed computing not only reduces the computing power pressure of a single institution, but also breaks down data silos, making cross-industry collaboration prediction such as finance, payment, and e-commerce possible. When a new institution is added, only the initial parameters need to be synchronized and the local gradient needs to be calculated, without the need to reconstruct the prediction model, which is suitable for dynamically expanding multi-institutional collaboration scenarios, effectively balancing data security and model performance, and improving the efficiency and practicality of cross-institutional modeling.
[0054] In an embodiment of the present disclosure, after the step S670 of generating the cross-institution prediction model according to the final model parameters, the following steps are further included, as shown in Figure 7 The specific content is as follows: Step S710: determining whether the prediction accuracy of the cross-institution prediction model exceeds the first preset percentage threshold; Step S720: if it is determined that the first preset percentage threshold is not exceeded, a fourth data flow path in the optimized blood relation information set is obtained; Step S730: based on the graph-based traceability algorithm, a directed acyclic graph is constructed according to the fourth data flow path; Step S740: determining whether the edge weight value of the directed acyclic graph is greater than a preset path weight value; Step S750: if it is determined that the preset path weight value is exceeded, the path is determined as an abnormal flow path; Step S760: determining an abnormal training feature according to the abnormal flow path; Step S770: based on a feature reconstruction algorithm, the abnormal training feature is optimized to generate an optimized training feature set.
[0055] Specifically, the first preset percentage threshold (which can be 85%) is a benchmark for measuring model performance to determine whether the cross-institution prediction model needs to be optimized. If the prediction accuracy of the model is lower than the threshold, the reason needs to be traced back through the optimized blood relationship information set. The fourth data flow path can be converted into a directed acyclic graph based on a graph-based tracing algorithm (such as a path restoration based on depth-first search), where the nodes are business links, the directed edges represent the flow direction, and the edge weights represent the normal flow probability of the link (calculated based on historical data). For example, the fourth data flow path extracted from the blood relationship information set is "user registration → product browsing → ordering → payment → product delivery → confirming receipt". The above data flow path is constructed as a directed acyclic graph, where the nodes are user registration, product browsing, ordering, payment, product delivery, and confirming receipt, and the directed edges and weights are as follows: user registration → product browsing (weight 0.9, 90% of registered users will browse products), product browsing → ordering (weight 0.3, 30% of browsing users will place an order), ordering → payment (weight 0.95, 95% of users will pay), payment → logistics delivery (weight 0.98, 98% of products will be delivered), and logistics delivery → confirming receipt (weight 0.99, 99% of products will be received). The preset path weight value is 0.25, and it is found that the edge weight of one of the "product browsing → ordering" is 0.3 > 0.25 (normal flow path), but the edge weight of the other "product browsing → adding to shopping cart → ordering" is 0.15 < 0.25 (abnormal flow path). The abnormal flow path can include abnormal features such as shopping cart stay duration, product discount intensity, product out of stock, and product price increase. At this time, the "product discount intensity" can be quantified based on a feature reconstruction algorithm (such as 0-10%, 10%-30%, and above 30%), and a cross-feature "discount intensity x stay duration" can be constructed with "shopping cart stay duration" (less than 5 minutes, 5-15 minutes, and more than 15 minutes) to generate an optimized training feature set.
[0056] In the above method, by taking the cross-institution prediction model accuracy not reaching the first budget percentage threshold as the judgment point, and by using the graph-based tracing algorithm to cut into the fourth data flow path in the optimized blood relationship information set, the abnormal flow path and the corresponding abnormal training features can be accurately located, and the blindness of model optimization is avoided. This not only ensures the rationality of cross-institution data flow, but also effectively improves the prediction accuracy of the prediction model, and enhances the reliability of the prediction model in a multi-institution cooperation scenario.
[0057] In an embodiment of the present disclosure, after the step S770 of optimizing the abnormal training features based on the feature reconstruction algorithm to generate an optimized training feature set, the following steps are further included, as shown in Figure 8 The specific content is as follows: Step S810: Obtain the key features in the optimized training feature set to generate a simplified feature set. Step S820: training the reduced feature set based on the incremental learning algorithm to generate second initial model parameters; Step S830: updating the cross-institution prediction model according to the second initial model parameters; Step S840: determining whether the prediction accuracy of the updated cross-institution prediction model exceeds the second preset percentage threshold; Step S850: if it is determined that the second preset percentage threshold is not exceeded, dynamically allocating computing resources to continuously optimize the second initial model parameters.
[0058] Specifically, the optimized training feature set is extracted, and three key features, “discount intensity x stay time” (importance score 0.35), “logistics timeliness” (importance score 0.28), and “payment method” (importance score 0.22), are selected according to the size of their importance scores, with the larger the value, the more important. A reduced feature set is generated. Based on the incremental learning algorithm, the reduced feature set is trained through three rounds of local iteration to generate second initial model parameters, for example, w1=0.22 (corresponding to discount x stay), w2=0.18 (logistics timeliness), and w3=0.15 (payment method). The cross-institution prediction model is updated accordingly, and the prediction accuracy is 87%, which does not exceed the second preset percentage threshold of 90%. Therefore, the computing resources are dynamically allocated, the number of GPU cores is increased from 2 to 4, the number of iterations is expanded from 10 to 20, and the learning rate is fine-tuned from 0.01 to 0.005. The second initial model parameters are continuously optimized, and the prediction model accuracy is improved to 95%, which exceeds the second budget percentage threshold of 90%.
[0059] In the above method, the key features are extracted to generate a reduced feature set, which can reduce the interference of redundant information on the cross-institution prediction model, reduce the computational complexity and improve the training efficiency. Based on the incremental learning algorithm, the second initial model parameters are trained and generated, which can efficiently update the original model and avoid resource waste caused by repeated training. Then, the cross-institution prediction model is updated according to the updated model parameters, which can quickly iterate and optimize the performance of the cross-institution prediction model, effectively enhance the performance stability and practical value of the cross-institution prediction model, and ensure that it plays a more reliable prediction role in multiple institution cooperation scenarios.
[0060] Also provided in the embodiments of the present disclosure is a machine learning-based multi-source heterogeneous data analysis system, which can include a first acquisition module, a first determination module, a first generation module, a second generation module, a second determination module, a first judgment module, a second acquisition module, a third determination module, a repair module, a construction module, a second judgment module, and a third generation module. The first acquisition module is configured to acquire data flow logs in a multi-source heterogeneous database, the data flow logs including a timestamp, an IP address of a data source, an ID of a flow node, and a size of a data packet. The first determination module is configured to determine a blood relationship information set according to the data flow logs, the blood relationship information set including a data source, a first data flow path, and a flow node dependency relationship. The first generation module is configured to store the first data flow path and the flow node dependency relationship in a graph database to generate a graph database structure. The second generation module is configured to assign a corresponding unique identifier to each flow node in the graph database structure based on a node label assignment algorithm to generate a node identifier set. The second determination module is configured to determine a dynamic association relationship set between each flow node based on a graph traversal algorithm according to the node identifier set. The first judgment module is configured to determine whether there is an inconsistent dependency relationship in the dynamic association relationship set. The second acquisition module is configured to acquire an inconsistent dependency field if it is determined that there is an inconsistent dependency relationship. The third determination module is configured to determine an error field set according to the inconsistent dependency field, the error field including a missing value, an abnormal value, and a format inconsistency. The repair module is configured to repair the error field set based on a data cleaning algorithm to generate a repaired data set. The construction module is configured to construct a second data flow path according to the repaired data set. The second judgment module is configured to determine whether the second data flow path has path integrity based on a path traversal algorithm. The third generation module is configured to generate an optimized blood relationship information set according to the second data flow path if it is determined that the second data flow path has path integrity.
[0061] It should be noted that the embodiments of the machine learning-based multi-source heterogeneous data analysis system provided in the present application can be specifically used to execute the processing procedures of the embodiments of the machine learning-based multi-source heterogeneous data analysis method in the above embodiments, and the functions thereof will not be repeated here. Please refer to the detailed description of the above method embodiments.
[0062] As can be known from the above description, the multi-source heterogeneous data analysis system based on machine learning provided by the embodiments of the present disclosure can accurately extract information such as timestamps, IP addresses of data sources, IDs of flow nodes, and the like by analyzing and processing the data flow transfer logs in the multi-source heterogeneous database, and can store the first flow transfer path and the node dependency relationship in the blood relationship information set in the graph database, thereby realizing the visualization and traceability of the blood relationship information, and making the data sources and flow transfer context clear and transparent. With the help of the node label allocation algorithm and the graph traversal algorithm, the dynamic association relationship set between the flow nodes can be determined, the association relationship between the flow nodes can be dynamically captured, the inconsistent dependency relationship can be discovered in time, and the inconsistent dependency field can be located, thereby ensuring the accuracy of data association. The error field is determined through the inconsistent dependency field, and the error field is repaired through the data cleaning algorithm to generate a repaired data set, and a second data flow transfer path is constructed according to the repaired data set, and after the path integrity is verified, an optimized blood relationship information set is generated, thereby effectively improving the accuracy of the blood relationship information, making it possible to accurately capture the real-time changes of data flow transfer, and thereby improving the accuracy of credit approval or risk assessment. It provides a high-quality and reliable basis for subsequent cross-institutional data sharing, federal learning modeling, and the like, while enhancing the intelligentization and automation level of data governance, and helping to efficiently cope with the dynamic changes and complex association challenges in the multi-source heterogeneous data environment.
[0063] In the embodiments of the present disclosure, an electronic device is also provided, which includes one or more processors, and a memory resource represented by a memory for storing instructions executable by the processors, such as an application program. The application program stored in the memory can include one or more than one module each corresponding to a set of instructions. In addition, the processors are configured to execute the instructions to perform the above-mentioned multi-source heterogeneous data analysis method based on machine learning.
[0064] The electronic device can also include a power supply component configured to perform power management of the electronic device, a wired or wireless network interface configured to connect the electronic device to a network, and an input / output (I / O) interface. The electronic device can be operated based on an operating system stored in the memory, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0065] In one embodiment, a computer device is also provided, which can be a server. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is configured to be executed by the processor to implement a machine learning-based multi-source heterogeneous data analysis method.
[0066] In one embodiment, a computer device is also provided, which can be a server. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is configured to be executed by the processor to implement a machine learning-based multi-source heterogeneous data analysis method.
[0067] Also provided in the embodiments of the present disclosure is a non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the instructions enable the electronic device to perform a method for analyzing multi-source heterogeneous data based on machine learning, including: obtaining a data flow transfer log in a multi-source heterogeneous database, the data flow transfer log including a timestamp, an IP address of a data source, an ID of a transfer node, and a size of a data packet; determining a blood relationship information set according to the data flow transfer log, the blood relationship information set including a data source, a first data flow transfer path, and a transfer node dependency relationship; storing the first data flow transfer path and the transfer node dependency relationship by using a graph database to generate a graph database structure; assigning a corresponding unique identifier to each transfer node in the graph database structure based on a node label assignment algorithm to generate a node identifier set; determining a dynamic association relationship set between each transfer node based on the node identifier set according to a graph traversal algorithm; determining whether there is an inconsistent dependency relationship in the dynamic association relationship set; if it is determined that there is an inconsistent dependency relationship, obtaining an inconsistent dependency field; determining an error field set according to the inconsistent dependency field, the error field including a missing value, an abnormal value, and a format inconsistency; repairing the error field set based on a data cleaning algorithm to generate a repaired data set; constructing a second data flow transfer path according to the repaired data set; determining whether the second data flow transfer path has path integrity based on a path traversal algorithm; and if it is determined that the second data flow transfer path has path integrity, generating an optimized blood relationship information set according to the second data flow transfer path.
[0068] The present disclosure can take the form of a computer program product implemented on one or more storage media (including but not limited to magnetic storage media, CD-ROM, optical storage media, etc.) including program code. The computer-readable storage media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0069] It is to be understood that even though various steps of the machine learning based multi-source heterogeneous data analysis method of the present disclosure are described in a particular order in the drawings, this is not required or implied as to the order of execution of the steps, nor is it required that all of the steps be performed to achieve the desired results. Additional or alternative steps can be omitted, multiple steps can be combined into a single step, a single step can be broken into multiple steps, etc., all of which are considered to be within the scope of the present disclosure.
[0070] It is to be understood that the present disclosure does not limit its application to the details of the machine learning based multi-source heterogeneous data analysis system modules set forth in the specification. The present disclosure is capable of other embodiments and of being practiced or being carried out in various ways. Variations and modifications of the foregoing are within the scope of the present disclosure. It is to be understood that the present disclosure extends to all alternative combinations of two or more of the individual features mentioned or evident from the description and / or drawings. All these different combinations constitute various alternative aspects of the present disclosure. The embodiments of the present disclosure are illustrative of the best mode of putting the present disclosure into practice and are presented by way of example only. They are not limiting of the present disclosure.
Claims
1. A multi-source heterogeneous data analysis method based on machine learning, characterized in that: include: Obtain data flow logs from multi-source heterogeneous databases, including timestamps, IP addresses of data sources, IDs of flow nodes, and sizes of data packets; Determine a lineage information set based on the data flow log, wherein the lineage information set includes a data source, a first data flow path, and a flow node dependency relationship; Using a graph database, storing the first data flow path and the flow node dependency to generate a graph database structure; Based on a node label assignment algorithm, assigning a corresponding unique identifier to each of the flow nodes in the graph database structure to generate a node identifier set; Based on a graph traversal algorithm, determining a dynamic association relationship set between each of the flow nodes according to the node identification set; Determining whether there is an inconsistent dependency relationship in the dynamic association relationship set; If it is determined that the inconsistent dependency relationship exists, obtaining the inconsistent dependency field; Determine an error field set based on the inconsistent dependent fields, where the error fields include missing values, abnormal values, and inconsistent formats; Repairing the error field set based on a data cleaning algorithm to generate a repaired data set; Constructing a second data flow path according to the repair data set; Based on a path traversal algorithm, determining whether the second data flow path has path integrity; If it is determined that the path is complete, an optimized lineage information set is generated based on the second data flow path.
2. The multi-source heterogeneous data analysis method based on machine learning according to claim 1 is characterized in that: If it is determined that the inconsistent dependency relationship exists, obtaining the inconsistent dependency field includes: Obtain edge weight values between current flow nodes, where the edge weight values are used to represent the dependency strength between each flow node; Based on a depth-first search algorithm, traverse the graph database structure to obtain a plurality of initial edge weight values; Determine whether the difference between the current edge weight value and the initial edge weight value is greater than a first preset deviation threshold; If the difference is greater than the first preset deviation threshold, it is determined that an inconsistent dependency exists; Taking any of the flow nodes as a starting point, recording all third data flow paths corresponding thereto, and calculating a path weight value corresponding to each of the third data flow paths; Determining whether a difference between a path weight value corresponding to each of the third data flow paths and a path weight value corresponding to a historical data flow path is greater than a second preset deviation threshold; If it is determined that the deviation is greater than the second preset deviation threshold, it is marked as a deviation flow path; The inconsistent dependent fields are determined according to the deviation flow path.
3. The multi-source heterogeneous data analysis method based on machine learning according to claim 1 is characterized in that: The step of repairing the error field set based on the data cleaning algorithm to generate a repaired data set includes: performing a standardization process on the error field set to generate a first standardized data set; Calculating a covariance matrix of the first standardized data set to determine eigenvalues and eigenvectors of the covariance matrix; Sorting the eigenvalues in order from the largest value to the smallest value, and selecting the eigenvectors corresponding to the first five eigenvalues in the sorting results to generate a characteristic matrix; A corresponding repair solution is determined according to the type of the error field in the feature matrix.
4. The multi-source heterogeneous data analysis method based on machine learning according to claim 3 is characterized in that: The determining a corresponding repair solution according to the type of the error field in the feature matrix includes: When the feature matrix has the missing value, obtaining multiple neighbor values corresponding to the missing value based on the K-nearest neighbor interpolation algorithm; Performing a weighted average operation on the multiple neighbor values, and using the operation result as the missing value; When there are outliers in the feature matrix, calculating a standard score of the outlier; Determining whether the standard score of the abnormal value exceeds a preset standard score threshold; If it is determined that the score exceeds the preset standard threshold, the corresponding abnormal value is eliminated; When the feature matrix has inconsistent formats, regular expressions are used to adjust the inconsistent formats to a unified format.
5. The multi-source heterogeneous data analysis method based on machine learning according to claim 1 is characterized in that: After generating an optimized lineage information set according to the second data flow path if it is determined that the path is complete, the method further includes: performing standardization on the optimized blood relationship information set to generate a second standardized data set; Calculating the privacy budget of the second standardized data set based on a differential privacy algorithm; Determine whether the privacy budget exceeds the preset privacy budget threshold; If it is determined that the data does not exceed the preset privacy threshold, adding Gaussian noise to the second normalized data set to generate a perturbed data set; Based on a federated learning framework, local model parameters of each institution are fused in the perturbation dataset to generate a shared model parameter set; Determine whether the shared model parameter set has passed the consistency verification of each institution; If it is determined that the consistency verification of each institution is satisfied, a secure shared data set is generated.
6. The multi-source heterogeneous data analysis method based on machine learning according to claim 5, characterized in that: After generating a secure shared data set if it is determined that the consistency verification of each institution is satisfied, the method further includes: determining a training feature set based on the secure shared dataset; Determining first initial model parameters based on the training feature set based on a distributed gradient descent algorithm; sending the first initial model parameters to each organization; Based on the local data, each mechanism calculates the corresponding local gradient update value; Aggregating the local gradient update values to generate a global gradient update value; Updating the first initial model parameters according to the global gradient update value to generate final model parameters; A cross-institutional prediction model is generated based on the final model parameters.
7. The multi-source heterogeneous data analysis method based on machine learning according to claim 6, characterized in that: After generating the cross-institutional prediction model according to the final model parameters, the method further includes: Determining whether the prediction accuracy of the cross-institutional prediction model does not exceed a first preset percentage threshold; If it is determined that the first preset percentage threshold is not exceeded, obtaining a fourth data flow path in the optimized blood relationship information set; Based on the graph tracing algorithm, a directed acyclic graph is constructed according to the fourth data flow path; Determining whether the edge weight value of the directed acyclic graph is greater than a preset path weight value; If the value is greater than the preset path weight, it is determined to be an abnormal flow path; determining abnormal training features based on the abnormal flow path; Based on a feature reconstruction algorithm, the abnormal training features are optimized to generate an optimized training feature set.
8. The multi-source heterogeneous data analysis method based on machine learning according to claim 7 is characterized in that: After optimizing the abnormal training features based on the feature reconstruction algorithm to generate an optimized training feature set, the method further includes: Obtain key features from the optimized training feature set to generate a streamlined feature set; Training the reduced feature set based on an incremental learning algorithm to generate second initial model parameters; updating the cross-institutional prediction model according to the second initial model parameters; Determining whether the updated prediction accuracy of the cross-institutional prediction model does not exceed a second preset percentage threshold; If it is determined that the second preset percentage threshold is not exceeded, computing resources are dynamically allocated to continuously optimize the second initial model parameters.
9. A multi-source heterogeneous data analysis system based on machine learning, characterized in that: include: The first acquisition module is used to obtain data flow logs in multi-source heterogeneous databases, wherein the data flow logs include a timestamp, an IP address of a data source, an ID of a flow node, and a size of a data packet; A first determining module is configured to determine a lineage information set based on the data flow log, wherein the lineage information set includes a data source, a first data flow path, and a flow node dependency relationship; A first generating module is configured to use a graph database to store the first data flow path and the flow node dependency relationship to generate a graph database structure; A second generating module is configured to assign a corresponding unique identifier to each of the flow nodes in the graph database structure based on a node label assignment algorithm to generate a node identifier set; A second determining module is configured to determine a dynamic association relationship set between each of the flow nodes according to the node identification set based on a graph traversal algorithm; A first judgment module is used to judge whether there is an inconsistent dependency relationship in the dynamic association relationship set; A second acquisition module is configured to acquire an inconsistent dependency field if it is determined that the inconsistent dependency relationship exists; A third determining module is configured to determine an error field set based on the inconsistent dependent fields, where the error fields include missing values, abnormal values, and format inconsistencies; A repair module, configured to repair the error field set based on a data cleaning algorithm to generate a repair data set; A construction module, configured to construct a second data flow path according to the repair data set; A second judgment module is used to judge whether the second data flow path has path integrity based on a path traversal algorithm; The third generation module is used to generate an optimized blood relationship information set according to the second data flow path if it is determined that the path integrity is met.
10. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Blood relationship information visual representation method
CN113722310A
Data circulation and traceability evidence obtaining method and device, equipment and storage medium
CN118827190A
Informatization processing method based on big data
CN119003495A
Data lineage parsing method and system
WO2025060581A1
Cited By
Security risk prediction method and system for mobile application data in full life cycle
CN121765753A