Federal learning-based data element analysis system and method

By designing a federated learning data element analysis system that includes modules of data preparation, federated learning training, association analysis, trend capture and suggestions generation, the problem of existing systems being unable to distinguish contingents from trends is solved, and more accurate and interpretable decision recommendations are achieved.

CN120197658AInactive Publication Date: 2025-06-24BEIJING HAIZHIYAN ADVERTISEMENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510678929.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data element analysis systems based on federated learning cannot distinguish accidental associations in data elements, resulting in reduced decision accuracy and interpretability, while being unable to capture the trend evolution of data elements.

Method used

A data element analysis system based on federated learning is designed, including data preparation module, federated learning training module, association analysis module, trend capture module and suggestions generation module. Through these modules, the system can distinguish between accidental and inevitable relationships in data elements, and capture the trend changes of data elements in real time to generate decision-making suggestions.

Benefits of technology

Effectively distinguish between accidental relationships, avoid misleading model decisions, improve decision-making accuracy and interpretability, and capture trend changes in data elements in real time to improve analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197658A_ABST
    Figure CN120197658A_ABST
Patent Text Reader

Abstract

The invention discloses a data element analysis system and method based on federated learning. The system comprises a data preparation module, a federated learning training module, a correlation analysis module, a trend capture module and a suggestion generation module. The data preparation module is used for collecting data elements from all participants and preprocessing the collected data, and the collected data elements comprise numeric data and non-numeric data; and the federated learning training module is used for constructing and training a data element analysis model based on federated learning. The invention belongs to the technical field of data analysis, and aims to solve the problems that in the prior art, accidental association in data elements cannot be distinguished, and meanwhile, trend evolution of the data elements cannot be captured through analysis of related parameters of the data elements. The method has the technical effects that the accidental association in the data elements can be effectively distinguished, and meanwhile, the trend evolution of the data elements can be conveniently analyzed and captured according to related parameters of the data elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data analysis, and specifically relates to a data element analysis system and method based on federated learning. Background Art

[0002] With the increasing prominence of data privacy and security issues, as well as the growing demand for data sharing and collaboration among people. The core idea of federated learning is to achieve model training and knowledge sharing among multiple participants while protecting data privacy. Therefore, a data element analysis system based on federated learning can store data on each local device or data owner, without centralized data sharing, but through encrypted communication and distributed computing methods, jointly train models among multiple participants, and the model parameters are encrypted during transmission, greatly reducing the risk of data privacy being violated.

[0003] However, most current data element analysis systems based on federated learning mainly rely on static statistical data element-related parameters, unable to distinguish accidental associations in data elements, causing accidental associations to mislead the model's decision-making, making it make judgments based on incorrect information, thereby reducing the accuracy and interpretability of decision-making. At the same time, the analysis of relevant parameters of data elements cannot capture the trend evolution of data elements, easily reducing the effect of data element analysis. Summary of the Invention

[0004] This application provides a data element analysis system and method based on federated learning, aiming to solve the problems in the prior art that accidental associations in data elements cannot be distinguished, and at the same time, the analysis of relevant parameters of data elements cannot capture the trend evolution of data elements.

[0005] In a first aspect, a data element analysis system based on federated learning includes a data preparation module, a federated learning training module, an association analysis module, a trend capture module, and a recommendation generation module;

[0006] The data preparation module is used to collect data elements from each participant and preprocess the collected data. The collected data elements include numerical data and non-numerical data;

[0007] The federated learning training module is used to construct and train a data element analysis model based on federated learning, and output relevant parameters including time weights and feature importance according to the preprocessed data elements for use by the association analysis module;

[0008] The association analysis module is used to monitor and dynamically analyze the relevant parameters between data elements in real time, distinguish accidental associations and necessary associations in data elements. The association analysis module includes a division calculation unit, a correlation analysis unit, and an association determination unit;

[0009] The trend capture module is used to capture the inevitable correlation trend evolution and accidental correlation trend evolution of data elements, integrate them and predict the future change trend of data elements to provide information for decision-making;

[0010] The recommendation generation module is used to input the relevant parameters of the obtained data elements and the trend prediction results into the decision-making rules to generate decision-making recommendations.

[0011] Further, the data preparation module includes a collection unit and a preprocessing unit. The collection unit is used to establish a secure data connection with each participating party and extract the required data elements from each data source in accordance with the predetermined data collection rules from the databases of the participating parties. The preprocessing unit is used to perform standardization, encoding, and data partitioning processing operations on the collected data elements to obtain preprocessed data.

[0012] Further, the partitioning calculation unit is used to select appropriate coefficients according to the update frequency and content of the data elements to calculate the relevant parameters of the data elements. The correlation analysis unit is used to compare the relevant parameter r between data elements within different time windows, analyze the change trend of the relevant parameters, and draw a line chart of the change of the relevant parameter coefficients over time.

[0013] Further, the association determination unit is used to extract the data elements with accidental associations from the correlation analysis unit and label them.

[0014] Further, the trend capture module includes an inevitable correlation trend analysis unit, an accidental correlation trend analysis unit, and an integration unit. The inevitable correlation trend analysis unit can analyze the causal relationship between inevitably correlated data elements through the Granger causality test method and find out the main factors affecting the inevitable change trend of data elements.

[0015] Further, the accidental correlation trend analysis unit is used to pay attention to the sudden events and random factors related to the data elements and analyze the short-term impact of the factors on the data elements. The integration unit is constructed based on a neural network model and is used to synthesize the impacts of the inevitable correlation trend and the accidental correlation trend to generate trend features.

[0016] Further, the federated learning training module specifically includes model initialization, model training and parameter aggregation, model update, dynamic weight allocation mechanism, and feature importance evaluation.

[0017] In a second aspect, a data element analysis method based on federated learning specifically includes the following steps:

[0018] S1: Extract numerical and non-numerical data elements from each participating party through an encryption protocol and preprocess the data elements;

[0019] S2: Construct a data element analysis model based on a decision tree model and a federated learning architecture, and analyze the real-time element data of each participant through the data element analysis model;

[0020] S3: Divide the data elements into time windows by the hour, calculate the correlation parameter coefficients of the data elements within each window, and determine the necessary associations and accidental associations based on the correlation parameters of the data elements;

[0021] S4: Capture the evolution of the necessary association trend and the accidental association trend of the data elements according to the necessary associations and accidental associations, and generate a trend prediction result by synthesizing the influences of the necessary association trend and the accidental association trend;

[0022] S5: Input the relevant parameters and the trend prediction result into the decision rules, and output specific decision suggestions according to the decision rules.

[0023] Further, the specific working steps of S1 are as follows:

[0024] S1.1: Find the minimum and maximum values in each data source among the preprocessed data elements, and use the minimum and maximum values to apply the Min - Max normalization formula to transform each data point;

[0025] S1.2: Perform encoding processing on the transformed data elements;

[0026] S1.3: Re - sort and divide the encoded data according to a random sequence in proportion.

[0027] Further, the calculation formula of the correlation parameter r of the relevant parameter coefficients is specifically as follows:

[0028]

[0029] Where, and are the values of two data elements at the i - th sample point respectively, x and y are the means of the two data elements respectively, and n is the number of samples.

[0030] Compared with the prior art, the present application has at least the following beneficial effects:

[0031] Based on further analysis and research of the problems in the prior art, by adopting the framework of federated learning and combining with an association analysis module, the present application can distinguish accidental associations in data elements, avoid misleading model decisions due to accidental associations, thereby reducing the problems of decision accuracy and interpretability. At the same time, the trend capture module can capture the trend changes of data elements in real time, predict the future trends of data elements, and overall improve the analysis effect of data elements. Description of the Drawings

[0032] Figure 1 A schematic diagram of a module of a data element analysis system based on federated learning provided in one embodiment of the present application;

[0033] Figure 2 A flowchart of a data element analysis method based on federated learning provided for one embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0035] like Figure 1 As shown, the data element analysis system based on federated learning provided by the present application includes a data preparation module, a federated learning training module, a correlation analysis module, a trend capture module and a suggestion generation module;

[0036] The data preparation module is used to collect data elements from each participant and pre-process the collected data to ensure the quality and consistency of the data, providing a reliable data basis for subsequent analysis and modeling. The data preparation module includes a collection unit and a pre-processing unit. The collection unit is used to establish a secure data connection with each participant, using an encrypted transmission protocol to extract the required data elements from the participant's database according to the predetermined data collection rules and from each data source. The collected data elements include numerical data and non-numerical data.

[0037] The preprocessing unit is used to perform standardization, encoding and data division processing operations on the collected data elements to ensure the quality of the data elements and obtain preprocessed data.

[0038] Among them, standardization processing: find the minimum and maximum values ​​in each data source in the preprocessed data elements, use the minimum and maximum values, and apply the Min-Max standardization formula to each data point for conversion.

[0039] Coding processing: Assign a unique integer label to each data source in the preprocessed data element, establish a mapping dictionary between categories and labels, and for each data source, find the corresponding label in the mapping dictionary according to its category, and use the label as the encoding value of the data point.

[0040] Data partitioning: Use a random number generator to generate a random sequence of the same size as the preprocessed data, then reorder the encoded data according to the sequence, and divide the randomly shuffled data into training set, validation set, and test set according to the partition ratio. The ratio of training set, validation set, and test set is 7:1.5:1.5. The training set is used for model training, the validation set is used to adjust model parameters and evaluate model performance, and the test set is used for final model testing and evaluation.

[0041] The federated learning training module is used to construct and train a data element analysis model based on federated learning. Under the premise of protecting the data privacy of each participating party, it makes full use of the data resources of multiple participating parties to improve the performance and generalization ability of the model, enabling the model to more accurately analyze the relationships and laws among data elements. It dynamically adjusts the aggregation weights of model parameters according to the time stability and cross-party consistency of data elements. It outputs relevant parameters including time weights and feature importance based on the preprocessed data elements for use by the association analysis module. The specific content is as follows:

[0042] a) Model initialization

[0043] Construct an analysis model based on a decision tree machine learning model and a federated learning architecture and perform initialization. Initialization includes setting the initial parameters of the model, selecting an optimization algorithm, and a loss function entropy loss;

[0044] b) Model training and parameter aggregation

[0045] Send the initialized model parameters to each participating party. After receiving the model parameters, the participating party loads them into the local model. Each participating party uses the locally preprocessed data elements to train the local model. During the training process, the participating party calculates the gradient of the local model, that is, the partial derivative of the loss function with respect to the model parameters. Then, use the local data and the calculated gradient to update the parameters of the local model.

[0046] During the local model training process, the participating party does not share the original data, but only shares the update information of the model parameters. To further protect data privacy, the update information of the model parameters can be encrypted, such as using homomorphic encryption technology, so that the central node cannot obtain the original data of the participating party when aggregating parameter updates.

[0047] Each participating party uploads the encrypted model parameter update information to the central node. The central node decrypts the encrypted parameter update information using the corresponding decryption key. The federated average algorithm is used to aggregate the parameter updates of each participating party.

[0048] c) Model update

[0049] The central node uses the aggregated global model parameter updates to update the global model. Add the aggregated parameter updates to the current parameters of the global model to obtain the updated global model parameters.

[0050] The central node sends the updated global model parameters to each participating party for the next round of local model training. Repeat the above process of local model training, model parameter aggregation, and global model update until the global model converges or reaches the preset number of training rounds.

[0051] d) Dynamic weight allocation mechanism

[0052] Introduce a time decay factor, assign low weights to short-term accidental associations and high weights to long-term stable associations. Based on the multi-party voting mechanism, only aggregate the characteristic parameters verified as inevitable associations by the majority of parties to filter out noise associations. Thus, reduce the impact of accidental associations on model decision-making and improve the robustness of the model.

[0053] e) Feature importance evaluation

[0054] Use the Shapley value or attention weights to calculate the global importance of the features of each party and filter out low-importance features. Take the feature importance as an additional explanatory layer of the model output to support the traceability of decision-making suggestions. Facilitate enhancing the interpretability of model decision-making and reducing the misleading risk of accidental associations.

[0055] The association analysis module is used to monitor and dynamically analyze the correlation parameters between data elements in real time, distinguish accidental associations and inevitable associations in data elements, and provide a more accurate and interpretable basis for the model's decision-making. Through the changes in the correlation parameters of data elements, the changing trend of the relationship between data elements can be detected in a timely manner, helping decision-makers better understand the interaction between data elements. The association analysis module includes a partitioning calculation unit, a correlation analysis unit, and an association determination unit.

[0056] The partitioning calculation unit is used to select appropriate coefficients according to the update frequency and content of data elements to calculate the correlation parameters of data elements. The specific content is as follows:

[0057] a) Divide the data elements into multiple consecutive time windows in chronological order per hour, and each time window contains the data elements within the time range of each hour.

[0058] b) Calculate the correlation parameter r of the data elements within the time window per hour according to the correlation parameter calculation formula. The specific calculation formula of the correlation parameter r is as follows:

[0059]

[0060] where and are the values of two data elements at the i-th sample point respectively, x and y are the means of the two data elements respectively, and n is the number of samples.

[0061] The correlation analysis unit is used to compare the correlation parameters r between data elements in different time windows, analyze the changing trend of the correlation parameters, draw a line chart of the correlation parameter coefficients changing with time, and visually observe the changes in the correlation parameters.

[0062] If the relevant parameter r remains stable and at a relatively high level within several consecutive time windows, it indicates an inevitable correlation between data elements; if the relevant parameter r suddenly increases or decreases within a certain time window, it indicates an accidental correlation between data elements.

[0063] The correlation determination unit is used to extract the data elements with accidental correlations from the correlation analysis unit and label them.

[0064] The trend capture module is used to capture the evolution of the inevitable correlation trend and the accidental correlation trend of data elements, integrate them and predict the future change trend of data elements to provide information for decision-making. Among them, the inevitable correlation trend reflects the internal law and stable direction of data elements in the long-term development process; the accidental correlation trend reflects the short-term fluctuations of data elements affected by sudden events or random factors. By integrating and analyzing the evolution of the inevitable and accidental correlation trends, the future change trend of data elements is predicted to provide accurate and timely information for decision-makers. The trend capture module includes an inevitable correlation trend analysis unit, an accidental correlation trend analysis unit, and an integration unit.

[0065] The inevitable correlation trend analysis unit can analyze the causal relationship between inevitably correlated data elements through the Granger causality test method and find out the main factors affecting the inevitable change trend of data elements.

[0066] The accidental correlation trend analysis unit is used to pay attention to sudden events and random factors related to data elements and analyze the short-term impact of factors on data elements. By collecting data before and after the event and comparing and analyzing the changes in the data, the impact degree and duration of the event are evaluated.

[0067] At the same time, an anomaly detection algorithm is used to identify the abnormal fluctuation points in the data. The abnormal fluctuation points reflect the emergence of the accidental correlation trend.

[0068] The integration unit is constructed based on a neural network model and is used to synthesize the impacts of the inevitable correlation trend and the accidental correlation trend to generate trend features. The specific content is as follows:

[0069] a) Preset trend weights: Inevitable correlation trend weight: Determined according to the stability of the trend and the goodness of fit of historical data.

[0070] For example, if the inevitable relationship trend is verified by long-term data, the weight can be set to 0.7.

[0071] Accidental correlation trend weight: Determined according to the frequency of accidental events.

[0072] For example, if the accidental relationship trend is driven by short-term events, the weight can be set to 0.3.

[0073] b) Calculate the comprehensive trend: Calculate the comprehensive trend T based on the inevitable relationship trend weight and the accidental relationship trend weight. The calculation formula for the comprehensive trend T is as follows:

[0074]

[0075] Among them, is the volatility of the inevitable correlation trend value in the line chart of the relevant parameter coefficients of this time window changing with time, is the volatility of the accidental correlation trend value in the line chart of the relevant parameter coefficients of this time window changing with time, is the inevitable relationship trend weight, is the accidental relationship trend weight.

[0076] c) Trend prediction: Select a neural network model, use the integrated trend information as feature data, and the future value of the data element as target data. Divide the data, divide the data set into a training set, a validation set, and a test set, and use the training set to train the selected prediction model. During the training process, by adjusting the parameters of the model, make the loss function value of the model on the training set the smallest.

[0077] Use the trained model to predict the future data elements. Take the comprehensive trend T of the data elements as the input and input it into the model to obtain the trend prediction result of the data elements.

[0078] The recommendation generation module is used to input the relevant parameters and trend prediction results of the obtained data elements into the decision rule to generate decision recommendations. The decision recommendations can include specific investment amounts, investment terms, risk assessments, etc. And feedback the decision recommendations to the decision maker, which helps the decision maker make decisions based on the decision recommendations.

[0079] In the above data element analysis system and method based on federated learning, by adopting the framework of federated learning and combining the association analysis module, it is possible to distinguish accidental associations in data elements, avoid misleading model decisions due to accidental associations, and thus reduce the problems of decision accuracy and interpretability. At the same time, the trend capture module can capture the trend changes of data elements in real time, predict the future trends of data elements, and overall improve the analysis effect of data elements.

[0080] As Figure 2 shown, the data element analysis method based on federated learning provided by this application includes the following steps:

[0081] S1: Extract numerical and non-numerical data elements from each participant through an encryption protocol, and preprocess the data elements to ensure the quality and consistency of the data;

[0082] The specific working steps of S1 are as follows:

[0083] S1.1: Find the minimum and maximum values in each data source among the preprocessed data elements, and use the minimum and maximum values to apply the Min-Max normalization formula to each data point for transformation;

[0084] S1.2: Perform encoding processing on the transformed data elements;

[0085] S1.3: Reorder and divide the encoded data according to a random sequence in proportion, and divide it into a training set, a validation set, and a test set according to the ratio of 7:1.5:1.5.

[0086] S2: Construct a data element analysis model based on a decision tree model and a federated learning architecture, and analyze the real-time element data of each participant through the data element analysis model to facilitate the determination of the relationships and rules between data elements;

[0087] S3: Divide the data elements into time windows by hour, calculate the correlation parameter coefficients of the data elements within each window, and determine the necessary associations and accidental associations according to the correlation parameters of the data elements.

[0088] S4: Capture the evolution of the necessary association trend and the accidental association trend of the data elements according to the necessary associations and accidental associations, and generate a trend prediction result by synthesizing the influences of the necessary association trend and the accidental association trend.

[0089] S5: Input the relevant parameters and the trend prediction result into the decision rule, and output specific decision suggestions according to the decision rule.

[0090] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

Claims

1. A data element analysis system based on federated learning, characterized in that It includes a data preparation module, a federated learning training module, an association analysis module, a trend capture module, and a recommendation generation module; The data preparation module is used to collect data elements from each participant and preprocess the collected data. The collected data elements include numerical data and non-numerical data; The federated learning training module is used to construct and train a data element analysis model based on federated learning, and output relevant parameters including time weights and feature importance according to the preprocessed data elements for use by the association analysis module; The association analysis module is used to monitor and dynamically analyze the relevant parameters between data elements in real time, and distinguish the accidental associations and necessary associations in the data elements. The association analysis module includes a division calculation unit, a correlation analysis unit, and an association determination unit; The trend capture module is used to capture the evolution of the necessary association trend and the accidental association trend of data elements, and predict the future change trend of data elements after integration; The recommendation generation module is used to input the relevant parameters and trend prediction results of the obtained data elements into the decision rule to generate decision recommendations.

2. The data element analysis system based on federated learning according to claim 1, wherein The data preparation module includes a collection unit and a preprocessing unit. The collection unit is used to establish a secure data connection with each participant, and extract the required data elements from each data source in the participant's database according to the predetermined data collection rules. The preprocessing unit is used to perform standardization, encoding, and data division processing operations on the collected data elements to obtain preprocessed data.

3. The data element analysis system based on federated learning according to claim 1, characterized in that The division calculation unit is used to calculate the relevant parameters of data elements by selecting appropriate coefficients according to the update frequency and content of data elements. The correlation analysis unit is used to compare the relevant parameters r between data elements in different time windows, analyze the change trend of the relevant parameters, and draw a line chart of the change of the relevant parameter coefficients over time.

4. The data element analysis system based on federated learning according to claim 1, wherein The association determination unit is used to extract the data elements with accidental associations from the correlation analysis unit and label them.

5. The data element analysis system based on federated learning according to claim 1, wherein The trend capture module includes a necessary association trend analysis unit, an accidental association trend analysis unit, and an integration unit. The necessary association trend analysis unit can analyze the causal relationship between necessarily associated data elements through the Granger causality test method, and find out the main factors affecting the necessary change trend of data elements.

6. The data element analysis system based on federated learning according to claim 5, wherein The accidental association trend analysis unit is used to pay attention to the sudden events and random factors related to data elements, and analyze the short-term impact of the factors on data elements. The integration unit is constructed based on a neural network model and is used to synthesize the impacts of the necessary association trend and the accidental association trend to generate trend features.

7. The data element analysis system based on federated learning according to claim 1, characterized in that, The federated learning training module specifically includes model initialization, model training and parameter aggregation, model update, a dynamic weight allocation mechanism, and feature importance evaluation.

8. A data element analysis method based on federated learning, characterized in that Specifically, it includes the following steps: S1: Extract numerical and non-numerical data elements from each participant through an encryption protocol, and preprocess the data elements; S2: Construct a data element analysis model based on a decision tree model and a federated learning architecture, and analyze the real-time element data of each participant through the data element analysis model; S3: Divide the data elements into time windows by the hour, calculate the correlation parameter coefficients of the data elements within each window, and determine the necessary associations and accidental associations based on the correlation parameters of the data elements; S4: Capture the evolution of the necessary association trends and accidental association trends of the data elements according to the necessary associations and accidental associations, and generate a trend prediction result by synthesizing the influences of the necessary association trends and accidental association trends; S5: Input the correlation parameters and the trend prediction result into the decision rules, and output specific decision suggestions according to the decision rules.

9. The method for analyzing data elements based on federated learning according to claim 8, wherein, The specific working steps of S1 are as follows: S1.1: Find the minimum value and the maximum value in each data source among the preprocessed data elements, and use the minimum value and the maximum value to apply the Min-Max normalization formula to each data point for transformation; S1.2: Perform encoding processing on the transformed data elements; S1.3: Reorder and proportionally divide the encoded data according to the random sequence.

10. The method for analyzing data elements based on federated learning according to claim 8, wherein, The calculation formula of the correlation parameter r of the correlation parameter coefficients is specifically as follows: ; wherein, and are respectively the values of two data elements at the i-th sample point, x and y are respectively the means of the two data elements, and n is the number of samples.

Citation Information

Patent Citations

  • Transverse federal learning method and device, computer equipment and storage medium

    CN113515760A

  • Federal learning model training method and system, computer equipment and storage medium

    CN116796831A

  • Federal learning-based unmanned mine card health diagnosis analysis method and system

    CN118551258A

  • Federal learning model-based data processing method, apparatus and device, and medium

    CN118551413A

  • Tunnel ellipticity historical data collaborative analysis system and method based on federated learning

    CN119762025A