Asset management coding model modeling method based on federated learning

Through heterogeneous asset coding alignment and dynamic weight adjustment, the accuracy reduction problem caused by inconsistent coding in traditional asset management models is solved, and efficient and accurate asset management model construction is achieved to adapt to heterogeneous data environments.

CN120542266AActive Publication Date: 2025-08-26CHINA NAT INST OF STANDARDIZATION

Patent Information

Application Number
CN202510699289.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-26
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

When traditional asset management models face data silos, privacy and compliance requirements and asset data heterogeneity, they cause federal models to fail, and inconsistent coding reduces the global model accuracy.

Method used

Through heterogeneous asset encoding alignment, dynamic federal aggregation mechanism, small sample data source optimization and dynamic weight adjustment, the encoding output distribution alignment of each data source is achieved, model training and weights are dynamically adjusted, and model parameter aggregation is optimized.

Benefits of technology

Improve model accuracy, reduce invalid calculations, quickly respond to asset status changes, alleviate the long-tail distribution problem of data, and improve model accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542266A_ABST
    Figure CN120542266A_ABST
Patent Text Reader

Abstract

The invention relates to the field, and particularly discloses an asset management coding model modeling method based on federated learning, and the method comprises the steps of heterogeneous asset coding alignment, a dynamic federated aggregation mechanism, small sample data source optimization and dynamic weight adjustment. Through a dynamic federation aggregation mechanism, resources are efficiently utilized, the real-time response capability is improved, through comparison between a historical state and a current state, short-term noise fluctuation is filtered, a model is prevented from being misled by an abnormal value, and rapid modeling is achieved for a newly-accessed data source through data synthesis and model pre-training. Quality-driven aggregation is improved, a high-weight data source is dominant in parameter updating, and the accuracy of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of asset data security management, and in particular to a method for modeling an asset management coding model based on federated learning. Background Art

[0002] Currently, with the development of financial technology, the Internet of Things, and supply chain digitalization, asset management has gradually relied on data-driven modeling to improve asset valuation, risk prediction, and transaction efficiency. However, traditional asset management models face problems such as data silos, privacy and compliance requirements, and asset data heterogeneity. Therefore, federated learning has become a key technology to resolve the contradiction between data privacy and collaboration. Combined with asset coding models, it can realize asset data value mining in a distributed environment.

[0003] For example, the invention patent with publication number CN118709803A discloses a data asset security management method, apparatus, device, and medium based on federated learning, which relates to the field of data security management technology. The method includes obtaining raw data sent by a data source; storing the raw data in two computing nodes in a federated learning system; when one of the computing nodes in the federated learning system fails, determining first target data that is not backed up and stored in all valid computing nodes in the federated learning system; backing up the first target data and storing the backup data of the first target data in a target computing node in the federated learning system, where the target computing node is a valid computing node other than the valid computing node storing the first target data.

[0004] However, in the process of implementing the technical solutions of the embodiments of the present application, the present application discovered that the above technology has at least the following technical problems: In existing data asset management methods, the focus is directly on static data backup and computing node disaster recovery. The differences in asset data structures among different institutions will cause the federated model to fail. When directly aggregating raw data or model parameters, the global model accuracy will be reduced due to inconsistent encoding. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides an asset management coding model modeling method based on federated learning, which can effectively solve the problems involved in the above-mentioned background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: The present invention provides an asset management coding model modeling method based on federated learning, including: S1. Heterogeneous asset coding alignment: register the metadata of the asset field of each data source in the central server, insert an adaptive feature conversion layer before the local coding model, unify the heterogeneous inputs into a federated shared feature space, add feature distribution consistency constraints during aggregation, and force the coding output distribution of each data source to be aligned.

[0007] S2. Dynamic federated aggregation mechanism: Obtain the historical status data of each asset and the status data of each asset at the current time point, conduct a comprehensive analysis to obtain the status change index of each asset, and compare it with the asset status change index threshold preset in the modeling database to obtain the comparison result. Based on the comparison result, trigger the local model secondary training and upload the parameters.

[0008] S3. Small Sample Data Source Optimization: Compare the metadata of each data source with the sample size preset in the modeling database. Based on the comparison results, determine whether to generate synthetic asset data. Provide a pre-trained universal asset encoder through the central server and fine-tune the synthetic asset data to adapt to local data.

[0009] S4. Dynamic weight adjustment: Obtain the quality data of each data source, analyze and obtain the quality score index of each data source, conduct a comprehensive analysis of the quality score index of each data source and the status change index of each asset, and finally obtain the aggregation weight. Match the aggregation weight with the model parameter aggregation process in the optimized federated learning corresponding to each aggregation weight interval preset in the modeling database, and finally optimize the model parameter aggregation process.

[0010] As a further method, the encoding output distributions of the data sources are aligned, and the specific analysis process is as follows: An adaptive feature conversion layer is inserted before the local encoding model to linearly map each field to a uniform interval.

[0011] When aggregating at the central server, distribution alignment is enforced through the distribution alignment loss function, the federated aggregation process, and the dynamic alignment weight mechanism.

[0012] As a further method, the specific analysis process of the status change index of each asset is as follows: A comprehensive analysis is conducted on the historical average market value, historical average annualized rate of return, historical average volatility, historical average daily trading volume of each asset during the historical management cycle, and the average market value, annualized rate of return, volatility and average daily trading volume of each asset during the management cycle to obtain the status change index of each asset.

[0013] As a further method, the asset status change index threshold preset with the modeling database is compared to obtain a comparison result, and the local model secondary training and parameter upload are triggered according to the comparison result. The specific comparison process is as follows: If the state change index of an asset is less than or equal to the asset state change index threshold preset in the modeling database, the comparison result is recorded as the first comparison result.

[0014] If the state change index of an asset is greater than the asset state change index threshold preset in the modeling database, the comparison result is recorded as the second comparison result.

[0015] When the comparison result is the second comparison result, the local model is triggered to perform secondary training and upload parameters.

[0016] As a further method, the metadata of each data source is compared with the sample size preset in the modeling database, and whether to generate synthetic asset data is determined based on the comparison result. The specific comparison process is as follows: If the metadata of a data source is greater than or equal to the sample size preset in the modeling database, the comparison result is recorded as the first metadata comparison result.

[0017] When the comparison result is the metadata first comparison result, there is no need to generate synthetic asset data.

[0018] If the metadata of a data source is smaller than the sample size preset in the modeling database, the comparison result will be recorded as the second metadata comparison result.

[0019] When the comparison result is the metadata second comparison result, synthetic asset data needs to be generated.

[0020] As a further method, the quality score index of each data source is analyzed in the following specific process: The quality data of each data source specifically includes the field missing rate, time coverage, outlier ratio and external verification accuracy of each data source.

[0021] The field missing rate, time coverage, outlier ratio and external verification accuracy of each data source are comprehensively analyzed to obtain the quality score index of each data source.

[0022] As a further method, the quality score index of each data source is analyzed in the following specific process: The quality data of each data source specifically includes the field missing rate, time coverage, outlier ratio and external verification accuracy of each data source.

[0023] The field missing rate, time coverage, outlier ratio and external verification accuracy of each data source are comprehensively analyzed to obtain the quality score index of each data source.

[0024] As a further method, the specific analysis process of the aggregation weight is as follows: Comprehensively analyze the quality score index of each data source and the status change index of each asset to obtain the aggregate weight. The specific analysis method is as follows: ; Where, is the aggregation weight, is the state change index of the a-th asset, a is the number of each asset, , b is the total assets, The weight factor corresponding to the unit value of the asset status change index preset in the modeling database, is the quality score index of the mth data source, m is the number of each data source, , n is the total number of data sources, The weight factor corresponding to the quality score index unit value of the data source preset for the modeling database.

[0025] As a further method, the model parameter aggregation process is optimized, and the specific optimization process is: Compare the aggregate weight with the aggregate weight interval values ​​preset in the modeling database to determine the specific interval corresponding to the aggregate weight and obtain the interval where the aggregate weight is located.

[0026] According to the interval of the aggregation weight, the model parameter aggregation process in the federated learning is optimized, and finally the model parameter aggregation process is optimized.

[0027] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) The present invention solves the problem of incompatibility of data formats of different data source systems by forcing the distribution of the encoded output of each data source to be aligned, and mapping the heterogeneous asset fields of different data sources (such as text, numerical values, and images) to a unified feature space. Feature distribution constraints (such as MMD distance and KL divergence) ensure that the data of each participant is distributed and aligned in the federated shared space, avoiding the model bias towards data sources with large data volumes.

[0028] (2) The present invention uses a dynamic federated aggregation mechanism to achieve efficient resource utilization. Secondary training is triggered only when the asset status change index exceeds a threshold, reducing invalid calculations and quickly responding to status changes of highly volatile assets. The model prediction delay is shortened from minutes to seconds. By comparing historical status with current status, short-term noise fluctuations are filtered out to prevent the model from being misled by outliers.

[0029] (3) The present invention can solve the cold start problem by optimizing small sample data sources. For newly connected data sources or niche asset categories, rapid modeling can be achieved through synthetic data and pre-trained models. Synthetic data supplements small sample categories and alleviates the problem of long-tail data distribution. The universal asset encoder provided by the central server (such as a BERT-based text encoder) can quickly adapt to local data through fine-tuning, reducing local training costs.

[0030] (4) The present invention improves quality-driven aggregation through dynamic weight adjustment and aggregation optimization. High-weight data sources (such as those with high quality scores and large state changes) dominate parameter updates, thereby improving model accuracy. Through weight interval management, the impact of low-quality or malicious data sources is automatically reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.

[0032] Figure 1 Schematic diagram of the method steps of the present invention. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0034] Reference Figure 1 As shown, the present invention provides a modeling method for an asset management coding model based on federated learning, including: S1. Heterogeneous asset coding alignment: registering the metadata of the asset field of each data source in the central server, inserting an adaptive feature conversion layer before the local coding model, unifying the heterogeneous input into a federated shared feature space, adding feature distribution consistency constraints during aggregation, and forcing the coding output distribution of each data source to be aligned.

[0035] In this embodiment, the central server refers to the coordination node in the federated learning system, responsible for global model management, parameter aggregation, policy control, and inter-node communication. It stores asset field metadata (such as field name, data type, and business meaning) registered by each data source, serves as a mapping benchmark for heterogeneous data alignment, collects model updates (such as gradients and parameter increments) uploaded by each node, executes a federated aggregation algorithm (such as FedAvg) to generate a global model, and distributes it to each node. During the aggregation process, feature distribution consistency constraints are imposed to ensure that the encoding output distributions of different nodes are aligned. The local encoding model is a machine learning model deployed on each data source node (such as a bank or securities firm). It is responsible for encoding local heterogeneous asset data into a unified vector in the federated shared feature space, without any data leaving the local node. The federated shared feature space is a unified feature representation space that maps asset data from different data sources into feature vectors with consistent dimensions, semantics, and distribution through heterogeneous asset encoding alignment, supporting cross-node model collaborative training and joint analysis. Through this architecture, the federated learning system can efficiently integrate multi-source heterogeneous asset data while protecting data privacy, and build a unified asset management encoding model.

[0036] S2. Dynamic federated aggregation mechanism: Obtain the historical status data of each asset and the status data of each asset at the current time point, conduct a comprehensive analysis to obtain the status change index of each asset, and compare it with the asset status change index threshold preset in the modeling database to obtain the comparison result. Based on the comparison result, trigger the local model secondary training and upload the parameters.

[0037] S3. Small Sample Data Source Optimization: Compare the metadata of each data source with the sample size preset in the modeling database. Based on the comparison results, determine whether to generate synthetic asset data. Provide a pre-trained universal asset encoder through the central server and fine-tune the synthetic asset data to adapt to local data.

[0038] S4. Dynamic weight adjustment: Obtain the quality data of each data source, analyze and obtain the quality score index of each data source, conduct a comprehensive analysis of the quality score index of each data source and the status change index of each asset, and finally obtain the aggregate weight.

[0039] Furthermore, the metadata of the asset fields registered by each data source in the central server specifically include the name, type, value range, physical unit, data distribution and missing rate of the asset fields registered by each data source in the central server.

[0040] It should be explained that the names, types, value ranges, physical units, data distribution, and missing rates of the asset fields registered by the above-mentioned data sources on the central server are used to support subsequent feature alignment and federated training. Metadata must be defined through standardized protocols (such as Protobuf Schema), and the central server builds a global metadata knowledge graph to achieve automatic matching.

[0041] Specifically, the encoding output distribution of each data source is aligned, and the specific analysis process is as follows: An adaptive feature conversion layer is inserted before the local encoding model to linearly map each field to a uniform interval.

[0042] It should be explained that the Adaptive Feature Transformer (AFT) layer is inserted before the local encoding model (such as a neural network). Its core components and process are as follows: class AdaptiveFeatureTransformer(nn.Module): def__init__(self,metadata): super().__init__() #Dynamic initialization based on metadata self.scaler = DynamicScaler(metadata["value_range"]) # Range scaling self.embedder = CategoryEmbedder(metadata["categories"]) #Category field embedding self.missing_handler=LearnableImputer(metadata["missing_rate"]) # Missing value learning filling def forward(self, x): x = self.missing_handler(x) x = self.scaler(x) # Normalize to the federated shared space x = self.embedder(x) # discrete value continuous return x In this embodiment, each field is linearly mapped to a uniform interval (eg, [0, 1]) according to the registered value_range.

[0043] In a specific embodiment, for example, for bank A's "PROFIT [0, 10 million] → [0, 1]" and brokerage firm B's "PROFIT [0, 50 million] → [0, 1]", categorical fields (such as industry classification) are converted into dense vectors through a trainable embedding layer, and missing values ​​are predicted based on the attention mechanism rather than simply filling in zero or the mean.

[0044] When aggregating at the central server, distribution alignment is enforced through the distribution alignment loss function, the federated aggregation process, and the dynamic alignment weight mechanism.

[0045] It should be explained that the above-mentioned distribution alignment loss function refers to a global objective function that forces the distribution of encoded outputs (such as asset feature vectors) from different data sources to be consistent in the federated shared feature space, thereby eliminating model bias caused by differences in data distribution. In this embodiment, the types of loss functions used include but are not limited to Maximum Mean Discrepancy (MMD), which measures the difference between two distributions through a kernel function. Minimizing MMD can make the two distributions indistinguishable in RKHS, and KL Divergence (Kullback-Leibler Divergence), which is used to measure the difference between two probability distributions and is typically used to constrain the encoded output to follow a standard normal distribution (such as the latent variable constraint in VAE). The federated aggregation process is that the central server initializes the global model parameters and distributes them to each node. Each node uses local data to train the model and calculate the gradient. The central server collects the gradient and averages it weighted by the amount of data. The dynamic alignment weighting mechanism dynamically adjusts the weight of each node during aggregation based on the degree of difference between the data distribution of each node and the global distribution, thereby reducing the impact of nodes with large distribution differences on the global model.

[0046] Furthermore, the historical status data of each asset specifically includes the historical average market value, historical average annualized rate of return, historical average volatility and historical average daily trading volume of each asset during the historical management cycle.

[0047] It should be explained that the above historical status data reflects the long-term performance of an asset in a fixed period of time in the past and is used to analyze trend changes and historical patterns. The historical management cycle refers to a fixed period of time in the past, usually a complete natural cycle (such as 1 calendar year). The long-term statistical characteristics of the current asset are randomly set according to the data of each asset in this embodiment; the above historical average market value refers to the average market value of the asset on each trading day during the historical management cycle. It is obtained by collecting the closing market value of each day in the historical management cycle (market value = asset price × issuance quantity), summing the market value of all trading days in the cycle, and dividing it by the total number of trading days to obtain the historical average market value; the historical average annualized rate of return refers to the conversion of the asset rate of return in the historical management cycle into an annualized level, reflecting the long-term profitability, which is obtained by calculating the total rate of return of the asset in the cycle ( ), annualize the total return, If multiple historical periods are involved (such as the annualized rate of return for the past three years), the average value is taken. The historical average volatility refers to the degree of fluctuation of asset prices within the historical management period. It is usually measured by the annualized standard deviation, which is calculated by calculating the standard deviation of daily returns within the period and annualizing the daily standard deviation. , and get the annualized volatility. Similarly, if there are multiple historical period data, take the average of the annualized volatility of each period; the historical average daily trading volume refers to the average of daily trading volume in the historical management period, reflecting asset liquidity. It is obtained by collecting the trading volume of each trading day in the period (such as the number of stock traded shares, the amount of bond traded, etc.), summing it and dividing it by the total number of trading days.

[0048] Furthermore, the specific analysis process of the status change index of each asset is as follows: The historical average market value, historical average annualized rate of return, historical average volatility, historical average daily trading volume of each asset during the historical management cycle are comprehensively analyzed to obtain the status change index of each asset. The specific analysis method is as follows: , Where, is the state change index of the a-th asset, a is the number of each asset, , b is the total assets, is the historical average market value of the a-th asset in the historical management cycle, is the average market value of the a-th asset during the management cycle, The market value deviation reference value preset for the modeling database, The impact factor corresponding to the market value unit value preset in the modeling database, is the historical average annualized rate of return of the a-th asset during the historical management period, is the annualized rate of return of the a-th asset during the management period, The annualized return deviation reference rate preset for the modeling database, The impact factor corresponding to the annualized rate of return unit value preset for the modeling database, is the historical average volatility of the a-th asset during the historical management period, is the volatility of the a-th asset during the management period, The volatility deviation reference value preset for the modeling database, The impact factor corresponding to the volatility unit value preset in the modeling database, is the historical average daily trading volume of the a-th asset during the historical management period, is the average daily trading volume of the a-th asset during the management period, The average daily trading volume deviation value preset for the modeling database, The impact factor corresponding to the daily average trading volume unit value preset in the modeling database.

[0049] In this embodiment, the state change index of each asset is used to comprehensively reflect the degree of deviation of a certain asset from its long-term historical state at the current moment, and is condensed into a single scalar value by mathematical modeling. The impact factor corresponding to the market value unit value, the impact factor corresponding to the annualized rate of return unit value, the impact factor corresponding to the volatility unit value, and the impact factor corresponding to the average daily trading volume unit value preset in the modeling database respectively represent the degree of influence of the market value unit value, the annualized rate of return unit value, the volatility unit value, and the average daily trading volume unit value on the asset's state change index, where , , , ,and .

[0050] In this embodiment, if the historical average market value of each asset during the historical management cycle deviates significantly from the average market value of each asset during the management cycle, or even exceeds the preset market value deviation reference value, the abnormal growth of market value may reflect a bubble, and a sudden drop may indicate a financial crisis. With the increase in volatility, the asset's state change index increases significantly; if the historical average annualized rate of return of each asset during the historical management cycle deviates significantly from the annualized rate of return of each asset during the management cycle, or even exceeds the preset annualized rate of return deviation reference rate, the abnormal surge in rate of return may be short-term manipulation, and the plunge may reflect the deterioration of fundamentals; each asset during the historical management cycle When the historical average volatility of each asset deviates significantly from the volatility of each asset during the management cycle, or even exceeds the preset volatility deviation reference value, a sudden increase in volatility marks a risk event, and a sudden drop may indicate liquidity depletion; when the historical average daily trading volume of each asset during the historical management cycle deviates significantly from the average daily trading volume of each asset during the management cycle, or even exceeds the preset average daily trading volume deviation value, an abnormal increase in trading volume may be a back-to-back transaction, and a decrease may indicate liquidity risk; therefore, through a detailed analysis of each parameter in the asset's state change index, accurate risk identification can be achieved and a basis for subsequent dynamic federated learning optimization can be provided.

[0051] Specifically, the asset status change index threshold preset in the modeling database is compared to obtain a comparison result, and the local model secondary training and parameter upload are triggered according to the comparison result. The specific comparison process is as follows: If the state change index of an asset is less than or equal to the asset state change index threshold value preset in the modeling database, the comparison result is recorded as the first comparison result.

[0052] It should be explained that the above-mentioned preset asset status change index threshold is a pre-set critical value used to determine whether the asset status has changed significantly, and can be obtained using the industry risk benchmark issued by the regulatory agency or industry association; in this embodiment, when the status change index of an asset is less than or equal to the asset status change index threshold preset in the modeling database, it means that the asset status is within a normal fluctuation range, no significant risks or opportunities are detected, the parameter distribution is consistent with the historical pattern (such as market value fluctuations within ±10%), the prediction error of the federated model has not increased significantly, and the current model is maintained, and the next round of training cycle is extended.

[0053] If the state change index of an asset is greater than the asset state change index threshold preset in the modeling database, the comparison result is recorded as the second comparison result.

[0054] It should be explained that when the state change index of an asset exceeds the asset state change index threshold preset in the modeling database, the asset state has changed abnormally, which may be caused by financial crisis, liquidity depletion, market manipulation, etc. When the parameter distribution deviates from the historical range (for example, the volatility changes from 15% to 40%), the local model will be immediately trained again, and an early warning will be sent to the risk control system.

[0055] When the comparison result is the second comparison result, the local model is triggered to perform secondary training and upload parameters.

[0056] In this embodiment, the local model secondary training operation steps are: def local_retrain(local_data, global_model): 1. Data Preparation train_data = preprocess(local_data) # Contains the latest asset status data #2. Model initialization local_model = clone(global_model) # inherit global model parameters #3. Differentiated Training if > 0.5: # Extremely abnormal train_params = {"lr": 0.01, "epochs": 20} # High learning rate for fast adaptation else: #General exception train_params={"lr": 0.001, "epochs": 10} # 4. Training Execution local_model.fit(train_data,**train_params) # 5. Gradient clipping (anti-overfitting) torch.nn.utils.clip_grad_norm_(local_model.parameters(),5.0) return local_model In this embodiment, the uploaded parameters include but are not limited to parameter type model gradients (gradient tensors of each layer of the neural network), feature distribution statistics (mean / variance of encoder output), and training metadata (hyperparameters such as learning rate, batch size, and training time).

[0057] It should be explained that the specific process of uploading parameters is: parameter encryption, metadata packaging, secure transmission, uploading to the central server through the gRPC channel, and using TLS1.3 to encrypt the transmission link.

[0058] Furthermore, the metadata of each data source is compared with the sample size preset in the modeling database, and whether to generate synthetic asset data is determined based on the comparison result. The specific comparison process is as follows: If the metadata of a data source is greater than or equal to the sample size preset in the modeling database, the comparison result is recorded as the first metadata comparison result.

[0059] When the comparison result is the metadata first comparison result, there is no need to generate synthetic asset data.

[0060] It should be explained that the sample size in the data source metadata refers to the number of valid asset data records currently available in the data source. When the metadata of a data source is greater than or equal to the sample size preset in the modeling database, it means that the data source has sufficient data to support local model training and no additional data enhancement is required.

[0061] If the metadata of a data source is smaller than the sample size preset in the modeling database, the comparison result will be recorded as the second metadata comparison result.

[0062] When the comparison result is the metadata second comparison result, synthetic asset data needs to be generated.

[0063] It should be explained that when the metadata of a data source is smaller than the sample size preset in the modeling database, it means that the data source has a small sample problem, which will lead to underfitting of the model, high training error and high test error, and the federation aggregation will be dominated by the big data participants, so it is necessary to generate synthetic asset data.

[0064] In this embodiment, by generating synthetic asset data, it is possible to balance local data distribution, improve feature diversity, and meet the minimum data requirements of federated learning.

[0065] Specifically, the quality score index of each data source is analyzed in the following steps: The quality data of each data source includes the field missing rate, time coverage, outlier ratio and external verification accuracy of each data source; It should be explained that the above-mentioned field missing rate refers to the missing ratio of asset fields in the data source. It is obtained by counting the non-empty values ​​of the data source fields to obtain the number of non-empty fields, subtracting the number of non-empty fields from the total number of fields to obtain the number of empty fields, and dividing it by the total number of fields to obtain the field missing rate; time coverage refers to the completeness of the asset status data in the time dimension in the data source, reflecting the continuity of the data in the time series. The time field in the data can be extracted through timestamp analysis, the date distribution can be statistically analyzed, and the number of days to be collected can be compared with the data dictionary. The outlier ratio refers to the sample ratio of asset data in the data source that deviates from the normal range, reflecting the data Reliability can be assessed by using statistical methods (such as Z-score and interquartile range (IQR)) to detect outliers in numeric fields (such as market capitalization and rate of return). The number of abnormal samples and the total number of samples are counted separately, and the outlier ratio is obtained by dividing the number of abnormal samples by the total number of samples. The external verification accuracy refers to the accuracy ratio of asset data in the data source verified by external authoritative channels, reflecting the authenticity of the data. This can be achieved by randomly sampling a certain number of asset samples from the data source, comparing the sample data with authoritative external data sources (such as financial data platforms and government public data), counting the number of consistent samples, and dividing the result by the asset sample.

[0066] A comprehensive analysis of the field missing rate, time coverage, outlier ratio, and external verification accuracy of each data source was performed to obtain the quality score index of each data source. The specific analysis method is as follows: , Where, is the quality score index of the mth data source, m is the number of each data source, m=1,2,3,...,n, n is the total number of data sources, e is a natural constant, is the field missing rate of the mth data source, The impact factor corresponding to the unit value of the field missing rate preset for the modeling database, is the outlier ratio of the mth data source, The impact factor corresponding to the abnormal value ratio unit value preset for the modeling database, is the time coverage of the mth data source, Preset temporal reference coverage for the modeling database, The impact factor corresponding to the time coverage unit value preset for the modeling database, is the external verification accuracy of the mth data source, The impact factor corresponding to the external validation accuracy unit value preset for the modeling database.

[0067] It should be explained that the time reference coverage preset in the above-mentioned modeling database refers to the standard value preset in the federated learning modeling process, which is used to measure the integrity and continuity of asset data in the time dimension. The impact factor corresponding to the unit value of the field missing rate, the impact factor corresponding to the unit value of the outlier ratio, the impact factor corresponding to the unit value of the time coverage, and the impact factor corresponding to the unit value of the external verification accuracy preset in the modeling database respectively represent the degree of influence of the unit value of the field missing rate, the unit value of the outlier ratio, and the unit value of the external verification accuracy on the quality score index of the data source, where , , , ,and .

[0068] In this embodiment, a higher field missing rate means lower data integrity of key attributes (such as asset market value, yield, etc.) in the data source. A larger outlier ratio directly reflects data collection errors (such as sensor failure) or business anomalies (such as data transmission errors) in the data source, resulting in a decrease in the quality score index of the data source. A larger or smaller time coverage, that is, when the degree of deviation from the preset time reference coverage is large, a smaller time coverage means that there is a gap in the data time series, which cannot capture the long-term trend or cyclical characteristics of the asset, resulting in insufficient sensitivity of the model to dynamic changes. Excessive time coverage may introduce outdated information, increase noise interference, and reduce the timeliness of the model. A lower external verification accuracy indicates that the data credibility of the data source is poor, which directly affects the authenticity of the data and causes serious deviations between the prediction results and the actual business. Therefore, by analyzing each parameter in the quality score index of the data source and characterizing the quality of the data source through different dimensions, not only can the reliability of the federated learning model be guaranteed, but also the data source provider can be guided to actively improve data quality through the dynamic weight mechanism, forming a positive cycle of the data ecosystem.

[0069] Specifically, the aggregation weight and the specific analysis process are as follows: Comprehensively analyze the quality score index of each data source and the status change index of each asset to obtain the aggregate weight. The specific analysis method is as follows: ; Where, is the aggregation weight, is the state change index of the a-th asset, a is the number of each asset, , b is the total assets, The weight factor corresponding to the unit value of the asset status change index preset in the modeling database, is the quality score index of the mth data source, m is the number of each data source, , n is the total number of data sources, The weight factor corresponding to the quality score index unit value of the data source preset for the modeling database.

[0070] It should be explained that the weight factor corresponding to the unit value of the asset's state change index and the weight factor corresponding to the unit value of the data source's quality score index preset in the above-mentioned modeling database respectively represent the degree of influence of the unit value of the asset's state change index and the unit value of the data source's quality score index on the aggregation weight. In this embodiment, , ,and .

[0071] In this embodiment, a large state change index reflects drastic fluctuations in asset status (such as a stock price plunge or a credit rating downgrade), which will quickly transmit local risks to the global model. Abnormal asset status may be accompanied by data quality issues (such as a surge in missing values ​​caused by trading interruptions). Data sources with low quality scores (such as incomplete data or a large amount of noise) may contain erroneous information. If given a high weight, this noise will be introduced into the global model. Therefore, through detailed analysis of each parameter in the aggregate weight, the accuracy, robustness, efficiency and privacy protection of federated learning in asset management scenarios can be significantly improved.

[0072] Specifically, the model parameter aggregation process is optimized, and the specific optimization process is: Compare the aggregate weight with the aggregate weight interval values ​​preset in the modeling database to determine the specific interval corresponding to the aggregate weight and obtain the interval where the aggregate weight is located.

[0073] In this embodiment, the modeling database divides the aggregation weight interval into different intervals such as (0-0.5], (0.5, 1], etc. Each interval corresponds to a different optimization of the model parameter aggregation process in federated learning, and ultimately optimizes the model parameter aggregation process.

[0074] According to the interval of the aggregation weight, the model parameter aggregation process in the federated learning is optimized, and finally the model parameter aggregation process is optimized.

[0075] In this embodiment, the aggregation weight interval (0-0.5] belongs to the low-weight interval. The corresponding optimization method is to control the impact of noise and reduce resource consumption. Federated learning only participates in parameter updates once every 5-10 rounds to reduce the interference of low-quality data. The interval (0.5, 1] ​​belongs to the high-weight interval. The corresponding optimization method is to enhance effective information transmission, accelerate convergence, ensure the immediate impact of high-quality data, accelerate model convergence, upload complete model parameters instead of gradients, and retain more information.

[0076] The above content is merely an example and explanation of the structure of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined in this specification, they should all fall within the scope of protection of the present invention.

Claims

1. A method for modeling an asset management coding model based on federated learning, characterized in that: include: S1. Heterogeneous Asset Coding Alignment: Each data source registers the metadata of the asset field in the central server, inserts an adaptive feature conversion layer before the local encoding model, unifies the heterogeneous inputs into a federated shared feature space, and adds feature distribution consistency constraints during aggregation to force the alignment of the encoding output distributions of each data source. S2. Dynamic federated aggregation mechanism: This mechanism obtains historical and current state data for each asset, performs comprehensive analysis to determine each asset's state change index, and compares this index with the asset state change index threshold preset in the modeling database. The comparison result triggers secondary training of the local model and uploads the parameters based on the comparison result. S3. Small Sample Data Source Optimization: Compare the metadata of each data source with the sample size preset in the modeling database. Based on the comparison results, determine whether to generate synthetic asset data. A pre-trained universal asset encoder is provided by the central server, and the synthetic asset data is fine-tuned to adapt to local data. S4. Dynamic weight adjustment: Obtain the quality data of each data source, analyze and obtain the quality score index of each data source, conduct a comprehensive analysis of the quality score index of each data source and the status change index of each asset, and finally obtain the aggregation weight. Match the aggregation weight with the model parameter aggregation process in the optimized federated learning corresponding to each aggregation weight interval preset in the modeling database, and finally optimize the model parameter aggregation process.

2. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The metadata of the asset fields registered by each data source in the central server specifically includes the name, type, value range, physical unit, data distribution and missing rate of the asset fields registered by each data source in the central server.

3. The method for modeling an asset management coding model based on federated learning according to claim 2, characterized in that: The encoding output distribution of each data source is aligned, and the specific analysis process is as follows: Insert an adaptive feature conversion layer before the local encoding model to linearly map each field to a uniform interval; When aggregating at the central server, distribution alignment is enforced through the distribution alignment loss function, the federated aggregation process, and the dynamic alignment weight mechanism.

4. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The historical status data of each asset, specifically including the historical average market value, historical average annualized rate of return, historical average volatility, and historical average daily trading volume of each asset during its historical management cycle; The current status data of each asset at a specific time point includes the average market value, annualized rate of return, volatility, and average daily trading volume of each asset during the management cycle.

5. The method for modeling an asset management coding model based on federated learning according to claim 4, characterized in that: The specific analysis process of the status change index of each asset is as follows: A comprehensive analysis is conducted on the historical average market value, historical average annualized rate of return, historical average volatility, historical average daily trading volume of each asset during the historical management cycle, and the average market value, annualized rate of return, volatility and average daily trading volume of each asset during the management cycle to obtain the status change index of each asset.

6. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The asset status change index threshold preset in the modeling database is compared to obtain a comparison result, and the local model secondary training and parameter upload are triggered according to the comparison result. The specific comparison process is as follows: If the state change index of an asset is less than or equal to the asset state change index threshold preset in the modeling database, the comparison result is recorded as the first comparison result; If the state change index of an asset is greater than the asset state change index threshold preset in the modeling database, the comparison result is recorded as the second comparison result; When the comparison result is the second comparison result, the local model is triggered to perform secondary training and upload parameters.

7. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The metadata of each data source is compared with the sample size preset in the modeling database, and whether to generate synthetic asset data is determined based on the comparison result. The specific comparison process is as follows: If the metadata of a data source is greater than or equal to the sample size preset in the modeling database, the comparison result is recorded as the first metadata comparison result; When the comparison result is the metadata first comparison result, there is no need to generate synthetic asset data; If the metadata of a data source is smaller than the sample size preset in the modeling database, the comparison result is recorded as the second comparison result of the metadata; When the comparison result is the metadata second comparison result, synthetic asset data needs to be generated.

8. The method for modeling an asset management coding model based on federated learning according to claim 7, characterized in that: The specific analysis process of the quality scoring index of each data source is as follows: The quality data of each data source includes the field missing rate, time coverage, outlier ratio and external verification accuracy of each data source; The field missing rate, time coverage, outlier ratio and external verification accuracy of each data source are comprehensively analyzed to obtain the quality score index of each data source.

9. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The specific analysis process of the aggregation weight is as follows: Comprehensively analyze the quality score index of each data source and the status change index of each asset to obtain the aggregate weight. The specific analysis method is as follows: ; Where, is the aggregation weight, is the state change index of the a-th asset, a is the number of each asset, , b is the total assets, The weight factor corresponding to the unit value of the asset status change index preset in the modeling database, is the quality score index of the mth data source, m is the number of each data source, , n is the total number of data sources, The weight factor corresponding to the quality score index unit value of the data source preset for the modeling database.

10. The method for modeling an asset management coding model based on federated learning according to claim 1, characterized in that: The model parameter aggregation process is optimized, and the specific optimization process is as follows: Compare the aggregation weight with the values ​​of each aggregation weight interval preset in the modeling database to determine the specific interval corresponding to the aggregation weight and obtain the interval where the aggregation weight is located; According to the interval of the aggregation weight, the model parameter aggregation process in the federated learning is optimized, and finally the model parameter aggregation process is optimized.

Citation Information

Patent Citations

  • Federal domain adaptation method applied to data isomerism

    CN114881134A

  • Federal learning method based on fairness enhancement

    CN119962635A

Cited By

  • Prediction model training method, text data prediction method, computing device and storage medium

    CN121935972A

  • Predictive model training, text data prediction method, computing device, and storage medium

    CN121935972B