Federal learning small and micro enterprise credit portrait system

Through federated learning and transfer learning technology, the integration of multi-source heterogeneous data has been solved, and data silos and privacy compliance issues in credit assessment of small and micro enterprises has been achieved, dynamic and secure credit assessment has been achieved, comprehensiveness and accuracy of credit assessment has been improved, and real-time response to economic cycle fluctuations has been supported, and the coverage of inclusive finance has been expanded.

CN120471705APending Publication Date: 2025-08-12颜子淳
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510597145.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

There is a data island problem in the credit assessment of small and micro enterprises. The traditional credit assessment system relies on a single financial data, is difficult to capture dynamic operating conditions, and lacks a privacy protection mechanism, making it difficult for financial institutions to accurately evaluate and provide financial services.

Method used

Federated learning technology is used to integrate multi-source heterogeneous data, and data fusion is carried out through horizontal and vertical federated learning models, combining transfer learning and macroeconomic factors to generate dynamic credit assessment models, and ensure data privacy through edge computing and blockchain audits.

Benefits of technology

It realizes efficient integration of multi-source heterogeneous data without sharing original data, improves the comprehensiveness and accuracy of credit assessments, supports real-time response to economic cycle fluctuations, reduces credit risks, expands the coverage of inclusive finance, and meets privacy compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471705A_ABST
    Figure CN120471705A_ABST
Patent Text Reader

Abstract

The invention discloses a federal learning small and micro enterprise credit portrait system, and relates to the technical field of financial science and technology. The data access module deeply fuses the data through semantic analysis and standardized preprocessing of the multi-source heterogeneous data, and breaks through the limitation of a traditional data island; the federation calculation module dynamically selects a transverse or longitudinal federation learning mode according to the data type, and realizes cross-mechanism and cross-modal data feature collaborative optimization by using an attention mechanism and a dynamic alignment algorithm; the credit evaluation module dynamically corrects credit score deviation in a sparse data scene through transfer learning and macroeconomic factor embedding, and generates an interpretability report to improve model transparency; and the application service module is combined with lightweight deployment and edge computing technologies, so that credit scores quickly respond to business requirements. All the modules guarantee data privacy through hierarchical encryption and block chain auditing, a complete closed loop from data collection to credit output is formed, and safe, real-time and high-precision credit evaluation service is provided for credit portraits of small and micro enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of financial technology, and in particular to a federated learning credit profiling system for small and micro enterprises. Background Art

[0002] Small and micro enterprises (SMEs) are a vital pillar of the national economy, but their financing challenges remain unresolved. Traditional credit assessment systems rely heavily on collateral and financial statements. However, due to their asset-light nature, lack of financial statements, and inadequate records, most SMEs are considered "credit-free," making it difficult for them to access formal financial services. Financial institutions face a dual dilemma: on the one hand, SMEs' core operating data is scattered across tax, electricity, logistics, and other agencies, creating data silos that prevent traditional credit reporting models from capturing true risks. On the other hand, high due diligence costs and sluggish manual review processes further exacerbate financial institutions' reluctance to lend. Existing scoring models rely on single-source financial data, making it difficult to capture the dynamic operating conditions of SMEs. This is particularly true for businesses lacking collateral and historical records, leading to significant bias in assessments, severely restricting the targeted allocation of financial resources.

[0003] In recent years, technological breakthroughs have provided new paths to address this dilemma. The maturity of federated learning technology has enabled collaborative data modeling across institutions, unlocking the value of dormant multi-source data while ensuring privacy and security. Advances in multimodal learning technology have broken down the semantic barriers between heterogeneous data such as text, time series, and tables, significantly improving the comprehensiveness of credit assessments. Meanwhile, the widespread adoption of edge computing and high-speed networks provides the underlying support for real-time credit inference. At the policy level, regulators are vigorously promoting the digital transformation of inclusive finance, explicitly requiring financial institutions to improve the efficiency of small and micro service coverage. Driven by both market demand and technological substitution, the demand for a new generation of credit assessment systems is urgent. A secure, dynamic, and low-cost technical solution is urgently needed to transform dispersed data footprints into quantifiable credit assets and build a cross-institutional and cross-data type credit assessment platform. Summary of the Invention

[0004] In response to the above-mentioned shortcomings of the existing technology, the present invention provides a federated learning credit profiling system for small and micro enterprises.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0006] The Federated Learning Small and Micro Enterprise Credit Profiling System includes:

[0007] A data access module is used to obtain heterogeneous data from multiple data sources and perform standardized preprocessing on the heterogeneous data;

[0008] The federated computing module is used to select a horizontal or vertical federated learning mode according to the data type of the heterogeneous data, schedule the local models of multiple participants for training, and generate a global credit assessment model through parameter aggregation, wherein:

[0009] If the horizontal federated learning model is selected, cross-institutional data with the same characteristics but different samples will be jointly trained, and aggregation weights will be dynamically allocated according to the data integrity indicators and feature redundancy of the participants.

[0010] If the longitudinal federated learning mode is selected, cross-modal data with different features of the same sample are aligned and trained, and sample space alignment is achieved through a dynamic time alignment loss function;

[0011] A credit assessment module, configured to output a dynamic credit score based on the global credit assessment model and integrate transfer learning, and generate an interpretable report including feature contribution;

[0012] The application service module is used to deploy the dynamic credit score to low-configuration devices and provide an application programming interface (API) for preset scenarios.

[0013] Beneficial effects

[0014] Compared with the known public technology, the technical solution provided by the present invention has the following beneficial effects:

[0015] 1. The present invention effectively solves the problems of data silos and privacy compliance in the credit assessment of small and micro enterprises through a federated learning framework. Without sharing the original data, the system integrates multi-source heterogeneous data such as bank statements, tax invoices, supply chain evaluations and IoT sensors, and uses horizontal federated learning to jointly train cross-institutional data with different samples of the same features, and realizes time-space alignment of cross-modal data through vertical federated learning. For example, in the vertical federation mode, a sliding window interpolation algorithm is used to dynamically align low-frequency tax data with high-frequency bank statements, and the dynamic time alignment loss function is combined to ensure the consistency of feature space, thereby breaking through the limitations of traditional models for multi-source data fusion. This design not only complies with the requirements of privacy regulations such as GDPR, but also significantly improves data utilization, enabling financial institutions to cover "three-no enterprises" (no mortgage, no report, no credit record) that are difficult to reach with traditional credit investigation.

[0016] 2. Through transfer learning technology, pre-trained industry-wide model knowledge is transferred to small and micro enterprises (SMEs) with sparse data. Feature weights are dynamically adjusted based on macroeconomic factors such as the PMI index and industry prosperity, enabling credit scores to respond in real time to economic cycle fluctuations. The PMI change rate is embedded in the gated network to dynamically adjust the credit weights of SMEs in different industries. Furthermore, the model compression unit uses channel pruning and 8-bit integer quantization to reduce the size of the global credit assessment model, enabling real-time execution on ARM-based edge devices. This allows county-level financial institutions to output credit scores in seconds without the need for high-performance hardware, significantly expanding the reach of inclusive finance.

[0017] 3. This invention establishes a secure and efficient ecosystem for credit assessment of small and micro enterprises. The blockchain audit module records the data contributions and model versions of each participant through a hash chain, ensuring traceability and compliance of data usage and providing a transparent audit interface for regulators. Furthermore, the cold start compensation module uses transfer learning to integrate industry benchmark parameters with the legal representative's historical credit history to generate an initial score with confidence correction for newly registered companies, resolving the assessment challenge in data-deficient scenarios. This technological paradigm not only helps financial institutions accurately allocate credit resources but also promotes digital guidance of small and micro enterprise business operations, injecting sustainable financial vitality into the development of the real economy. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0019] Figure 1 This is a system module connection diagram of the present invention;

[0020] Figure 2 This is a diagram of the blockchain audit module and cold start compensation module of the present invention. DETAILED DESCRIPTION

[0021] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, a person skilled in the art can propose a variety of interchangeable structural modes and implementation modes. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention and should not be regarded as the entire invention or as a limitation or restriction of the technical solution of the present invention.

[0022] Application Overview:

[0023] In existing technologies, credit assessments for small and micro enterprises mostly rely on static financial data or manual due diligence information from a single financial institution, making it difficult to integrate multi-source heterogeneous data (such as bank statements, electronic invoices, logistics evaluations, etc.) scattered across taxation, supply chain, the Internet of Things, and other fields. Traditional methods have weak processing capabilities for unstructured data and rely on manual labeling or simple rules to extract features, which is inefficient and difficult to capture semantic associations. Static methods such as mean replacement are often used to fill in missing values in structured data, which can easily introduce statistical biases, and the manual cleaning and normalization processes are time-consuming and labor-intensive. In addition, traditional model updates lag behind economic cycle fluctuations and cannot respond to dynamic data (such as changes in industry prosperity) in real time. In addition, there is a lack of privacy protection mechanisms, and cross-institutional data sharing carries the risk of leaking sensitive information (such as corporate ID numbers and bank account numbers).

[0024] The inventors found that the core difficulty in credit assessment of small and micro enterprises lies in the problem of data silos and dynamic scenario adaptation. By analyzing the horizontal and vertical data fusion mechanisms under the federated learning framework, it was found that the data quality (such as integrity and redundancy) of different participants in horizontal federated learning has a significant impact on model performance, while the time series alignment capability of cross-modal data in vertical federated learning is the key to improving assessment accuracy. Further experiments show that the dynamic aggregation weight distribution based on the multi-head attention mechanism can effectively screen high-quality data sources, and the time-space alignment loss function can significantly reduce the feature deviation of cross-modal data. In addition, the fusion of transfer learning and macroeconomic factors (such as the PMI index) can solve the problem of data sparsity of small and micro enterprises and enhance the trust of financial institutions through explainable reports.

[0025] Compared to traditional methods, this application uses federated learning technology to construct a credit profiling system for small and micro enterprises, effectively addressing the challenges of data silos and privacy compliance in traditional credit assessments. Without sharing the original data, the system integrates heterogeneous data from multiple sources, such as banks, tax authorities, and supply chains. Using a multi-head attention mechanism and a dynamic time series alignment algorithm, it overcomes the semantic barriers of cross-institutional and multimodal data, significantly improving the comprehensiveness and accuracy of credit assessments. Based on transfer learning and macroeconomic sensitive factor embedding technology, the model can automatically adapt to the sparsity of small and micro enterprise operating data and respond in real time to external variables such as industry prosperity and economic cycle fluctuations, enabling dynamic adjustment of credit scores. Furthermore, through edge computing architecture and model pruning and quantification techniques, the system compresses complex federated models, enabling real-time execution on low-profile equipment at county-level financial institutions, eliminating the traditional solution's reliance on high-performance hardware. This "safe, efficient, and dynamically adaptable" approach not only meets the real-time risk control needs of financial institutions but also provides a scalable technical foundation for inclusive financial services in lower-tier markets.

[0026] This invention restructures the underlying logic of credit assessment for small and micro enterprises, promoting the deep integration of financial technology and the real economy. By activating the "data footprint" scattered across government, logistics, e-commerce, and other fields, the system transforms the daily operations of small and micro enterprises into quantifiable credit assets, guiding them to standardize operations and enhance their financing capabilities. Furthermore, leveraging blockchain auditing and standardized data interfaces, the system lays the foundation for cross-industry and cross-regional credit recognition.

[0027] After introducing the basic concept of the present invention, embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0028] Example

[0029] The federated learning small and micro enterprise credit profiling system in this embodiment is as follows: Figure 1 Shown, including:

[0030] Data access module, used to obtain heterogeneous data from multiple data sources and perform standardized preprocessing on the heterogeneous data;

[0031] Heterogeneous data includes bank transaction data, tax invoice data, supply chain evaluation data, and IoT sensor data; preprocessing includes semantic analysis of unstructured data and missing value filling and normalization of structured data.

[0032] Among them, bank transaction data records the flow of funds in and out of bank accounts, including information such as transaction time, transaction amount, and counterparty. It can be used to analyze the cash flow and financial status of enterprises or individuals. Tax invoice data reflects a company's business activities and tax payment status, including invoice number, invoice date, product or service name, amount, tax amount, etc., which assists in tax supervision, corporate financial accounting, and tax planning. Supply chain evaluation data is used to assess the performance and capabilities of various links and partners in the supply chain, such as suppliers' on-time delivery rate, product quality, and price competitiveness. It is important for optimizing supply chain management, reducing costs, and improving efficiency. IoT sensor data is real-time data collected by various IoT devices (such as temperature sensors, humidity sensors, and pressure sensors). It reflects the status of the physical environment or equipment and can be applied in various fields such as enterprise industrial monitoring, smart homes, and environmental monitoring.

[0033] Unstructured data, such as shopkeepers' social media data, typically lacks a fixed format or structure. Semantic parsing uses technologies like natural language processing and image recognition to understand the meaning of this unstructured data and convert it into structured information that computers can process and analyze.

[0034] Structured data, such as bank statements, is data with a fixed format and structure. Missing value filling involves using appropriate methods (such as mean filling, median filling, and model prediction filling) to fill missing values when certain fields are missing in the data, ensuring data integrity. Normalization involves scaling data according to specific rules to bring them into a consistent scale range, eliminating the impact of different dimensions between different features and improving model stability and accuracy.

[0035] After acquiring this data, the data access module begins preliminary preprocessing. For unstructured data, semantic parsing methods from natural language processing (NLP) are utilized. For example, word vector models are used to convert the words in evaluation statements into vector representations. Deep learning models such as recurrent neural networks (RNNs) are then combined to understand the meaning of the statements, extract valuable information, and convert it into structured data. For structured data, missing values are primarily checked and filled. For example, some transaction records in bank flow data may lack transaction notes. In this case, appropriate filling can be performed based on the historical transaction habits of the micro-enterprise or the general situation in the same industry. For example, for routine small transfers, missing notes can be filled in as "daily operating expenses." All data is also normalized. For example, transaction amounts in bank flow data are unified into a specific range through mathematical transformations to facilitate subsequent calculations and processing.

[0036] Semantic parsing of unstructured data: Text data processing (supply chain evaluation data) uses a pre-trained BERT model to extract semantic features, such as mapping "delivery delay of 3 days" to a numerical label (such as "delivery timeliness score: 2 / 5"); extracts key entities (such as supplier name, product specifications) through named entity recognition (NER) and generates structured fields; uses an LSTM model to determine the evaluation polarity (positive / negative / neutral) and outputs a comprehensive score (such as "service quality: 4.2 / 5"); processes sensor data and parses the original signal into structured time series data, such as converting the vibration frequency signal into frequency domain features through Fourier transform and annotating the timestamp and device ID.

[0037] Normalization of structured data: For missing fields in the "Transaction Notes" section of bank statements, virtual values are generated based on historical data distribution (e.g., "Small and Micro Enterprise Transfer" is filled in by default); for missing fields in the "Product Name" section of tax invoices, the industry type of the associated invoicing enterprise is inferred (e.g., "Raw Material Purchase" is filled in by default for manufacturing enterprises).

[0038] Normalization of structured data: Bank transaction amounts are scaled to the [0, 1] range using the Min-Max algorithm, and sensor data (such as temperature values) are normalized using the Z-Score to eliminate dimensional differences.

[0039] The data access module also includes a data quality verification unit for detecting the missing rate and noise level in heterogeneous data. If the missing rate exceeds a preset missing threshold, the data source is triggered to re-upload or generate virtual fill data.

[0040] The data quality verification unit conducts comprehensive quality checks on the acquired heterogeneous graph data, focusing on two key metrics: missingness rate and noise level. For missingness rate testing, using bank transaction data as an example, the missingness of each field is counted and its contribution to the overall data is calculated. If the missingness rate of a field exceeds a pre-set threshold, such as a missingness rate of more than 30% for counterparty information, a corresponding mechanism is triggered. If the data is obtained from a retransmittable data source, such as some banking systems that support data re-upload, a retransmission request is sent to the data source, requesting that it re-upload complete and accurate data. The data provider is notified via an API to retransmit the missing data, and a retransmission log is recorded. For data that cannot be retransmitted or is costly to retransmit, virtual filler data is generated to address missing data. This process may utilize data generation models, such as those based on variational autoencoders (VAEs), which learn the distribution characteristics of existing complete data and generate virtual data that conforms to that distribution to fill in the missing data.

[0041] For noise level detection, taking IoT sensor data as an example, due to the influence of various factors such as the environment, some outliers may exist. Statistical methods and machine learning algorithms are used to identify these noise points. These noise points are then appropriately corrected or removed to ensure data quality and meet the requirements of subsequent federated learning model training.

[0042] Compared with traditional technologies, traditional credit assessment methods primarily rely on the financial data of a single financial institution or manual due diligence information, making it difficult to integrate heterogeneous data from multiple sources across sectors such as taxation, supply chain, and the Internet of Things, leading to prominent data silos. Furthermore, traditional methods are weak in processing unstructured data (such as supply chain evaluation data), often relying on manual labeling or simple rule-based feature extraction, which is inefficient and difficult to capture semantic associations. Missing values in structured data are often filled using static methods such as mean substitution, which can easily introduce statistical biases. Manual data cleaning and normalization processes are time-consuming and labor-intensive, making them difficult to support real-time credit assessment needs.

[0043] Through the above technical solutions, core problems such as cross-institutional data security integration, unstructured semantic analysis, and dynamic missing value filling have been solved, and full-dimensional coverage and real-time updates of small and micro-enterprise credit profiles have been achieved. This application realizes the efficient integration and secure collaboration of small and micro-enterprise credit assessment processes. The data access module deeply integrates structured data such as bank statements and tax invoices with supply chain evaluations and sensor time series data through semantic analysis and standardized preprocessing of multi-source heterogeneous data, breaking through the limitations of traditional data silos. The data access module realizes the automation of the entire process from multi-source heterogeneous data collection, quality repair to federation-ready format conversion. Among them, semantic analysis and virtual filling technology solve the processing problems of unstructured data and missing data, while the federation-ready format design significantly improves the efficiency of subsequent joint modeling, providing a reliable data input foundation for the small and micro-enterprise credit profile system.

[0044] The federated computing module is used to select the horizontal / vertical federated learning mode based on the data type of heterogeneous data, schedule the local models of multiple participants for training, and generate a global credit assessment model through parameter aggregation, where:

[0045] If the horizontal federated learning model is selected, cross-institutional data with the same characteristics but different samples will be jointly trained, and aggregation weights will be dynamically allocated according to the data integrity indicators and feature redundancy of the participants;

[0046] If the longitudinal federated learning mode is selected, cross-modal data with the same sample but different features are aligned and trained, and sample space alignment is achieved through the dynamic time alignment loss function.

[0047] The federated computing module includes:

[0048] The multi-head attention feature extraction unit is used in the horizontal federated learning model to extract the feature vectors of cross-institutional data through the Transformer architecture, generate data integrity indicators and feature redundancy, and calculate the cross-institutional attention weight. The calculation formula of the attention weight is:

[0049]

[0050] Among them, Q i , K j are query vectors and key vectors of different modalities, d k is the preset dimension;

[0051] The time-space alignment unit is used in the longitudinal federated learning model to generate pseudo data aligned with the high-frequency data timestamps of the cross-modal data through sliding window interpolation for the low-frequency data of the cross-modal data, and constrain the feature space consistency through the dynamic time alignment loss function. The loss function is:

[0052]

[0053] in, and are the feature vectors of high-frequency data and interpolated low-frequency data respectively.

[0054] The anomaly detection subunit is used by the time-space alignment unit to detect abnormal points of data fluctuation during the interpolation process and correct them through linear interpolation to ensure time series continuity;

[0055] The layered encryption unit is used to decrypt and process highly sensitive data in a trusted execution environment (TEE), and to transmit gradient parameters using a homomorphic encryption algorithm. Highly sensitive data includes corporate ID numbers and bank account numbers. The homomorphic encryption algorithm is the Paillier algorithm, and the single encryption time is less than the preset time threshold.

[0056] The horizontal federated learning model focuses on joint training of cross-institutional data with the same features but different samples. In this model, data from each participating institution shares the same feature dimensions but uses different samples. By combining these data for training, we can fully leverage the data resources of each institution and improve the model's generalization capabilities. During the training process, dynamically assigning aggregation weights based on the data integrity indicators and feature redundancy of each participant is a key step. Data integrity indicators reflect the quality and availability of data. Data with high integrity contributes more to model training and should be assigned a higher weight. Feature redundancy, on the other hand, measures the degree of duplication of data features. Data with low redundancy can bring more new information to the model and should also be given an appropriate weight. This dynamic weighting ensures that each participant's data plays a reasonable role in global model training, preventing model performance from being affected by data quality or feature duplication.

[0057] In the horizontal federated learning model, the multi-head attention feature extraction unit plays a key role, leveraging the Transformer architecture to accomplish a series of important tasks. First, the unit extracts feature vectors from cross-institutional data. The Transformer architecture, with its powerful parallel computing capabilities and ability to process sequential data, can deeply mine the data for latent features and transform the raw data into representative feature vectors. Based on the extracted feature vectors, it further generates data integrity metrics and feature redundancy. Data integrity metrics reflect the completeness and availability of data, helping to assess the quality of each participant's data. Feature redundancy measures the degree of duplication of data features, preventing the introduction of excessive duplicate information during model training. Furthermore, the unit calculates cross-institutional attention weights by measuring the similarity between query and key vectors to determine the degree of correlation between data from different modalities. Higher similarity values result in greater attention weights, meaning that data will receive more attention during model training. In this way, the multi-head attention feature extraction unit effectively integrates cross-institutional data, providing strong support for subsequent model training.

[0058] The vertical federated learning model aligns and trains cross-modal data with the same sample but different features. In practical applications, different organizations may possess different types of data on the same object, which may differ in time and feature space. For example, a vertical federated learning model can be established for a tax bureau (corporate tax data) and a power company (power stability data). The core task of the vertical federated learning model is to address these differences and achieve effective data integration. During training, cross-modal data must first be aligned. For low-frequency data, pseudo data is generated using a sliding window interpolation method that aligns the timestamps of the high-frequency data, ensuring temporal consistency across the different modalities. A dynamic time alignment loss function is then used to constrain feature space consistency. This loss function measures the difference between the feature vectors of the high-frequency data and the interpolated low-frequency data. By minimizing this loss function, data from different modalities are brought closer together in feature space, thereby achieving alignment in the sample space. This alignment method fully exploits the information from data from different modalities, improving model accuracy and reliability.

[0059] In the longitudinal federated learning model, the time-space alignment unit is primarily responsible for processing low-frequency data in cross-modal data. Due to differences in temporal frequency between data from different modalities, the timestamps of low-frequency data often differ from those of high-frequency data. To address this issue, the unit uses sliding window interpolation to process the low-frequency data and generate pseudo-data that is aligned with the timestamps of the high-frequency data. Sliding window interpolation is an interpolation method based on local data. It slides a fixed-size window across the time series of the low-frequency data and uses the data within the window to estimate the values of missing time points. This ensures that the low-frequency data is temporally consistent with the high-frequency data, laying the foundation for subsequent feature space alignment. To constrain feature space consistency, the unit uses a dynamic time alignment loss function. This loss function measures the difference in feature space between the high-frequency data and the interpolated low-frequency data. By minimizing this loss function, the data from different modalities are brought closer together in the feature space, thereby achieving sample space alignment and improving model accuracy and reliability.

[0060] The anomaly detection subunit plays a crucial role in the interpolation process performed by the time-space alignment unit. Its primary task is to detect outliers in data fluctuations and correct them to ensure time series continuity. The anomaly detection subunit employs various methods to detect outliers. For example, it sets a threshold to determine whether a data point deviates from the normal range. If a data point's value exceeds the preset threshold, it is considered an outlier. Statistical analysis methods, such as mean and variance, can also be used to identify unusual fluctuations in the data. Once an outlier is detected, the anomaly detection subunit corrects it using linear interpolation. Linear interpolation is a simple yet effective method that estimates a reasonable value for the outlier based on a linear relationship between the normal data points before and after the outlier. This method eliminates the impact of outliers on the data time series and ensures data continuity and stability. Ensuring time series continuity is crucial for model training. The presence of outliers in the data can cause the model to learn erroneous information, impacting its performance and accuracy. Therefore, the work of the anomaly detection subunit improves data quality and provides a reliable data foundation for subsequent model training.

[0061] The layered encryption unit is primarily responsible for the encrypted transmission of highly sensitive data and gradient parameters. Within a Trusted Execution Environment (TEE), this unit securely decrypts and processes highly sensitive data. The TEE divides the processor into independent secure enclaves (such as Intel SGX). After decryption, data is processed only within this "secure enclave," preventing external intrusion. Hardware-level encryption ensures data security for all participants and supports multi-platform deployments, including Intel SGX and ARM TrustZone. Highly sensitive data includes important information such as corporate ID numbers and bank account numbers, the security of which is directly related to the interests of both businesses and individuals. To ensure data security, the layered encryption unit uses homomorphic encryption for gradient parameter transmission. Homomorphic encryption is a specialized encryption algorithm that allows specific computations to be performed on encrypted data without first decrypting the data. In this module, the Paillier algorithm is used. Paillier homomorphic encryption supports addition operations in a ciphertext state, meeting the gradient aggregation requirements of federated learning. The Paillier algorithm has excellent encryption performance and security, and can effectively protect the privacy of gradient parameters during transmission. A single calculation takes less than 0.3 seconds (traditional methods take more than 2 seconds). At the same time, to meet the needs of practical applications, the algorithm's single encryption time is less than the preset time threshold. This means that while ensuring data security, encryption operations can be completed quickly, improving the system's operating efficiency. Through the work of layered encryption units, the federated computing module can achieve secure data transmission and processing while protecting data privacy, providing reliable protection for the training of credit assessment models.

[0062] In addition, in addition to horizontal federation (i.e., situations with a lot of sample overlap) and vertical federation (i.e., situations with a lot of feature overlap), the federated computing module also adopts a variety of flexible methods such as hybrid federation. In a more complex heterogeneous data environment, hybrid federation comprehensively utilizes the characteristics of horizontal and vertical federation, and flexibly switches or combines the two modes according to the actual data structure and needs to achieve efficient processing and utilization of heterogeneous data, thereby providing strong data support for the generation of accurate and reliable global credit assessment models.

[0063] Practical application cases:

[0064] In the field of credit assessment, federated computing modules have demonstrated significant application value. For example, a large financial institution alliance, comprised of multiple banks and fintech companies, holds a vast amount of customer data, but due to privacy and security concerns, direct sharing is difficult. By introducing federated computing modules, these institutions can jointly train credit assessment models without leaking the original data.

[0065] In the horizontal federated learning model, institutions jointly train customer data on samples with the same characteristics but different characteristics. For example, different banks possess credit data on different customer groups. By dynamically assigning aggregate weights, the model leverages the data strengths of each institution and improves its generalization capabilities. In practical applications, this model has improved the accuracy of credit assessments for new customers by 15% compared to traditional models, effectively reducing credit risk.

[0066] In a vertical federated learning model, financial institutions collaborate with e-commerce platforms to align cross-modal data with similar but different characteristics. Financial institutions possess customer financial information, while e-commerce platforms possess customer consumption behavior data. Through processing using a temporal-spatial alignment unit and anomaly detection sub-units, this data is effectively integrated. The jointly trained credit assessment model enables a more comprehensive assessment of a customer's creditworthiness, improving the identification rate of potential defaulters by 20%.

[0067] Compared with traditional technologies, traditional credit assessment mainly relies on static data from a single institution or limited data sources (such as financial statements and historical credit records). It is difficult to integrate heterogeneous data across institutions and modalities (such as bank statements, tax invoices, supply chain evaluations, etc.), resulting in prominent data silos. Traditional methods have weak processing capabilities for unstructured data and often rely on manual labeling or simple rules to extract features. This is inefficient and difficult to capture semantic associations. Missing values in structured data are often filled with static methods such as mean replacement, which can easily introduce statistical biases, and the manual data cleaning and normalization process is time-consuming and labor-intensive. In addition, traditional model updates lag behind business changes and cannot respond to dynamic data (such as IoT sensor data) in real time. In addition, there is a lack of privacy protection mechanisms, and there is a risk of sensitive information leakage during data sharing.

[0068] By integrating heterogeneous data from multiple sources through this technical solution, the model can obtain more comprehensive information, enabling a more accurate assessment of a customer's credit status. This helps financial institutions more accurately identify credit risk, reduce default rates, and ensure the stable operation of the financial market. Secondly, the federated computing module effectively protects data privacy. In traditional credit assessment models, data sharing often carries the risk of privacy leakage. However, the federated computing module utilizes layered encryption units to process highly sensitive data in a trusted execution environment and employs homomorphic encryption algorithms to transmit gradient parameters, ensuring data security and privacy throughout the entire process. Furthermore, this module promotes industry collaboration and development. Different institutions can collaborate without leaking core data, enabling resource sharing and complementary strengths.

[0069] This cross-institutional federated learning framework integrates multimodal information, including bank statements (time series data), tax invoices (tabular data), supply chain reviews (text data), and IoT devices (sensor data), without leaving the data domain. Through dynamic weight allocation and feature alignment algorithms, it generates a 360-degree credit profile of small and micro enterprises. In practical applications, this module has demonstrated strong adaptability and flexibility. It intelligently matches the most appropriate learning model to different types of small and micro enterprise data, ensuring effective data utilization and efficient model training. By combining local models from multiple participants, the federated computing module breaks down data silos, enabling data sharing and collaboration, thereby improving model accuracy and reliability. In the field of credit assessment, the federated computing module is of great significance. It integrates data from multiple sources to comprehensively assess credit status, providing more accurate credit assessment results for financial institutions and other institutions, reducing credit risk and promoting the healthy development of financial markets.

[0070] The credit assessment module is used to output dynamic credit scores based on the global credit assessment model and integrate transfer learning, and generate an explainable report including feature contribution.

[0071] The credit assessment module includes a transfer learning unit, which pre-trains a basic model based on pre-set industry common data. Through federated learning, the knowledge of the basic model is transferred to small and micro enterprises with sparse data, generating virtual feature vectors and macroeconomic sensitivity factors. The loss function of transfer learning is:

[0072]

[0073] Among them, λ is the preset migration weight coefficient, and KL divergence is used to constrain the feature distribution difference between the source domain and the target domain.

[0074] Virtual feature vectors are generated by extracting and transforming features from the data during the transfer learning process. This can supplement the deficiencies in small and micro-enterprise data and improve the model's generalization capabilities. The KL divergence in the transfer loss function measures the difference in feature distribution between the source domain (preset industry-wide data) and the target domain (small and micro-enterprise data). By minimizing the KL divergence, the feature distributions of the source and target domains are made as close as possible, thereby improving the effectiveness of transfer learning. λ, as a preset transfer weight coefficient, balances the effects of the KL divergence and other loss terms, ensuring the stability and effectiveness of the transfer process.

[0075] The strategy for correcting feature distribution discrepancies is key to addressing data sparsity. By constraining the KL divergence in the transfer loss function to account for the difference in feature distributions between the source and target domains, the model can better adapt to the data characteristics of small and micro enterprises. The λ weight coefficient plays a balancing role in this process, adjusting the influence of the KL divergence based on actual conditions. This ensures that while correcting for feature distribution discrepancies, the model avoids overfitting to the source domain data, thereby improving the model's generalization ability on small and micro enterprise data.

[0076] The credit assessment module also includes a macroeconomic factor embedding unit, which is used to input the Purchasing Managers' Index (PMI) and industry prosperity into the gated network and dynamically adjust the feature weights. The weight update formula is:

[0077] w t =sigmoi(W·[h t ;ΔPMI t ])

[0078] Among them, ΔPMI t is the rate of change of PMI index, h t is the current eigenvector, W is a weight matrix, which plays the role of linear transformation of other vectors in the entire calculation process. t ;ΔPMI t ] can be multiplied to integrate and weight the information contained in these vectors.

[0079] In the macroeconomic factor embedding unit, the fusion path for macroeconomic factors and feature vectors is as follows: First, the PMI index change rate and industry prosperity index are fed into the gating network as input. The gating network calculates the adjustment coefficient based on pre-set rules. Then, according to the weight update formula described above, the macroeconomic factors are fused with the feature vectors to achieve dynamic adjustment of feature weights. This allows the credit assessment model to adjust feature weights in real time based on macroeconomic changes, ensuring that credit scores are more consistent with actual economic conditions.

[0080] The macroeconomic factor embedding unit also includes an industry classification sub-unit, which is used to classify and code small and micro enterprises according to preset industries and load the prosperity indicators of the corresponding industries. The classification codes include retail, manufacturing and logistics.

[0081] The classification and coding rules for the retail, manufacturing, and logistics industries are based on the industry's business characteristics and economic attributes. The retail industry primarily involves the sale of goods, with coding focusing on factors such as sales channels and product types. The manufacturing industry focuses on product production, with coding taking into account production processes, production capacity, and scale. The logistics industry focuses on the transportation and storage of goods, with coding related to transportation methods and storage capacity. The mechanism for loading industry prosperity indicators extracts prosperity data for the corresponding industry from relevant databases based on classification codes. For example, for the retail industry, indicators such as the consumer confidence index and sales growth rate are loaded; the manufacturing industry focuses on industrial added value and order volume; and the logistics industry prioritizes freight volume and warehouse utilization.

[0082] Differentiated scoring strategies across industries are reflected in their varying sensitivities to macroeconomic factors. For example, in the retail sector, during periods of economic prosperity, consumer confidence rises, sales increase, and credit scores improve accordingly. Conversely, during periods of economic downturn, consumer spending declines, and scores may decline. The manufacturing industry is significantly influenced by the PMI index. A rising PMI indicates robust manufacturing activity, which in turn boosts credit scores. The logistics industry, on the other hand, is closely tied to freight volume, a key indicator of industry prosperity. Increased freight volume also correlates with higher credit scores. This differentiated scoring strategy allows for a more accurate assessment of the credit standing of small and micro enterprises across different industries.

[0083] The credit assessment module also includes an explainability generation unit, which uses SHAP values to analyze feature contributions and generate an explainability report containing key influencing factors. This explainability report can also utilize methods such as LIME to output key factors influencing the credit score (such as the weight of tax growth rate), thereby enhancing the trust of financial institutions.

[0084] In credit scoring, quantifying key influencing factors, such as tax growth rates, is achieved using SHAP values. SHAP values, based on the Shapley value in game theory, measure the marginal contribution of each feature to the model's predictions. SHAP and LIME differ in their application scenarios. SHAP values are suitable for global explanations, comprehensively demonstrating the contribution of each feature to the entire model. They are very effective for understanding the overall model behavior and interactions between features. LIME, on the other hand, focuses more on local explanations. It explains individual predictions by generating local approximate models near the prediction point, making it suitable for detailed interpretation of specific samples. Feature contribution visualization involves displaying SHAP values graphically. Common visualization methods include waterfall charts and bee swarm plots. Waterfall charts clearly demonstrate the order and magnitude of each feature's contribution to the final prediction, starting with a baseline value and gradually adding up the contributions of each feature until the final prediction value is reached. Bee swarm plots display the SHAP values of each feature in each sample as scattered points, providing a visual representation of the distribution of feature contributions and helping analysts quickly identify key influencing factors.

[0085] Explainable reports are crucial for credit risk management decisions. They provide financial institutions with a clear basis for decision-making, enabling them to understand the credit score calculation process and key influencing factors. For example, through explainable reports, financial institutions can clearly identify the impact of factors such as tax growth rates on corporate credit, enabling more accurate risk assessments. This helps financial institutions avoid blind decisions and reduce credit risk. The correlation between dynamic scores and macroeconomic cycles can be presented in charts and reports. For example, a time series graph of credit scores against macroeconomic indicators such as the PMI index and industry prosperity can be used to visually demonstrate the changing relationship between the two. This allows financial institutions to clearly visualize how corporate credit scores fluctuate over the macroeconomic cycle, better understanding risk trends.

[0086] The logical framework for dynamically adjusting credit scores is as follows: Macroeconomically sensitive factors, such as the PMI index and industry prosperity, are embedded in the dynamic credit assessment model, enabling credit scores to adjust dynamically with the economic cycle. A transfer learning mechanism is also introduced to address the data sparsity issue for small and micro enterprises. The interactive process between the various units is as follows: The transfer learning unit first completes training of the basic model and knowledge transfer, generating relevant feature vectors and factors. The macroeconomic factor embedding unit receives this information and dynamically adjusts feature weights based on the PMI index and industry prosperity. Finally, the interpretability generation unit analyzes the adjusted results and generates an interpretable report to inform financial institutions' decision-making. The transfer learning unit is the foundational support for the module. It pre-trains a basic model based on pre-set industry-wide data and then uses federated learning to transfer the knowledge of this basic model to data-sparse small and micro enterprises. This process generates virtual feature vectors and macroeconomic sensitive factors. In the transfer learning loss function, λ is used as the pre-set transfer weight coefficient, and the KL divergence is used to constrain the difference in feature distribution between the source and target domains to ensure the effectiveness of the transfer. The macroeconomic factor embedding unit is responsible for inputting the Purchasing Managers' Index (PMI) and industry prosperity indicators into the gated network, dynamically adjusting feature weights. This unit also includes an industry classification subunit, which can classify and code small and micro enterprises according to preset industries, such as retail, manufacturing, and logistics, and load the corresponding industry prosperity indicators. The interpretability generation unit analyzes feature contributions using SHAP values and generates interpretable reports containing key influencing factors. It can also use methods such as LIME to output key influencing factors of credit scores, thereby enhancing the trust of financial institutions.

[0087] Practical application cases:

[0088] The complete process for migrating the manufacturing industry's basic model to the logistics industry is as follows: First, pre-training is performed using deep learning algorithms to build a basic model that learns the common characteristics and patterns of the manufacturing industry.

[0089] Next, federated learning is used to migrate the knowledge of the underlying model to the logistics industry. During this migration process, each logistics company trains the model locally and only uploads the model parameters to a central server for aggregation to ensure data security. Simultaneously, virtual feature vectors are generated to supplement the deficiencies in logistics industry data.

[0090] To optimize the efficiency of generating virtual feature vectors, feature selection algorithms can be used to filter out features that have a greater impact on credit scores and reduce unnecessary calculations. Parallel computing technology can also be used to increase the speed of feature extraction and conversion.

[0091] The KL divergence is used as a quantitative indicator for feature distribution alignment. By minimizing the KL divergence of manufacturing and logistics data, the feature distributions of the two industries are brought as close as possible. When the KL divergence is less than 0.1, the feature distribution alignment is considered to be good, and transfer learning is more effective. After the transfer is completed, the credit scores for the logistics industry are evaluated and verified. By comparing with actual data, model parameters are continuously adjusted to ensure the accuracy and reliability of the credit scores.

[0092] Compared with traditional technologies, traditional models often face large prediction errors when processing sparse data from small and micro enterprises. Due to insufficient data volume, it is difficult for the model to learn comprehensive features and patterns, resulting in inaccurate predictions for data-sparse objects such as newly registered enterprises. Traditional credit assessment methods mainly rely on the financial data of a single financial institution or manual due diligence information, and it is difficult to integrate multi-source heterogeneous data such as bank statements, tax invoices, and supply chain evaluations, resulting in a prominent data silo problem. Traditional models have weak processing capabilities for unstructured data and rely on manual labeling or simple rules to extract features, which is inefficient and difficult to capture semantic associations. Missing values in structured data are often filled with static methods such as mean replacement, which is prone to introduce statistical bias, and the manual cleaning and normalization process is time-consuming and labor-intensive. In addition, traditional methods lack a dynamic response mechanism, and credit scores cannot adapt to economic cycle fluctuations (such as changes in industry prosperity) in real time. There is also a risk of leakage of sensitive information (such as corporate ID numbers and bank account numbers) during cross-institutional data sharing. Through the federated learning framework and transfer learning technology, this solution systematically solves core problems such as cross-modal data security fusion, sparse data modeling, and dynamic feature weight adjustment, achieving full-dimensional coverage and real-time updates of credit profiles of small and micro enterprises.

[0093] Through the above-mentioned technical solution, this application achieves dynamic collaboration and high-confidence analysis of the credit assessment process for small and micro enterprises. The federated learning framework integrates heterogeneous data from multiple sources, including banks and tax authorities. Using a transfer learning unit, it transfers industry-wide model knowledge to data-sparse small and micro enterprises, generating virtual feature vectors to compensate for data gaps. KL divergence is also used to constrain feature distribution differences between the source and target domains, ensuring model generalization. The macroeconomic factor embedding unit loads the PMI index and industry prosperity indicators in real time. A gated network dynamically adjusts feature weights, enabling credit scores to precisely adapt to the differentiated needs of different industries as the economic cycle fluctuates. The interpretability generation unit, based on SHAP values and LIME methods, quantifies the contribution of key factors such as tax growth rates. It uses visual logic to reveal the underlying decision-making process of credit scoring, providing financial institutions with a transparent and traceable assessment basis. Each module is collaboratively optimized through a closed-loop feedback mechanism, forming a complete technical chain from data fusion to dynamic scoring to interpretable output. This provides a real-time, accurate, and compliant solution for precise risk control and inclusive finance scenarios.

[0094] The application service module deploys dynamic credit scoring to low-profile devices and provides application programming interfaces (APIs) for pre-defined scenarios. This module includes a model compression unit, which performs channel pruning and 8-bit integer quantization on the global credit assessment model, reducing the model size and making it compatible with ARM-based edge devices. Edge deployment: Lightweight models (such as pruning and quantization techniques) for small and micro enterprises are adapted for real-time computing on mobile devices or IoT devices.

[0095] Channel pruning significantly reduces model size. By removing redundant channels from the model, the number of parameters can be significantly reduced, effectively shrinking the model size. To ensure the accuracy of the pruned model, a series of mechanisms are employed. First, channel importance is assessed before pruning, prioritizing channels with the greatest impact on model performance. Second, after pruning, the model is fine-tuned to adapt it to the new structure.

[0096] Quantization models convert floating-point parameters into integers, significantly improving computational efficiency. 8-bit integer quantization models require only 8 bits to store each parameter during computation, significantly reducing memory usage. On low-end devices with limited memory resources, the ARM architecture edge device's instruction set provides excellent support for integer operations, enabling efficient execution of 8-bit integer operations.

[0097] To comprehensively evaluate the performance of lightweight models, we established a comprehensive evaluation metric consisting of inference latency, memory usage, and prediction accuracy. Inference latency reflects the model's response speed in real-time computing, memory usage reflects the model's demand for device resources, and prediction accuracy measures the model's evaluation accuracy.

[0098] The selection of an embedded inference framework is crucial, requiring comprehensive consideration of multiple criteria to ensure compatibility with IoT device hardware. First, performance efficiency. The framework should possess efficient inference capabilities, enabling rapid calculation of dynamic credit scores on low-profile IoT devices. For example, choosing a framework that supports multi-threaded parallel computing can fully leverage the performance of the device's multi-core processors and improve inference speed. Secondly, resource usage. IoT devices typically have limited memory and storage, so the framework should minimize these requirements. Some frameworks utilize model quantization and compression techniques, effectively reducing resource requirements and making model deployment easier on edge devices. Finally, compatibility. The framework should be compatible with various hardware platforms and operating systems to ensure stable operation across a wide range of IoT devices. Furthermore, support for multiple deep learning model formats should facilitate the integration of diverse credit assessment models. Computing resource scheduling algorithms are also crucial. Computing resources should be dynamically allocated based on IoT device hardware characteristics, such as CPU usage and remaining memory. When device resources are limited, urgent credit assessment tasks should be prioritized to ensure that scenarios with high real-time requirements are not affected.

[0099] The application service module also includes a scenario-based API unit, which is used to output dynamic credit scores based on preset scenarios and is deployed to low-configuration devices through lightweight edge computing nodes. The preset scenarios include credit approval, post-loan monitoring and supply chain finance.

[0100] In the credit approval scenario, effective features for small and micro-enterprise financing can be mined from multiple dimensions. Firstly, in addition to traditional financial indicators, operational data such as order volume and customer reviews can be incorporated to more comprehensively reflect the company's operating status. Secondly, machine learning algorithms can be used to filter and combine features, removing redundant features and constructing a more representative feature set.

[0101] To control API response latency, we employ caching and asynchronous processing. We cache commonly used feature data and model parameters to reduce recalculation. For complex computations, we employ asynchronous processing, returning preliminary results before completing detailed calculations in the background, ensuring that API response latency remains within reasonable limits.

[0102] Designing a time series data processing pipeline for the dynamic early warning API for post-loan monitoring is fundamental to post-loan monitoring. First, the collected time series data is cleaned and preprocessed to remove noise and outliers. Then, time series analysis methods, such as the ARIMA model, are used to model and forecast the data. Finally, the forecast results are compared with the actual data to determine if there are any anomalies.

[0103] The linkage between anomaly detection models and credit scores is key to achieving dynamic early warning. When the anomaly detection model detects data anomalies, it promptly updates the company's credit score. For example, if a company's repayment record shows an anomaly, the credit score will decrease accordingly, triggering an early warning mechanism, alerting financial institutions to take action. This linkage mechanism enables timely identification of potential risks and safeguards the financial security of financial institutions.

[0104] In supply chain finance, balancing data privacy protection and model reasoning is a key issue. Homomorphic encryption and differential privacy technologies are used to enable model reasoning without leaking the original data. Homomorphic encryption allows computation to be performed on encrypted data, while differential privacy protects data privacy by adding noise.

[0105] The application service module also includes an adaptive pruning strategy subunit, which is used to dynamically adjust the pruning rate according to the data distribution of each participant to ensure the loss of model accuracy.

[0106] In the federated learning framework, there are significant differences in the data distribution of each participant, and it is crucial to quantify the differences in the characteristics of the participants. First, the characteristic statistical analysis method can be used to calculate the mean, variance, skewness and other statistics of the data characteristics of each participant, and by comparing these statistics, the degree of difference in the characteristic distribution can be measured. Secondly, clustering analysis technology is used to cluster the participants according to the data characteristics. The data distribution of participants in the same category is relatively similar, while there are obvious differences between different categories. The clustering results can provide a more intuitive understanding of the characteristic differences of each participant. In the derivation of the dynamic adjustment formula of the pruning rate, the balance between model accuracy and computational efficiency is taken into account. Assume that the initial pruning rate is (p0), and the quantitative value of the characteristic difference of participant (i) is (d i ), then the adjusted pruning rate (p i ) can be expressed as (p i =p0+α@d i ), where (α) is the adjustment coefficient, which can be adjusted according to actual conditions. For participants with significantly different data distributions, the pruning rate is appropriately increased to reduce model complexity. For participants with relatively similar data distributions, the pruning rate is kept low to ensure model accuracy.

[0107] The edge device model version management solution uses a version number management mechanism, assigning each model version a unique version number. When a model is updated, the content and time of the update are recorded to facilitate subsequent traceability and management. A model repository is also established to store different versions of models. Edge devices can download the appropriate version from the repository based on their needs.

[0108] When a model needs to be updated, only the changed parts are updated, rather than the entire model. This reduces data transmission and computational complexity, lowering the performance requirements for edge devices. For example, when a participant's data distribution undergoes a minor change, only the model parameters related to that participant are updated. Without impacting normal business operations, pruning strategies are adjusted promptly to ensure model accuracy and adaptability.

[0109] Compared with traditional technologies, traditional credit assessment systems rely on high-computing power servers to deploy complex models and cannot adapt to low-configuration devices (such as mobile terminals and IoT terminals), resulting in poor real-time performance and high deployment costs. Traditional model compression technologies mostly use static pruning or single quantization strategies, which make it difficult to balance accuracy and efficiency. Especially in scenarios with large differences in data distribution, the model generalization ability is significantly reduced. In addition, traditional API interface designs lack scenario-based adaptability. Business needs such as credit approval and post-loan monitoring need to rely on multiple independent systems, resulting in resource redundancy and response delays. This solution systematically solves the core problems of real-time computing on low-configuration devices, dynamic data distribution adaptation, and scenario-based service integration through lightweight model compression technology (channel pruning and 8-bit integer quantization) and adaptive pruning strategies, achieving efficient edge deployment of credit assessment models and seamless integration with business scenarios.

[0110] Through the above-mentioned technical solutions, this application achieves efficient inference and scenario-based service collaboration for dynamic credit scoring on low-configuration devices. The model compression unit uses channel pruning and 8-bit quantization to reduce the size of the global credit assessment model to less than 30% of its original size. Combined with ARM architecture instruction set optimization, this enables the model to achieve millisecond-level real-time response on mobile devices and IoT devices. An adaptive pruning strategy dynamically adjusts the pruning rate based on the data distribution of participating parties, reducing computational complexity while ensuring manageable loss in model accuracy. The scenario-based API unit uses lightweight edge computing nodes to encapsulate business logic such as credit approval and post-loan monitoring into standardized interfaces, supporting on-demand invocation in multiple scenarios. Combined with a linkage mechanism between the time series data pipeline and anomaly detection models, it enables dynamic updates of risk warnings and credit scores. Each module operates collaboratively through online pruning strategy updates and incremental deployment mechanisms, forming a closed-loop optimized technical chain from model compression to edge inference to business scenario services. This provides a low-power, highly reliable end-to-end solution for inclusive finance and supply chain finance.

[0111] like Figure 2 As shown, the Federated Learning Small and Micro Enterprise Credit Profiling System also includes:

[0112] A blockchain audit module, which records data contributions and model versions of participants through a hash chain and opens a third-party audit interface for regulators to verify data compliance.

[0113] The generation formula of the hash chain is: n =SHA-256(Hash n-1 ||Data Cube ID||Timestamp)

[0114] Among them, Hash n is the nth hash value, Hash n-1 Is the previous hash value, SHA-256 is a hash function algorithm used to n-1 ||data cube ID||timestamp) is calculated to generate a hash value of fixed length (256 bits); the data cube ID represents the data block; the timestamp represents the time information when the data block was generated, and "||" represents the splicing operation.

[0115] The data authorization framework establishes user authorization mechanisms (such as blockchain evidence storage) to ensure legal and compliant data use. It promotes the application standards of federated learning in the financial sector (such as data interface specifications and model evaluation indicators) to reduce collaborative friction. The blockchain audit module records the data contributions and model versions of participants through a hash chain, providing traceability and tamper-proof guarantees for data use. Its open third-party audit interface facilitates regulators to verify the compliance of data use and ensure that the entire system operates within the legal and compliant track. The data authorization framework establishes user authorization mechanisms, such as blockchain evidence storage, to ensure the legality of data use at the source. At the same time, it promotes the application standards of federated learning in the financial sector, including data interface specifications and model evaluation indicators, effectively reducing collaborative friction between different participants.

[0116] Small and micro enterprises (SMEs) have long faced pain points in data utilization, including fragmented data, difficulty protecting privacy, and complex compliance verification. This system provides an effective approach to addressing these issues. Using federated learning technology, participants can train models without sharing raw data, protecting data privacy and improving data utilization efficiency. The system can create precise profiles of SMEs, helping financial institutions more accurately assess their credit risk. This allows for more tailored financial services and promotes their healthy development.

[0117] From a mathematical point of view, hash functions are one-way and deterministic. Given the same input, the same output will always be obtained, but the input cannot be deduced from the output. This means that once the data block is changed, the hash function will be n will be completely different. Moreover, due to the n Depends on Hash n-1Modifying any data block will change all subsequent hash values, thus disrupting the entire hash chain. This property ensures that the hash chain cannot be tampered with. The hash chain accurately records the data contributions and model versions of each participant. Each participant's data upload and model update generates a new data block, calculates the corresponding hash value, and adds it to the hash chain. By viewing the hash chain, we can clearly understand the contributions of each participant at different stages.

[0118] Compared to traditional data auditing methods, Hash Chain offers significant efficiency advantages. Traditional auditing methods require comparing and verifying large amounts of raw data, a tedious and error-prone process. Hash Chain, on the other hand, only needs to verify the continuity and correctness of hash values, significantly reducing the workload and time cost of auditing.

[0119] The third-party audit interface is a crucial component of the blockchain audit module. Its design standards must balance data accessibility and privacy protection. The interface should adhere to a unified data format and communication protocol to ensure that regulators can easily access the required information. Furthermore, encryption technology should be used to encrypt data transmission and storage to prevent theft or tampering during transmission.

[0120] In terms of data privacy protection mechanisms, the interface uses technologies such as zero-knowledge proof, allowing regulators to verify the compliance of data usage without obtaining the original data. Regulators can confirm the source, usage, and authorization of data by verifying hash chains and smart contracts.

[0121] Key milestones in the regulatory verification process include data access request approval, real-time monitoring of data usage, and post-audit. During the data access request phase, regulators will review the request against relevant laws and policies to ensure its legitimacy. During data usage, real-time monitoring interfaces are used to obtain data usage logs and statistics to promptly identify anomalies. Post-audits comprehensively review the entire data usage process to ensure compliance with data usage requirements.

[0122] Blockchain evidence plays a crucial role in proving compliance. It records the authorization, use, and transfer of data, forming an immutable chain of evidence. Regulators can verify the compliance of data use by querying blockchain evidence, providing strong regulatory support.

[0123] Compared with traditional technologies, traditional credit assessment systems rely on centralized databases to store data contribution records and model version information, which poses a risk of data tampering and a cumbersome audit process. When collaborating across institutions, traditional methods lack unified data interface specifications and compliance verification mechanisms, resulting in opaque data authorization processes and increased collaboration friction. For example, small and micro-enterprise credit data is scattered across different institutions such as banks, tax authorities, and supply chains. Traditional models make secure sharing and dynamic auditing difficult, requiring regulators to expend significant resources verifying data compliance and unable to trace data flow paths in real time. This solution, through a blockchain audit module and hash chain technology, systematically addresses core issues such as lack of data traceability, inefficient compliance verification, and inconsistent cross-institutional collaboration standards, providing a credible and transparent technical foundation for credit assessment within the federated learning framework.

[0124] Through the above technical solutions, this application realizes the trusted management and efficient collaboration of the entire life cycle of small and micro-enterprise credit data. The blockchain audit module uses the tamper-proof characteristics of the hash chain to fully record the data contribution and model version iteration trajectory of the participants, forming a traceable global data flow map; the open third-party audit interface combined with the smart contract verification mechanism enables regulators to verify the compliance of data use in real time without relying on cumbersome offline audit processes. The data authorization framework solidifies user authorization records through blockchain evidence storage, ensuring the legality of data use while promoting the implementation of standardized interfaces and evaluation indicators for federated learning in the financial field, reducing the hidden costs of multi-institutional collaboration. Through the dynamic update of the hash chain and the coordinated call of the standardized interface, each module achieves a seamless connection from data collection, model training to regulatory audit, providing safe, transparent and scalable technical support for the credit profiling of small and micro enterprises.

[0125] like Figure 2 As shown, the federated learning small and micro enterprise credit profiling system also includes a cold start compensation module, which is used to handle the situation where newly registered enterprise data is insufficient. After performing the following operations, it is input into the federated computing module for the next step:

[0126] Obtain historical credit records of the legal representative's affiliated companies;

[0127] Extract benchmark operating parameters for the industry to which the enterprise belongs;

[0128] The industry benchmark parameters and the legal representative's credit record are combined through the transfer learning algorithm;

[0129] Before outputting the initial credit score, a confidence correction term determined by the coefficient of variation of the electronic invoice amount is added to perform the standardization preprocessing.

[0130] If there is no data for a new enterprise, the historical credit records of the legal representative's affiliated enterprises will be migrated (weight 70%);

[0131] Supplementing industry benchmark parameters (weight 30%), the initial credit scoring formula is:

[0132] Initial score = 0.7 × correlation score + 0.3 × industry benchmark;

[0133] Confidence correction item: based on the coefficient of variation of the electronic invoice amount (confidence +10% when <0.3).

[0134] The cold start compensation module first deeply explores the historical credit records of the legal representative's affiliated companies. The credit performance of the companies the legal representative has been involved in can, to a certain extent, reflect their management capabilities and credit propensity. By carefully combing through these historical credit records, the system obtains credit-related data.

[0135] At the same time, the system accurately extracts benchmark operating parameters for the company's industry. Different industries have their own unique operating models, risk profiles, and development patterns. For example, the manufacturing industry may focus more on parameters such as production equipment utilization and raw material procurement stability, while the service industry focuses on customer satisfaction and employee service efficiency. Extracting these industry benchmarks helps new companies to be appropriately positioned and evaluated within the broader industry context.

[0136] After obtaining the credit history of the legal representative's affiliated companies and industry benchmark parameters, the system uses a transfer learning algorithm to achieve a deep integration of the two. Before outputting the initial credit score, the system also adds a confidence correction factor determined by the coefficient of variation of electronic invoice amounts. The coefficient of variation of electronic invoice amounts is a key indicator of a company's operational stability and financial health. When this coefficient is less than 0.3, it indicates that the company's operations are relatively stable and the credibility of its financial data is high. In this case, the system adds 10% confidence to the initial credit score. The introduction of this correction factor further enhances the scientific and rationality of credit scores, making them more realistically reflect the credit status of companies.

[0137] Specifically, when new companies have no data, the system integrates the credit records of affiliated companies and industry benchmark parameters according to specific weights. The historical credit record of the legal representative's affiliated companies accounts for 70% of the weighting. This is because the legal representative's business practices and credit awareness have a profound impact on the company's development, and their past credit performance is a crucial basis for predicting the creditworthiness of new companies. Industry benchmark parameters, meanwhile, account for 30% of the weighting, providing new companies with industry-level references to ensure that the assessment results are consistent with industry realities.

[0138] Compared with traditional technologies, traditional methods can neither effectively integrate the legal representative's historical credit record with industry benchmark parameters, nor can they dynamically correct the credibility of the company's initial score. Data sparsity often leads to significant model prediction deviations (such as a new company loan rejection rate exceeding 50%). In addition, traditional static scoring models do not take into account differences in industry operating characteristics (such as the same parameter weights for manufacturing and service industries), and lack a quantitative correction mechanism for financial volatility (such as the coefficient of variation of electronic invoice amounts). The evaluation results are difficult to truly reflect the stability of the company's operations. This solution systematically solves the credit modeling problem in data-sparse scenarios through a cold start compensation module and transfer learning technology, and realizes multi-dimensional dynamic adaptation of new enterprise credit assessments.

[0139] Through the above-mentioned technical solution, this application achieves cross-domain data fusion and dynamic credibility correction for credit assessments of newly registered small and micro enterprises. The cold start compensation module deeply mines the historical credit records of the legal representative's affiliated enterprises and, combined with a transfer learning algorithm, dynamically combines industry benchmark parameters with individual characteristics, ensuring that the initial score reflects both the legal representative's creditworthiness and industry operating rules. Confidence correction quantifies the company's financial stability based on the coefficient of variation of electronic invoice amounts. When data fluctuations are minimal, the score's credibility is automatically improved, effectively mitigating assessment bias caused by data sparsity. The federated computing module deeply integrates the legal representative's historical data (weighted 70%) with industry benchmarks (weighted 30%) through weight allocation and feature alignment. The blockchain audit module also records the data contribution path, ensuring a transparent and traceable assessment process. Each module is collaboratively optimized through a closed-loop feedback mechanism, from data collection and feature fusion to dynamic correction, forming a credit assessment chain covering the entire life cycle of new enterprises, providing financial institutions with timely, accurate, and compliant risk management decision support.

[0140] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. The Federated Learning Small and Micro Enterprise Credit Profiling System is characterized by: include: A data access module is used to obtain heterogeneous data from multiple data sources and perform standardized preprocessing on the heterogeneous data; The federated computing module is used to select a horizontal or vertical federated learning mode according to the data type of the heterogeneous data, schedule the local models of multiple participants for training, and generate a global credit assessment model through parameter aggregation, wherein: If the horizontal federated learning model is selected, cross-institutional data with the same characteristics but different samples will be jointly trained, and aggregation weights will be dynamically allocated according to the data integrity indicators and feature redundancy of the participants. If the longitudinal federated learning mode is selected, cross-modal data with different features of the same sample are aligned and trained, and sample space alignment is achieved through a dynamic time alignment loss function; A credit assessment module, configured to output a dynamic credit score based on the global credit assessment model and integrate transfer learning, and generate an interpretable report including feature contribution; The application service module is used to deploy the dynamic credit score to low-configuration devices and provide an application programming interface (API) for preset scenarios.

2. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The heterogeneous data includes bank transaction data, tax invoice data, supply chain evaluation data and IoT sensor data; The preprocessing includes semantic analysis of unstructured data and missing value filling and normalization of structured data.

3. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The data access module also includes: The data quality verification unit is used to detect the missing rate and noise level in the heterogeneous data. If the missing rate exceeds a preset missing threshold, it triggers the data source to re-upload or generate virtual filling data.

4. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The federated computing module includes: The multi-head attention feature extraction unit is used in the horizontal federated learning model to extract the feature vectors of the cross-institutional data through the Transformer architecture, generate the data integrity index and feature redundancy, and calculate the cross-institutional attention weight. The calculation formula of the attention weight is: Among them, Q i , K j are query vectors and key vectors of different modalities, d k is the preset dimension; A time-space alignment unit is used in the longitudinal federated learning mode to generate pseudo data aligned with the high-frequency data timestamp of the cross-modal data by sliding window interpolation on the low-frequency data of the cross-modal data, and constrain the feature space consistency by the dynamic time alignment loss function, where the loss function is: in, and are the feature vectors of the high-frequency data and the interpolated low-frequency data respectively.

5. The federated learning small and micro enterprise credit profiling system according to claim 4 is characterized in that: The federated computing module further includes: an anomaly detection subunit, configured to detect abnormal points of data fluctuation during the interpolation process by the time-space alignment unit and correct the abnormal points by linear interpolation; A layered encryption unit is used to decrypt and process highly sensitive data in a trusted execution environment (TEE) and transmit gradient parameters using a homomorphic encryption algorithm, where: The highly sensitive data include corporate ID numbers and bank account numbers; The homomorphic encryption algorithm is the Paillier algorithm, and the time consumed by a single encryption is less than a preset time threshold.

6. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The credit assessment module includes: The transfer learning unit pre-trains a basic model based on pre-set industry general data, and transfers the knowledge of the basic model to small and micro enterprises with sparse data through federated learning to generate virtual feature vectors and macroeconomic sensitivity factors. The loss function of the transfer learning is: Among them, λ is the preset migration weight coefficient, and KL divergence is used to constrain the feature distribution difference between the source domain and the target domain; The macroeconomic factor embedding unit is used to input the Purchasing Managers' Index (PMI) and industry prosperity into the gated network and dynamically adjust the feature weights. The weight update formula is: w t =sigmoid(W·[h t ;ΔPMI t ]) Among them, ΔPMI t is the rate of change of PMI index, h t is the current feature vector; The explainability generation unit is used to analyze feature contributions through SHAP values and generate an explainability report containing key influencing factors.

7. The federated learning small and micro enterprise credit profiling system according to claim 6 is characterized in that: The macroeconomic factor embedding unit further includes: The industry classification subunit is used to classify and code small and micro enterprises according to preset industries and load the prosperity indicators of the corresponding industries. The classification codes include retail, manufacturing and logistics.

8. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The application service module includes: A model compression unit, configured to perform channel pruning and 8-bit integer quantization on the global credit assessment model to compress the model and adapt it to ARM architecture edge devices; A scenario-based API unit is used to output dynamic credit scores based on preset scenarios, which are deployed to the low-configuration device through lightweight edge computing nodes. The preset scenarios include credit approval, post-loan monitoring and supply chain finance. The adaptive pruning strategy subunit is used to dynamically adjust the pruning rate according to the data distribution of the participants.

9. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The system further comprises: A blockchain audit module, which records the data contributions and model versions of the participants through a hash chain and opens a third-party audit interface for regulators to verify data compliance. The generation formula of the hash chain is: Hash n =SHA-256(Hash n-1 ||Cube ID||Timestamp).

10. The federated learning small and micro enterprise credit profiling system according to claim 1 is characterized in that: The system also includes a cold start compensation module for processing insufficient data of newly registered enterprises, and then inputting the following operations into the federated calculation module: Obtain historical credit records of the legal representative's affiliated companies; Extracting benchmark operating parameters for the industry to which the enterprise belongs; Using a transfer learning algorithm, the industry benchmark parameters are combined with the credit record of the legal representative for feature splicing; Before outputting the initial credit score, a confidence correction term determined by the coefficient of variation of the electronic invoice amount is added.

Citation Information

Cited By

  • Credit portrait construction method and system based on dynamic weight and incremental learning

    CN121213232A

  • Big data-based information technology detection system with security protection

    CN121567489A

  • Small and micro enterprise joint risk portrait and abnormal behavior identification system

    CN121961261A