A data processing method, apparatus, device, medium, and product

By leveraging blockchain technology and a federated learning architecture, the challenges of data sharing and synchronization among multiple participants are resolved, enabling rapid and accurate global model training. This improves the efficiency and security of bank credit assessment and is applicable to bank credit assessment models.

CN122490575APending Publication Date: 2026-07-31AGRICULTURAL DEVELOPMENT BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AGRICULTURAL DEVELOPMENT BANK OF CHINA
Filing Date
2026-05-06
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In multi-party data sharing scenarios, how to quickly achieve data sharing and synchronization, train an accurate global model, and at the same time ensure data security and privacy, especially when banks assess customer credit, traditional methods suffer from difficulties in data sharing, low assessment efficiency, and data silos.

Method used

A data processing approach based on blockchain technology and federated learning architecture is adopted. The blockchain layer performs data alignment and privacy protection set intersection protocol, the federated learning layer performs local model training and parameter aggregation, and the application layer performs model application, thereby realizing data sharing and synchronization among multiple participants.

Benefits of technology

It enables rapid sharing and synchronization of data among multiple participants, trains an accurate global model, improves data processing efficiency, and enhances the accuracy and security of credit assessment while protecting data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490575A_ABST
    Figure CN122490575A_ABST
Patent Text Reader

Abstract

This invention discloses a data processing method, apparatus, device, medium, and product. Based on a federated learning architecture, which includes a blockchain layer, a federated learning layer, and an application layer, the method includes: the blockchain layer aligning the private datasets of each participant using a preset cryptographic protocol to obtain a private data intersection, which is then sent to the federated learning layer; the federated learning layer instructing each participant to interact with the blockchain layer based on the private data intersection to perform local model training to obtain local model parameters, and instructing the blockchain layer to aggregate the local model parameters to obtain global model parameters; the global model parameters are then fed back to each participant and the application layer; the application layer obtains the global model based on the global model parameters and responds to data processing requests by performing data processing based on the global prediction model. This invention can quickly achieve data sharing and synchronization among multiple participants, train an accurate global model, and improve data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of blockchain technology and big data, and in particular to a data processing method, apparatus, device, medium and product. Background Technology

[0002] With the continuous development of blockchain and big data technologies, multi-party data sharing for better data processing has become widely used, especially for banks. When facing massive amounts of personal consumer credit and micro-enterprise loans, the traditional "5C principle" (character, ability, capital, collateral, and conditions) is highly subjective and difficult to quantify. Banks urgently need multi-party data sharing models to accurately identify risks such as multiple borrowing, thereby reducing non-performing loan rates.

[0003] Therefore, in data sharing scenarios involving multiple parties, the urgent problems to be solved are how to quickly achieve data sharing and synchronization among multiple parties to train an accurate global model, and how to ensure data security and protect the data privacy of each party. Summary of the Invention

[0004] This invention provides a data processing method, apparatus, device, medium, and product to quickly achieve data sharing and synchronization among multiple participants, train an accurate global model, improve data processing efficiency, and ensure data security.

[0005] According to one aspect of the present invention, a data processing method is provided, based on a federated learning architecture, the federated learning architecture including a blockchain layer, a federated learning layer, and an application layer, the method comprising: The blockchain layer aligns the private datasets of each participant based on a pre-defined cryptographic protocol, obtains the intersection of private data, and sends it to the federated learning layer. The federated learning layer instructs each participant to interact with the blockchain layer based on the intersection of private data, to train the local model to obtain local model parameters, and instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, and then feeds back the global model parameters to each participant and the application layer. The application layer obtains the global model based on the global model parameters and responds to data processing requests by performing data processing based on the global prediction model.

[0006] According to another aspect of the present invention, a data processing apparatus is provided, implemented based on a federated learning architecture, the federated learning architecture including a blockchain layer, a federated learning layer, and an application layer, the apparatus comprising: The alignment module is used to instruct the blockchain layer to perform data alignment on the private datasets of each participant based on a preset cryptographic protocol, obtain the intersection of private data, and send it to the federated learning layer. The aggregation module is used to instruct the federated learning layer to interact with the blockchain layer based on the intersection of private data, to train the local model to obtain local model parameters, and to instruct the blockchain layer to aggregate the local model parameters to obtain global model parameters, and to feed back the global model parameters to each participant and the application layer. The processing module is used to instruct the application layer to obtain the global model based on the global model parameters, and in response to data processing requests, to perform data processing based on the global prediction model.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data processing method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is also provided, the computer program product including a computer program that, when executed by a processor, implements the data processing method of any embodiment of the present invention.

[0010] The technical solution of this invention involves a blockchain layer that aligns the private datasets of each participant based on a preset cryptographic protocol, obtaining a private data intersection, which is then sent to the federated learning layer. The federated learning layer instructs each participant to interact with the blockchain layer based on the private data intersection, performing local model training to obtain local model parameters. It then instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, which are fed back to each participant and the application layer. The application layer obtains the global model based on the global model parameters and responds to data processing requests, performing data processing based on the global prediction model. This approach enables rapid data sharing and synchronization among multiple participants, training an accurate global model, and improving data processing efficiency.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a data processing method provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of a data processing device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," "target," "candidate," and "alternative," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the invention described herein can be practiced in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The acquisition, storage, use, and processing of data in the technical solutions of this application comply with relevant laws and regulations.

[0016] It should be noted that the user information collected in this invention is information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, and necessary confidentiality measures have been taken. This process does not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or reject automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making process. In other words, the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of this data comply with the relevant laws, regulations, and standards of the relevant regions.

[0017] It should be noted that banks are institutions that manage risk. When responding to loan requests from businesses or individuals, traditional manual approval processes are slow (taking days or even weeks), failing to meet consumers' demand for fast loan processing. The common processes and problems in the banking industry for assessing customer credit are as follows: (1) The supporting materials are submitted by the customer, which can easily lead to the problem of single data dimension. In addition, many banks still rely too much on the central bank's credit information and internal transaction data. For "credit blank" or small and micro enterprises with light assets who lack credit history, it is difficult to fully capture their true credit level due to the lack of supplementary data such as e-commerce consumption and supply chain, which can easily lead to the phenomenon of "high-quality customers being wrongly rejected".

[0018] (2) Banks use manual methods or simple models to collect and verify customer information. This is because customer data is widely distributed across different institutions and various systems within the bank. Traditional manual cross-platform collection and verification methods are not only time-consuming and labor-intensive, but also prone to data loss or inconsistency, affecting the efficiency and accuracy of the assessment.

[0019] (3) Assess customer credit. Credit risk models are mostly trained based on historical data, and the update cycle is relatively long. In the face of the emergence of new business formats or sudden events, the adaptability and prediction accuracy of the model will drop rapidly, making it difficult to capture rapidly changing market risks in a timely manner.

[0020] To address the aforementioned issues, this invention provides a model training and data processing scheme based on blockchain technology and federated learning. This scheme can train an accurate and effective global model for data processing while protecting the security of user data from all participating parties. Specifically, it can be applied to building credit assessment models. By utilizing customer credit assessment models for automated credit evaluation, it effectively solves the problems of low efficiency, high cost, and data silos in the bank customer data assessment process. It can reduce approval time to a few minutes and extend financial services to micro and small enterprises that traditional banks have not served, thus achieving inclusive finance. Specific implementation methods will be described in detail in subsequent embodiments.

[0021] Example 1 Figure 1 This is a flowchart of a data processing method provided by an embodiment of the present invention. This embodiment is applicable to situations where a global model is obtained by training a model based on blockchain technology and federated learning strategy for data processing, and is particularly applicable to situations where a customer credit rating assessment model is trained for customer credit rating assessment. This method can be executed by a data processing device, which can be implemented in hardware and / or software. The data processing device can be configured in an electronic device, which can be configured with a data processing system based on a federated learning architecture, which includes a blockchain layer, a federated learning layer, and an application layer.

[0022] The blockchain layer comprises consensus nodes and ordering nodes. Each consensus node establishes connections with at least two participants; each participant establishes connections with two designated consensus nodes. Consensus nodes are responsible for verifying the accuracy of transactions and maintaining the public ledger, employing PBFT (Practical Byzantine Fault Tolerance) consensus as their mechanism. Ordering nodes are responsible for ordering transactions to generate blocks. Each block includes a header and a body; the header contains a pointer to the previous block, and the body contains the verified transaction information. The blockchain layer primarily participates in the preparation work for federated learning and assists in the aggregation of relevant parameters.

[0023] The federated learning layer can support different participants to carry out collaborative model training when there is overlap in samples and complementary feature dimensions. Specifically, each participant independently performs forward propagation and gradient calculation of the model locally based on its private feature space. The generated intermediate representations and gradient parameters are encrypted using privacy protection technologies such as homomorphic encryption or secure multi-party computation before being uploaded to the blockchain layer for aggregation.

[0024] After the blockchain layer aggregation is completed, the global model parameters are securely distributed to all participants for the next round of local training and joint optimization, thereby achieving joint modeling and optimization across the feature space without exposing the original data. This federated learning layer design achieves secure collaboration and efficient fusion of features from multiple parties, effectively balancing model performance and data privacy protection.

[0025] The application layer comprises the infrastructure, including edge nodes. Its primary function is to support model invocation through simplified application programming interfaces (APIs), enabling access to and management of the data and functions provided by the blockchain layer. This layer allows participants to establish connections based on the public address information of consensus nodes, thereby effectively coordinating, exchanging, and sharing machine learning models, parameter updates, and related metadata. The application layer also handles tasks such as client requests, task allocation, security verification, and data encryption and decryption, providing the necessary support and platform for the operation of the entire federated learning system. Through the design and implementation of the application layer, clients can easily utilize federated learning technology for model training, evaluation, and inference, while ensuring full protection of data privacy and security, providing a reliable foundation for intelligent decision-making and prediction.

[0026] like Figure 1 As shown, the data processing method includes: S101, the blockchain layer, based on a preset cryptographic protocol, aligns the private datasets of each participant to obtain the intersection of private data and sends it to the federated learning layer.

[0027] The pre-defined cryptographic protocol refers to a protocol that enables data alignment between at least two participating parties. The pre-defined cryptographic protocol is an optimized Privacy Set Intersection (PSI) protocol. Data alignment can be ID (identification) alignment. Participating parties can be banks, e-commerce platforms, or telecom operators. The private dataset can be a user ID dataset private to each participating party. The private data intersection refers to a user ID dataset shared or jointly owned by all participating parties.

[0028] It should be noted that existing PSI protocols can only align data between two participants and do not support multi-party participation. This invention proposes a blockchain-based PSI protocol that supports data alignment, such as ID alignment, between two or more participants. The multi-party PSI protocol employs a Diffie-Hellman commutative encryption algorithm (i.e., the Diffie-Hellman key exchange algorithm). Its core idea is to map each participant's private ID to elements on a cyclic group and utilize the commutativity of exponential operations to achieve bidirectional encrypted comparison of IDs. Introducing session random numbers into the commutative encryption algorithm ensures that encryption results across rounds are not correlated, further resisting replay and chaining attacks. This algorithm achieves secure alignment of IDs across multiple parties while effectively guaranteeing the privacy and robustness of the protocol. This allows multiple participants to ultimately align the same ID elements through multiple rounds of encryption and exchange without exposing their original IDs.

[0029] It should be noted that, to protect data security, each participant's private dataset is encrypted and specially processed before being sent to the blockchain layer. Therefore, the consensus node only receives a string of characters. Traditional federated learning uses a homomorphic encryption key determined by a third party. However, in the federated learning architecture of this invention, the blockchain layer generates the key, the consensus node completes the first alignment task, and the sorting node completes the second alignment task, obtaining the final intersection of private data for the federated learning layer to train the model.

[0030] It should be noted that this invention proposes a data collection scheme based on blockchain technology and creatively proposes a privacy-preserving set intersection protocol in the blockchain. This ensures that each participant can only know the index information of the intersection sample without disclosing the original data, thereby achieving privacy protection at the feature level, significantly reducing the cost of manual intervention and improving data credibility.

[0031] Optionally, the blockchain layer performs data alignment on the private datasets of each participant based on a preset cryptographic protocol to obtain a private data intersection. This includes: the blockchain layer instructing consensus nodes to perform a first alignment task on the private datasets of each participant based on an optimized privacy-preserving set intersection protocol, to obtain a first private data intersection; and, based on the first private data intersection, instructing sorting nodes to perform a second alignment task on the private datasets of each participant, to obtain a final private data intersection.

[0032] Optionally, the blockchain layer, based on a privacy-preserving set intersection protocol, instructs consensus nodes to perform the first alignment task of the private datasets of each participant to obtain the first private data intersection. This includes: the blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to obtain encrypted private data sent by the participants with whom they have established connections, to obtain encrypted private datasets; the consensus nodes select private data with the same encryption value from the encrypted private datasets to complete the first alignment task of the private datasets of each participant, thus obtaining the first private data intersection.

[0033] Here, encrypted private data refers to the data uploaded by each participant after encrypting their local dataset. The first private data intersection refers to the intersection of the data from all participants obtained after the first alignment task.

[0034] Optionally, each participant will randomly establish connections with multiple consensus nodes through their public addresses, and each consensus node will also establish connections with two or more participants. It is with the participating parties One of the connected consensus nodes. Participants Use exchangeable cryptographic public keys For its ID set hash value Encryption is performed to generate a set of private IDs. That is, a private dataset, which is then sent to the consensus node. .

[0035] For example, the private datasets of each participant This can be expressed by the following formula: =({ExEnc( , ),ExEnc( , ),ExEnc( , ),...}, ); Among them, ExEnc( , This indicates that the sample ID hash value is generated using the Diffie-Hellman-based commutative cryptographic algorithm CE-DH. With key The encrypted value obtained after encryption. This represents the nth participant.

[0036] For example, based on the connection relationship between consensus nodes and participants, consensus nodes are denoted as... The set of participants in the connection is ,in It is the index of the i-th participant in the set P. Consensus Node from A set of private sample IDs, i.e., a private dataset, was received, represented as follows: ={ , , ...}, the private sample ID uses the same key. Encrypted. Then, the consensus node... Perform the first ID alignment, i.e., the first alignment task, from... Select the identical parts from the encrypted private sample IDs, compare all encrypted elements, and choose the one that matches. The same ExEnc encrypted value forms the first ID alignment set. Each consensus node Unable to decrypt from The encrypted private ID set received there.

[0037] Optionally, the blockchain layer, based on the first private data intersection, instructs the sorting nodes to perform a second alignment task on the private datasets of each participant to obtain the final private data intersection. This includes: the blockchain layer instructing the consensus nodes to generate message verification codes based on the first private data intersection and the exchangeable encrypted public keys, and to generate transactions and send them to the sorting nodes based on the message verification codes, the exchangeable encrypted public keys, the first private data intersection, and the corresponding set of participants; the sorting nodes select private data with the same encryption value from the transaction sets generated by the transactions sent by each consensus node to complete the second alignment task on the private datasets of each participant and obtain the final private data intersection.

[0038] For example, after obtaining the first ID alignment set After that, each consensus node A key is required. and alignment results This is used to generate a hash-based message authentication code (HMAC). Then, the consensus nodes... Package the message verification code, the exchangeable encrypted public key, the intersection of the first private data and its corresponding set of participants into a transaction set. And send it to the sorting node. This can be expressed by the following formula: ; Optionally, sorting nodes can share received transactions with each other and adopt the corresponding consensus node. key The authenticity of each transaction is verified, followed by a second ID alignment algorithm, i.e., the second alignment task. The sorted nodes are updated via commutative cryptography. and To obtain the intersection, specifically, the sorted nodes can compare all encrypted elements in the transaction set and select the... The same ExEnc encrypted value forms a second ID alignment set. After calculating the intersection, sort the nodes and select the current sequence number to retain in the new round. Delete elements in each u that do not have the current sequence number. Then, the private intersection U will be updated and inserted. The sequence numbers are then rearranged to obtain the final private data intersection.

[0039] Optionally, the sorting node packages the new transaction containing the final private data intersection with other verified transactions into a new block and broadcasts it in the blockchain layer. Each consensus node verifies the block, checking whether the final private data intersection is a subset of its original, stored first ID alignment set. Consensus nodes reach consensus through voting, selecting a valid block as the proposal block. The voting results are computed under homomorphic encryption to ensure the privacy of the voting content. The winning consensus node broadcasts the encrypted block to other nodes in the network. After verifying the block's validity and decrypting the voting results, each node updates its local blockchain and shares the second ID alignment result, i.e., the private intersection, with connected participants, ensuring all nodes' new agreement and consistency on the blockchain.

[0040] S102. The federated learning layer instructs each participant to interact with the blockchain layer based on the intersection of private data, to train the local model to obtain local model parameters, and instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, and then feeds back the global model parameters to each participant and the application layer.

[0041] Local model parameters refer to the model parameters obtained by each participant after training the model locally using their own private dataset. All participants use the same initial model for local training.

[0042] Optionally, the blockchain layer broadcasts the final private data intersection, allowing each participant to update and synchronize their local private intersection. After that, each participant can interact with the blockchain layer to train their local model and obtain local model parameters.

[0043] Optionally, the federated learning layer can determine the final global model parameters through multiple rounds of training and iteration. Specifically, let the set of consensus nodes that win first in each round be denoted as . { , ... }, where max represents the maximum number of rounds. The homomorphic encryption key pair used in the i-th round is denoted as . Round 0 is defined as the process of determining the intersection of private data based on the PSI protocol. After determining the intersection of private data, the first round of PBFT (Practical Byzantine Fault Tolerance) key distribution process is as follows: Consensus Node Randomly select a pair of homomorphic encryption keys The public key in the key is broadcast to the blockchain layer so that each participant can interact with the consensus node of the blockchain layer to train local model parameters and then combine them to obtain global model parameters.

[0044] Optionally, after obtaining the global model parameters, the global model parameters can be distributed to each participant to instruct each participant to initialize or fine-tune their local model with the global model parameters as the starting point for the next round of training.

[0045] Optionally, the federated learning layer instructs each participant to perform local model training based on the intersection of private data to obtain local model parameters. This includes: the federated learning layer instructing each participant to perform local model training based on the intersection of private data to obtain intermediate output, and encrypting the intermediate output using a homomorphic key to obtain encrypted intermediate output; each participant determines the corresponding encryption gradient, determines a random mask based on a preset random number generator and the encryption gradient, and determines a privacy gradient based on the encryption gradient and the random mask; each participant uploads the privacy gradient to the consensus node connected to it in the blockchain layer, and performs local model training based on the decrypted privacy gradient decrypted by the consensus node to obtain local model parameters.

[0046] Optionally, the federated learning layer can obtain the final local model parameters trained by each participant through multiple rounds of iteration, thereby determining the final global model parameters. Specifically, each round of iteration includes consensus node distributing homomorphic keys, each participant training the model to obtain intermediate outputs, using homomorphic keys to encrypt and obtain encrypted intermediate outputs, determining privacy gradients, interacting with consensus nodes to obtain decrypted privacy gradients, and local model training to obtain the final local model parameters trained by each participant.

[0047] For example, the i-th iteration of the federated learning layer may include the following eight steps: (1) Key preparation: Before the start of the i-th round, each consensus node will receive a homomorphic key pair; (2) Key distribution: Each consensus node will distribute the public key in the homomorphic key pair to all consensus nodes; (3) Model training: Each participant trains its own local model and then obtains the intermediate output needed to update its local model. It can be expressed by the following formula: in, The model gradient for the data labels; y represents the dimension coefficient of the data feature corresponding to the label; the superscript "data" indicates that the data has been normalized; y is the true label value of the sample; 0.5 is the model offset; It is a scaling factor used to control the range of updated values ​​and avoid gradient explosion.

[0048] (4) Exchange intermediate outputs: The participants use the homomorphic key from the i-th round to encrypt the intermediate outputs obtained after training the model, and obtain encrypted intermediate outputs. These encrypted intermediate outputs will be used in the gradient uploading and model update process. Each participant n will receive an encrypted gradient. The encryption gradient can be represented by the following formula: Where n represents the number of participants; The data feature dimension coefficients corresponding to the labels; y is the transpose of the data feature dimension coefficients corresponding to the label; y is the true label value of the sample; 0.5 is the model offset; It is a scaling factor used to control the range of updated values ​​and avoid gradient explosion; This is the local model gradient of the nth participant, used as a supplement to the final gradient.

[0049] (5) Upload privacy gradient: Each participant can use a pseudo-random number generator to generate its own encrypted gradient based on a unique seed. Random mask of the same dimension as the vector And based on the formula The privacy gradient is obtained by adding the random mask to the local encryption gradient. It is uploaded to the consensus node connected to it in the blockchain layer.

[0050] In this process, the random mask is generated and stored locally, preventing external entities from directly inferring the true gradient information. Any two participants negotiate and generate a shared set of random seeds via a secure channel, which are then used to generate a pair of random masks with opposite values. As the encrypted vectors uploaded by all participants are aggregated in the blockchain network, these canceling masks automatically cancel each other out, ensuring that the final aggregation result matches the true gradient.

[0051] (6) Returning the gradient: The consensus node uses the homomorphic key to decrypt the privacy gradient, obtains the decrypted privacy gradient, and returns it to the corresponding participant. Due to the cryptographic masking, if any blockchain node attempts to decrypt it... Unable to obtain Only the participants can calculate the gradient. ; (7) Upload test results and update the model: Each participant uses The local model is updated and tested using a test dataset. The test results are sent to the connected consensus node, which packages the test results into transactions and sends them to the ranking node. The loss function is calculated through forward propagation, and then the gradient of the loss function with respect to the model parameters is obtained through backpropagation. The model parameters are updated using the gradient descent optimization algorithm, and then the updated parameters are used for the next iteration until the model converges. At this point, all participants can obtain the final local model parameters. (8) Block generation and preparation: The block generation phase adopts practical Byzantine fault-tolerant consensus to ensure that consensus can be reached within a limited time.

[0052] Optionally, after multiple rounds of distributed training and optimization, and combined with the verification and encryption mechanisms of blockchain, the parameters of each local model are gradually updated and converged with the collaboration of various participants in the banking industry (such as banks, regulatory agencies, and financial institutions), ultimately achieving the expected prediction accuracy. When the test results reach the set target or the maximum number of iterations, the training process terminates, generating a final set of local model parameters. These parameters can be used to obtain global model parameters to construct a global prediction model, which can achieve high-precision predictions on the dataset. The global prediction model can be a model that has been cryptographically verified by banks to assess customer credit ratings.

[0053] It should be noted that in traditional federated learning model training frameworks, collaborators often share data and perform computations based on a trust model. However, during this process, collaborators may encounter various attacks, such as data poisoning attacks, model reverse engineering attacks, and malicious computing node attacks. Attackers may tamper with training data or the computation process, leading to biases in the model training results or the leakage of private information. Attacks not only harm the training results of individual collaborators but may also undermine the integrity and trustworthiness of the entire federated learning process. For example, malicious nodes may manipulate data or the training process, potentially reducing the accuracy of the global model or even preventing the training process from converging. To address this issue, this invention designs a model training framework based on blockchain and federated learning. It aims to leverage the decentralized and immutable characteristics of blockchain to ensure effective protection of the data and model training process of each participant in a multi-party environment. Blockchain provides data integrity protection, a decentralized trust mechanism, and, combined with privacy protection technologies such as homomorphic encryption, ensures secure computation during training while protecting the privacy of all participants. The solution of this invention effectively resists the attack risks in traditional federated learning schemes and ensures the security, integrity and privacy of the model training process under multi-participant collaboration.

[0054] It should be noted that this invention optimizes a customer credit assessment model based on a federated learning algorithm, enabling multiple participants to jointly train a high-quality model without sharing the original data, thereby improving the accuracy of credit assessment and reducing labor costs.

[0055] S103. The application layer obtains the global model based on the global model parameters and responds to the data processing request by performing data processing based on the global prediction model.

[0056] The global model serves as a customer credit assessment model. For banking applications, upstream and downstream enterprises in the supply chain conduct supply and demand transactions on the platform, generating key business data. The blockchain layer automatically collects this business data based on smart contracts. When a buyer lacks sufficient funds, they can click the loan function at the application layer, upload the necessary personal information, loan information, and relevant credentials, triggering subsequent federated learning model training to assess the credit limit of the enterprise customer. The bank then uses this assessment as a crucial basis for determining the loan amount.

[0057] Optionally, in response to a data processing request, data processing is performed according to a global prediction model, including: in response to a credit assessment request for a target customer, determining credit assessment data for the target customer according to the global prediction model to provide feedback on the credit assessment request.

[0058] It should be noted that, to support ID alignment among multiple participants, this invention proposes a two-stage privacy set intersection (PSI) protocol based on blockchain, which can achieve global ID alignment after each participant completes local alignment. Simultaneously, a secure logistic regression training framework based on blockchain is designed, combining mechanisms such as random key distribution, random masking, and homomorphic encryption aggregation to improve the security and robustness of model training. This scheme not only effectively supports collaborative training among multiple participants but also significantly improves execution efficiency while maintaining high accuracy, and provides a promising solution for application in bank customer credit assessment.

[0059] It should be noted that this invention constructs a distributed model trusted training architecture based on blockchain technology, which can protect data privacy and improve the security of the training process. The blockchain is responsible for managing node identities, storing metadata, and recording the entire federated learning process, ensuring transparency and traceability. Simultaneously, federated learning utilizes the trusted environment of the blockchain for model training, parameter exchange, and aggregation. After training locally, each participant uploads encrypted model parameters and other information to the blockchain for other participants to verify their authenticity and integrity. Through the integration of these two aspects, the blockchain not only enhances the privacy and security of federated learning but also strengthens the system's robustness, ensuring the continuity and effectiveness of the training process. This allows for efficient and secure completion of banking transactions while protecting data security.

[0060] The technical solution of this invention involves a blockchain layer that aligns the private datasets of each participant based on a preset cryptographic protocol, obtaining a private data intersection, which is then sent to the federated learning layer. The federated learning layer instructs each participant to interact with the blockchain layer based on the private data intersection, performing local model training to obtain local model parameters. It then instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, which are fed back to each participant and the application layer. The application layer obtains the global model based on the global model parameters and responds to data processing requests, performing data processing based on the global prediction model. This approach enables rapid data sharing and synchronization among multiple participants, training an accurate global model, and improving data processing efficiency.

[0061] Example 2 Figure 2 This is a structural block diagram of a data processing device provided in an embodiment of the present invention. This embodiment is applicable to situations where a global model is obtained through model training based on blockchain technology and federated learning strategies for data processing, and is particularly applicable to situations where a customer credit rating assessment model is trained for customer credit rating assessment. The data processing device provided by the present invention can execute the data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method. The data processing device can be implemented in hardware and / or software and configured in an electronic device with data processing capabilities. It is implemented based on a federated learning architecture, which includes a blockchain layer, a federated learning layer, and an application layer, such as... Figure 2 As shown, the data processing device may specifically include: Alignment module 201 is used to instruct the blockchain layer to perform data alignment on the private datasets of each participant based on a preset cryptographic protocol, obtain the intersection of private data, and send it to the federated learning layer. The aggregation module 202 is used to instruct the federated learning layer to instruct each participant to interact with the blockchain layer based on the intersection of private data, to perform local model training to obtain local model parameters, and to instruct the blockchain layer to aggregate the local model parameters to obtain global model parameters, and to feed back the global model parameters to each participant and the application layer. The processing module 203 is used to instruct the application layer to obtain the global model based on the global model parameters, and in response to the data processing request, to perform data processing based on the global prediction model.

[0062] The technical solution of this invention involves a blockchain layer that aligns the private datasets of each participant based on a preset cryptographic protocol, obtaining a private data intersection, which is then sent to the federated learning layer. The federated learning layer instructs each participant to interact with the blockchain layer based on the private data intersection, performing local model training to obtain local model parameters. It then instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, which are fed back to each participant and the application layer. The application layer obtains the global model based on the global model parameters and responds to data processing requests, performing data processing based on the global prediction model. This approach enables rapid data sharing and synchronization among multiple participants, training an accurate global model, and improving data processing efficiency.

[0063] Furthermore, the blockchain layer includes consensus nodes and ordering nodes; each consensus node establishes connections with at least two participants; each participant establishes connections with two designated consensus nodes; the preset cryptographic protocol is an optimized privacy-preserving set intersection protocol. Alignment module 201 is specifically used for: The blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to perform the first alignment task of the private datasets of each participant, thereby obtaining the first private data intersection. Based on the first private data intersection, the sorting node is instructed to perform a second alignment task on the private datasets of each participant to obtain the final private data intersection.

[0064] Furthermore, the alignment module 201 is also used for: The blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to obtain encrypted private data sent by the participants with whom they have established connections, thus obtaining an encrypted private dataset. Consensus nodes select private data with the same encryption value from the encrypted private dataset to complete the first alignment task of the private datasets of each participant, and obtain the first private data intersection.

[0065] Furthermore, the alignment module 201 is also used for: The blockchain layer instructs consensus nodes to generate message verification codes based on the intersection of the first private data set and the exchangeable encrypted public key, and to generate transactions and send them to the sorting nodes based on the message verification codes, the exchangeable encrypted public key, the intersection of the first private data set and the corresponding set of participants. The sorting node selects private data with the same encryption value from the transaction set generated by the transactions sent by each consensus node to complete the second alignment task of the private datasets of each participant and obtain the final private data intersection.

[0066] Furthermore, the aggregation module 202 is specifically used for: The federated learning layer instructs each participant to train a local model based on the intersection of private data, obtains intermediate output, and encrypts the intermediate output using a homomorphic key to obtain encrypted intermediate output. Each participant determines its corresponding encryption gradient, determines a random mask based on a preset random number generator and encryption gradient, and determines a privacy gradient based on the encryption gradient and random mask. Each participant uploads its privacy gradient to the consensus node connected to it in the blockchain layer, and trains the local model based on the decrypted privacy gradient after it has been decrypted by the consensus node, thus obtaining the local model parameters.

[0067] Furthermore, the global model is a customer credit assessment model; Processing module 203 is specifically used for: In response to a credit assessment request from a target customer, the system determines the target customer's credit assessment data based on a global prediction model to provide feedback on the credit assessment request.

[0068] Example 3 Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Figure 3 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0069] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or loaded from storage unit 18 into the random access memory 13. The random access memory 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output interface 15 is also connected to the bus 14.

[0070] Multiple components in electronic device 10 are connected to input / output 15, including: input unit 16, such as a keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as a disk, optical disk, etc.; and communication unit 19, such as a network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0071] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data processing methods.

[0072] In some embodiments, the data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method by any other suitable means (e.g., by means of firmware).

[0073] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), system-on-a-chip (SoCs), complex programmable logic devices (PLCs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0074] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0075] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (e.g., voice input, speech input, or tactile input).

[0077] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0078] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system to address the shortcomings of traditional physical hosts and virtual reality services, such as high management difficulty and weak business scalability.

[0079] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the data processing method of any embodiment of the present invention.

[0080] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0081] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0082] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized in that, Based on a federated learning architecture, which includes a blockchain layer, a federated learning layer, and an application layer, the method includes: The blockchain layer aligns the private datasets of each participant based on a pre-defined cryptographic protocol, obtains the intersection of private data, and sends it to the federated learning layer. The federated learning layer instructs each participant to interact with the blockchain layer based on the intersection of private data, to train the local model to obtain local model parameters, and instructs the blockchain layer to aggregate the local model parameters to obtain global model parameters, and then feeds back the global model parameters to each participant and the application layer. The application layer obtains the global model based on the global model parameters and responds to data processing requests by performing data processing based on the global prediction model.

2. The method according to claim 1, characterized in that, in, The blockchain layer includes consensus nodes and ordering nodes; each consensus node establishes connections with at least two participants; each participant establishes connections with two designated consensus nodes; the preset cryptographic protocol is an optimized privacy-preserving set intersection protocol; Correspondingly, the blockchain layer, based on a pre-defined cryptographic protocol, aligns the private datasets of each participant to obtain the intersection of private data, including: The blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to perform the first alignment task on the private datasets of each participant, thereby obtaining the first private data intersection. Based on the first private data intersection, the sorting node is instructed to perform a second alignment task on the private datasets of each participant to obtain the final private data intersection.

3. The method according to claim 2, characterized in that, The blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to perform the first alignment task on the private datasets of each participant, obtaining the first private data intersection, including: The blockchain layer, based on an optimized privacy-preserving set intersection protocol, instructs consensus nodes to obtain encrypted private data sent by the participants with whom they have established connections, thus obtaining an encrypted private dataset. Consensus nodes select private data with the same encryption value from the encrypted private dataset to complete the first alignment task of the private datasets of each participant, and obtain the first private data intersection.

4. The method according to claim 2, characterized in that, Based on the first private data intersection, the blockchain layer instructs the sorting nodes to perform a second alignment task on the private datasets of each participant, resulting in the final private data intersection, including: The blockchain layer instructs consensus nodes to generate message verification codes based on the intersection of the first private data set and the exchangeable encrypted public key, and to generate transactions and send them to the sorting nodes based on the message verification codes, the exchangeable encrypted public key, the intersection of the first private data set and the corresponding set of participants. The sorting node selects private data with the same encryption value from the transaction set generated by the transactions sent by each consensus node to complete the second alignment task of the private datasets of each participant and obtain the final private data intersection.

5. The method according to claim 1, characterized in that, The federated learning layer instructs each participant to train its local model based on the intersection of private data, obtaining local model parameters, including: The federated learning layer instructs each participant to train a local model based on the intersection of private data, obtains intermediate output, and encrypts the intermediate output using a homomorphic key to obtain encrypted intermediate output. Each participant determines its corresponding encryption gradient, determines a random mask based on a preset random number generator and encryption gradient, and determines a privacy gradient based on the encryption gradient and random mask. Each participant uploads its privacy gradient to the consensus node connected to it in the blockchain layer, and trains the local model based on the decrypted privacy gradient after it has been decrypted by the consensus node, thus obtaining the local model parameters.

6. The method according to claim 1, characterized in that, in, The global model is a customer credit assessment model; Accordingly, in response to data processing requests, data processing is performed based on the global prediction model, including: In response to a credit assessment request from a target customer, the system determines the target customer's credit assessment data based on a global prediction model to provide feedback on the credit assessment request.

7. A data processing apparatus, characterized in that, Implemented based on a federated learning architecture, the federated learning architecture including a blockchain layer, a federated learning layer, and an application layer, the device includes: The alignment module is used to instruct the blockchain layer to perform data alignment on the private datasets of each participant based on a preset cryptographic protocol, obtain the intersection of private data, and send it to the federated learning layer. The aggregation module is used to instruct the federated learning layer to interact with the blockchain layer based on the intersection of private data, to train the local model to obtain local model parameters, and to instruct the blockchain layer to aggregate the local model parameters to obtain global model parameters, and to feed back the global model parameters to each participant and the application layer. The processing module is used to instruct the application layer to obtain the global model based on the global model parameters, and in response to data processing requests, to perform data processing based on the global prediction model.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that is executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data processing method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data processing method according to any one of claims 1-6.