Risk control credit monitoring method based on cloud computing
Through cross-modal data fusion, three-level privacy protection and intelligent hybrid cloud scheduling, the problems of data silos, privacy protection and resource scheduling in traditional risk control credit monitoring methods are solved, safe, efficient and transparent credit monitoring is achieved, and risk assessment and resource utilization are improved.
Patent Information
- Application Number
- CN202510396141.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the scenarios of cloud computing and multi-institution collaboration, traditional risk control credit monitoring methods have problems such as inefficient integration of data silos and heterogeneous integration, separation of privacy protection methods, rigid hybrid cloud resource scheduling and unreal-time compliance audits.
The cross-modal data deep fusion mechanism, three-level privacy protection system, intelligent hybrid cloud scheduling strategy and dynamic compliance closed loop are adopted, and multi-source heterogeneous data is integrated through federated learning, combined with homomorphic encryption and differential privacy to achieve hierarchical privacy protection, and hybrid cloud resource scheduling and dynamic model update technology are used to improve the real-time risk assessment, and a compliance audit closed loop is built based on blockchain and interpretability analysis.
It realizes the precise construction and real-time iteration of cross-institutional risk control models, provides a safe, efficient and transparent credit monitoring solution, breaks data silos, improves risk assessment dimensions, ensures data privacy and compliance, and optimizes resource utilization.
Smart Images

Figure CN120338944A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud computing, and particularly relates to a risk control credit monitoring method based on cloud computing. Background Art
[0002] Traditional risk control credit monitoring methods face the following key bottlenecks in the scenarios of cloud computing and multi-institution collaboration, which are specifically manifested as follows: data islands and inefficient heterogeneous integration, existing methods usually only rely on structured data of a single institution and lack the ability to deeply analyze unstructured text; fragmented privacy protection means, existing solutions mostly adopt isolated technologies, such as homomorphic encryption only protecting static data, while the federated learning framework does not cover the risk of gradient leakage; rigid hybrid cloud resource scheduling, traditional hybrid cloud architectures do not schedule tasks according to data sensitivity levels, resulting in high-security-demand computations being forced to be deployed in public clouds, or non-sensitive tasks occupying private cloud resources; compliance auditing and decision-making black boxes, existing methods rely on offline log auditing and lack real-time evidence storage and interpretability support, such as the risk scores of black-box models (such as deep learning) cannot be traced back to specific data features and it is difficult to pass the "right to explanation" clause review of GDPR.
[0003] In view of the above problems, based on the core innovations of cross-modal data deep fusion mechanism, three-level privacy protection system, intelligent hybrid cloud scheduling strategy and dynamic compliance closed-loop, on the premise of ensuring data privacy and compliance, the accurate construction and real-time iteration of cross-institution risk control models are realized, providing a safe, efficient and transparent credit monitoring solution for the fintech field. Summary of the Invention
[0004] The present invention discloses a risk control credit monitoring method based on cloud computing, which integrates multi-source heterogeneous data through federated learning, realizes hierarchical privacy protection by combining homomorphic encryption and differential privacy, improves the real-time performance of risk assessment by using hybrid cloud resource scheduling and dynamic model update technology, and constructs a compliance auditing closed-loop based on blockchain and interpretability analysis.
[0005] The technical solution of the present invention is realized as follows:
[0006] The present invention discloses a risk control credit monitoring method based on cloud computing, including the following steps:
[0007] S1. Real-time collect user transaction records, credit history and enterprise financial data through a structured data unit to generate structured data tags; integrate an NLP natural language processing engine through an unstructured text unit to parse free text in social media, enterprise announcements and judicial documents, extract entity keywords and generate text feature tags; in a cross-institution scenario, a federated learning client divides multi-source data according to institutional domains and generates encrypted feature indexes to provide metadata for distributed training;
[0008] S2. Based on the encrypted feature index generated in S1, use the homomorphic encryption algorithm to process sensitive data, support ciphertext transmission, storage, and partial calculations; the federated learning framework aligns the feature dimensions of multi-source encrypted data to generate cross-domain joint training intermediate parameters; the dynamic access control unit generates a dynamic token based on the unique hash value generated from the user role and device hardware information, authorizing entities to access the desensitized data subset within the specified time window;
[0009] S3. Based on the intermediate parameters and desensitized data in S2, the private cloud node stores the core sensitive data and performs redundant backups, and calls the elastic resources of the public cloud to perform non-sensitive feature extraction and model training; the federated learning coordinator splits the training task into local subtasks and global aggregation tasks; the task distribution controller allocates computing tasks based on the structured data labels, text feature labels generated in S1, and the public cloud resource status, and forcibly routes sensitive calculations to the private cloud node through the private cloud gateway policy;
[0010] S4. Based on the local subtask training results in S3, the rule engine calls the risk control rule library to output risk events and confidence levels. The machine learning unit aggregates the local base models of multiple participants and constructs a global scoring model through weighted average or model integration methods; the integrated SHAP analysis unit generates a feature contribution report, and the local differential privacy mechanism adds noise to the model gradient to obscure sensitive features; the outputs of the rule engine and the model unit are fused through weighting to generate a risk score, and the weights are dynamically adjusted according to the real-time calculated model accuracy and recall rate;
[0011] S5. Based on the risk score in S4, the stream data engine receives the user transaction records in S1 and the risk score in S4, processes the user behavior sequence to generate a dynamic portrait; the adaptive update unit monitors the model decay metrics and triggers incremental learning to update the model parameters; during the asynchronous update process of the federated framework, the parameter server periodically aggregates the global model parameters, and the federated learning coordinator manages task splitting and resource allocation;
[0012] S6. Aggregate the text features in S1, the risk metrics in S4, the dynamic portrait in S5, and the federated learning operation records to construct a multi-dimensional relationship network of users-enterprises-transactions and render a risk heat map; the compliance audit unit generates an operation chain with a time stamp through blockchain evidence storage, analyzes the mapping relationship between the federated learning records and the privacy protocol terms, and outputs an audit report integrating the evaluation items of the GDPR General Data Protection Regulation, the PIPL Personal Information Protection Law, and the traceable matrix based on the blockchain operation chain; the independent verification module verifies the identity authentication information, the integrity of gradient transmission, and the compliance of differential privacy noise addition.
[0013] Preferably, the specific steps of S1 are as follows:
[0014] S1-1. The structured data unit collects the transaction flow data of financial institutions in real time through a distributed message queue, cleans the missing values and outliers in the data using pattern matching technology based on regular expressions, maps the numerical features to a standardized interval using a normalization method, generates structured data labels including account identification, transaction timestamp, and amount dimensions, and persists the structured data labels in a columnar storage format;
[0015] S1-2. The unstructured text parsing module loads a pre-trained language model to perform semantic analysis on judicial documents and enterprise announcements, identifies legal entities and financial indicators in the text through a bidirectional neural network architecture, stores the extraction results in a graph structure data format, and establishes an entity association mapping with the structured data;
[0016] S1-3. The cross-institutional federated data catalog constructs a distributed index based on a cryptographic hash table, aligns the features of multi-source heterogeneous data in space, extracts the data distribution features through a deep neural network to generate encrypted metadata, and uses a decentralized synchronization protocol to ensure the consistency of metadata updates.
[0017] Preferably, the S2 specifically includes the following steps:
[0018] S2-1. The homomorphic encryption module performs polynomial transformation operations on sensitive fields, completes feature cross-operation in the ciphertext space, generates a verifiable encrypted feature matrix, and optimizes the matrix dimension through linear algebra methods;
[0019] S2-2. The feature hashing process uses a non-collision hash function to map multi-source data to a unified vector space, calculates the statistical distribution features in the hash bucket to generate joint training intermediate parameters, and serializes the parameters in a compact binary format to improve the transmission efficiency;
[0020] S2-3. The dynamic token generator fuses the device hardware fingerprint and digital certificate information, generates a time-limited access credential through the hash message authentication code algorithm, binds the access credential to the device geographical fence and implements a time window verification mechanism.
[0021] Preferably, the S3 specifically includes the following steps:
[0022] S3-1. The private cloud storage cluster is deployed using a distributed file system architecture, realizes redundant storage of data blocks based on the erasure code algorithm, integrates a hardware-level encryption module in the storage process to protect data security, and realizes the key life cycle management through a dedicated security device;
[0023] S3-2. The hybrid cloud resource scheduler monitors the computing load metrics of the graphics processor, dynamically allocates feature processing tasks to different architecture acceleration chips according to preset policies, and implements a load balancing and failover mechanism during the task distribution process;
[0024] S3-3. The trusted task routing controller enables a hardware isolation protection environment for highly sensitive tasks according to the data security level label, establishes an encrypted communication channel that complies with the transport layer security standard, and adopts an intelligent path selection algorithm for the network routing policy.
[0025] Preferably, the specific steps of S4 are as follows:
[0026] S4-1. The rule engine loads an anti-fraud rule library containing account abnormal behavior patterns, and uses the Rete algorithm of the Drools rule engine for multi-condition inference matching. The rule conditions include risk event feature combinations such as sudden increase in transaction frequency, cross-regional device login, and blacklist of large transfer payees. The risk events that match successfully trigger the alarm work order generation process;
[0027] S4-2. The federated model aggregator receives the local model parameter update amounts of each participant through a secure multi-party computing protocol, and globally aggregates the hidden layer weight matrix of the credit scoring model using a weighted average algorithm. During the aggregation process, bilinear mapping is used to verify the parameter integrity and source authenticity;
[0028] S4-3. The differential privacy module injects random noise that satisfies the probability density distribution into the forward model update amount before gradient transmission. The noise parameters are dynamically adjusted according to the data sensitivity level to ensure that the intermediate parameters during the training process cannot be reverse-derived to the original training samples.
[0029] Preferably, the specific steps of S5 are as follows:
[0030] S5-1. The user portrait generator processes the time-series transaction behavior data through a long short-term memory neural network. The network structure includes a bidirectional recurrent layer and an attention mechanism, and outputs a dynamic feature vector that represents the consumption periodicity and capital flow pattern. The dimension of the feature vector is automatically aligned with the input layer of the credit scoring model;
[0031] S5-2. The model decay monitor continuously calculates the relative change rate of the area under the ROC curve of the validation set. When it is detected that the model prediction performance drops by more than the preset threshold, an incremental training task containing samples in the recent time period is triggered. The incremental training uses the elastic weight consolidation algorithm to prevent catastrophic forgetting;
[0032] S5-3. The asynchronous update coordinator optimizes the parameter synchronization process using a historical gradient momentum buffering mechanism. The coordinator maintains an independent version control queue for each participant and determines the parameter fusion ratio by comparing the cosine similarity between the global model and the local model.
[0033] Preferably, the specific steps of S6 are as follows:
[0034] S6-1. The graph computing engine constructs a heterogeneous relationship graph including the enterprise shareholder chain, guarantee relationships, and the actual controller path, calculates the node centrality score using a graph embedding algorithm based on random walk, and the risk propagation model identifies hidden related-party credit risks by simulating the capital flow;
[0035] S6-2. The blockchain evidence storage module organizes data operation records into a hash chain structure in chronological order. Each block contains the operation type, the digital identity of the executor, and the fingerprint information of the operation object. The stored evidence data achieves distributed ledger consistency through the Byzantine fault-tolerant consensus algorithm;
[0036] S6-3. The compliance validator uses a non-interactive zero-knowledge proof protocol to verify the cross-domain data transmission path. The prover generates an evidence chain containing the data hash value and routing nodes, and the verifier confirms that no man-in-the-middle tampering has occurred during the transmission process through elliptic curve bilinear pair operations.
[0037] Compared with the prior art, the advantages of the present invention are as follows:
[0038] (1) Deep integration of multi-source heterogeneous data: Structured data units and unstructured text units work together. The NLP engine extracts entity keywords and generates feature tags, and combined with the federated learning client, cross-institutional data domain alignment is achieved, breaking data islands and enhancing the risk assessment dimension;
[0039] (2) Hierarchical design for privacy protection: Homomorphic encryption (data level), local differential privacy (gradient level), and federated learning (collaboration level) are integrated to form a multi-level protection system; Dynamic access control is based on device fingerprints and time-limited tokens to ensure minimum privilege access, balancing security and efficiency;
[0040] (3) Elastic resource scheduling for hybrid cloud: Through the task distribution controller and private cloud gateway policy, sensitive computing is forced to route to the private cloud, and non-sensitive tasks are elastically allocated to the public cloud, optimizing resource utilization; Redundant backup and erasure code technologies ensure high data availability. Brief Description of the Drawings
[0041] Figure 1 It is a flowchart of the overall method of the present invention. Detailed Embodiments
[0042] Embodiment
[0043] As Figure 1 shown, the embodiment of the present invention discloses a risk control credit monitoring method based on cloud computing, including the following steps:
[0044] S1. Real-time collect user transaction records, credit history, and enterprise financial data through structured data units to generate structured data tags; integrate the NLP natural language processing engine through unstructured text units to parse free text in social media, enterprise announcements, and judicial documents, extract entity keywords, and generate text feature tags; in a cross-institutional scenario, the federated learning client divides multi-source data by institutional domain and generates encrypted feature indexes to provide metadata for distributed training;
[0045] The specific steps of S1 are as follows:
[0046] S1-1. The structured data unit real-time collects the transaction flow data of financial institutions through a distributed message queue, uses pattern matching technology based on regular expressions to clean missing values and outliers in the data, and applies a normalization method to map numerical features to a standardized interval to generate structured data tags containing account identification, transaction timestamp, and amount dimensions. The structured data tags are persistently stored in a columnar storage format;
[0047] S1-2. The unstructured text parsing module loads a pre-trained language model to perform semantic analysis on judicial documents and enterprise announcements, identifies legal entities and financial indicators in the text through a bidirectional neural network architecture, stores the extraction results in a graph structure data format, and establishes an entity association mapping with the structured data;
[0048] S1-3. The cross-institutional federated data catalog constructs a distributed index based on an encrypted hash table, performs spatial alignment on multi-source heterogeneous data features, extracts data distribution features through a deep neural network to generate encrypted metadata, and uses a decentralized synchronization protocol to ensure consistency in metadata updates. The specific contents are as follows: clarify the fault tolerance mechanism and conflict resolution rules, and add adversarial generation networks and contrast learning techniques:
[0049] S1-3-1. Distributed index construction: The encrypted hash table uses a collision-resistant hash function based on the SM3 national cryptography algorithm to generate index key values, and each key value corresponds to the logical address of the data shard within the institutional domain; the index key values achieve fuzzy matching in the cross-institutional ciphertext state through the homomorphic encryption algorithm (Paillier encryption), supporting privacy-preserving joint queries;
[0050] S1-3-2. Spatial alignment operation: Use an adversarial generation network (GAN) to train a multi-source data feature space projector to map heterogeneous data from different institutions to a unified semantic space; the alignment error is optimized through a contrast learning loss function (InfoNCE) to ensure that the cosine similarity of the feature vectors of the same entity in different institutional domains is ≥0.85;
[0051] S1-3-3. The deep neural network adopts a residual convolution structure. The input is the data distribution histogram and statistical features within the institutional domain, and the output layer generates a 128-dimensional encrypted metadata vector. The metadata vector is fragmented and stored in multiple federated nodes through the threshold secret sharing algorithm, and no single node can restore the original data distribution alone;
[0052] S1-3-4. Decentralized synchronization protocol: An improved Raft consensus algorithm (supporting Byzantine fault tolerance) is adopted. When metadata is updated, a pre-commit vote is initiated, and nodes need to verify the consistency of the data hash value and digital signature. During the synchronization process, version conflicts are resolved through a timestamp-based vector clock mechanism, and the latest operation record is preferentially retained;
[0053] S2. Based on the encrypted feature index generated by S1, a homomorphic encryption algorithm is used to process sensitive data, supporting ciphertext transmission, storage, and partial calculations. The federated learning framework aligns the feature dimensions of multi-source encrypted data to generate cross-domain joint training intermediate parameters. The dynamic access control unit generates a dynamic token based on the unique hash value generated from the user role and device hardware information, authorizing entities to access the desensitized data subset within the specified time window;
[0054] The above S2 specifically includes the following steps:
[0055] S2-1. The homomorphic encryption module performs polynomial transformation operations on sensitive fields, completes feature cross-operation in the ciphertext space, and generates a verifiable encrypted feature matrix. The matrix dimension is compressed and optimized through linear algebra methods, specifically including the following:
[0056] Clarify the specific algorithm and parameter configuration of homomorphic encryption, and add a zero-knowledge proof verification mechanism:
[0057] S2-1-1. Polynomial transformation operation: For sensitive fields (such as ID numbers and transaction amounts), an RLWE (Ring Learning with Errors) homomorphic encryption scheme is adopted to map the plaintext to the ciphertext coefficients on the polynomial ring. The polynomial order is set to 2048, the modulus q is a safe prime number (q = 12289), and the noise distribution adopts a discrete Gaussian distribution (standard deviation σ = 3.19);
[0058] S2-1-2. Ciphertext feature cross-operation: Perform feature cross-multiplication operations in the ciphertext state, use polynomial multiplication to generate cross terms, and decompose the operation result into a low-dimensional polynomial group through the Chinese Remainder Theorem (CRT). Cross-term verification adopts a zero-knowledge proof protocol (zk-SNARK) to generate a verifiable proof of the correctness of the ciphertext operation;
[0059] S2-1-3. Matrix Compression Optimization: The encrypted feature matrix is approximated by singular value decomposition (SVD) for low rank, retaining the first k singular values (k = min(m,n) × 0.2), and the truncation error threshold is set to 1e-4; the compressed matrix is stored in a sparse matrix format, and the non-zero elements are further compressed for storage space through run-length encoding (RLE);
[0060] S2-2. The feature hashing process uses a non-collision hashing function to map multi-source data to a unified vector space, calculates the statistical distribution features within the hash bucket to generate intermediate parameters for joint training, and parameter serialization uses a compact binary format to improve transmission efficiency;
[0061] S2-3. The dynamic token generator fuses the device hardware fingerprint and digital certificate information, generates a time-sensitive access credential through the hash message authentication code algorithm, binds the access credential to the device geofence and implements a time window verification mechanism;
[0062] S3. Based on the intermediate parameters and desensitized data of S2, the private cloud node stores the core sensitive data and performs redundant backups, and calls the elastic resources of the public cloud to perform non-sensitive feature extraction and model training; the federated learning coordinator splits the training tasks into local subtasks and global aggregation tasks; the task distribution controller allocates computing tasks based on the structured data labels, text feature labels, and public cloud resource status generated by S1, and forcibly routes sensitive computations to the private cloud node through the private cloud gateway policy;
[0063] S3 specifically includes the following steps:
[0064] S3-1. The private cloud storage cluster is deployed using a distributed file system architecture, realizes redundant storage of data blocks based on the erasure code algorithm, integrates a hardware-level encryption module during the storage process to protect data security, and key lifecycle management is implemented through a dedicated security device;
[0065] S3-2. The hybrid cloud resource scheduler monitors the computing load metrics of the graphics processor, dynamically allocates feature processing tasks to different architecture acceleration chips according to preset policies, and implements a load balancing and failover mechanism during the task distribution process;
[0066] S3-3. The trusted task routing controller, based on the data security level label, enables a hardware isolation protection environment for highly sensitive tasks, establishes an encrypted communication channel that complies with the transport layer security standard, and the network routing policy uses an intelligent path selection algorithm;
[0067] S4. Based on the local subtask training results in S3, the rule engine calls the risk control rule library to output risk events and confidence levels. The machine learning unit aggregates the local base models of multiple parties and constructs a global scoring model through weighted average or model integration methods; the integrated SHAP analysis unit generates a feature contribution report, and the local differential privacy mechanism adds noise to the model gradients to obscure sensitive features; the outputs of the rule engine and the model unit are fused through weighting to generate a risk score, and the weights are dynamically adjusted according to the real-time calculated model accuracy and recall rate;
[0068] S4 specifically includes the following steps:
[0069] S4-1. The rule engine loads the anti-fraud rule library containing account abnormal behavior patterns and uses the Rete algorithm of the Drools rule engine for multi-condition inference matching. The rule conditions include risk event feature combinations such as sudden increase in transaction frequency, cross-regional device login, and blacklist of large transfer payees. The matched risk events trigger the process of generating an alarm work order;
[0070] S4-2. The federated model aggregator receives the local model parameter update amounts of each party through a secure multi-party computing protocol and globally aggregates the hidden layer weight matrix of the credit scoring model using the weighted average algorithm. During the aggregation process, bilinear mapping is used to verify the parameter integrity and source authenticity, which specifically includes the following:
[0071] Explain the mathematical basis for dynamic weight adjustment:
[0072] S4-2-1. Secure multi-party computing protocol: Using the Shamir secret sharing scheme, each party divides its local model parameters into n fragments (n≥3t+1, t is the maximum number of fault-tolerant nodes) and distributes them to other parties; the aggregator reconstructs the global parameters through Lagrange interpolation, and a threshold signature mechanism is implemented during the process to verify the legality of the fragment sources;
[0073] S4-2-2. Weighted average algorithm: The weights are calculated based on the F1-score of the local model on the validation set, and the dynamic adjustment formula is:
[0074]
[0075] Among them, the aggregation of the hidden layer weight matrix uses block parallel computing, and each matrix block independently performs weighted average operations, and the block size is set to 256×256;
[0076] F1 i : The F1-score (the harmonic mean of precision and recall) of the local model of the i-th party on the validation set; N: The total number of parties in federated learning - square operation (F1 2)It is used to amplify the weight contribution of the high-precision model; the denominator is the sum of the squares of the F1-scores of all participating parties, ensuring weight normalization;
[0077] S4-2-3. Bilinear mapping verification: Define the bilinear pairing function e: G1×G2→G T , where G1 and G2 are elliptic curve groups, and the participants need to provide the group element commitments of the model parameters;
[0078] Verification equation:
[0079]
[0080] where H(W i ) is the BLAKE2 hash value of the weight matrix. If the verification fails, the parameter rollback mechanism is triggered; g1, g2: are the generators of groups G1 and G2 respectively;
[0081] S4-3. The differential privacy module injects random noise that satisfies the probability density distribution into the forward model update amount before gradient transmission. The noise parameters are dynamically adjusted according to the data sensitivity level to ensure that the intermediate parameters during the training process cannot be reversely deduced to the original training samples;
[0082] S5. Based on the risk score of S4, the streaming data engine receives the user transaction records of S1 and the risk score of S4, processes the user behavior sequence to generate a dynamic portrait; the adaptive update unit monitors the model decay index and triggers incremental learning to update the model parameters; during the asynchronous update process of the federated framework, the parameter server periodically aggregates the global model parameters, and the federated learning coordinator manages task splitting and resource allocation;
[0083] The specific steps of S5 are as follows:
[0084] The user portrait generator processes the time-series transaction behavior data through a long short-term memory neural network. The network structure includes a bidirectional recurrent layer and an attention mechanism, and outputs a dynamic feature vector that characterizes the consumption periodicity and capital flow pattern. The dimension of the feature vector is automatically aligned with the input layer of the credit scoring model;
[0085] The model decay monitor continuously calculates the relative change rate of the area under the ROC curve of the validation set. When it is detected that the model prediction performance drops by more than the preset threshold, an incremental training task including samples in the recent time period is triggered. The incremental training uses the elastic weight consolidation algorithm to prevent catastrophic forgetting;
[0086] The asynchronous update coordinator uses the historical gradient momentum buffering mechanism to optimize the parameter synchronization process. The coordinator maintains an independent version control queue for each participating party and determines the parameter fusion ratio by comparing the cosine similarity between the global model and the local model;
[0087] S6. Aggregate the text features of S1, the risk indicators of S4, the dynamic portraits of S5, and the federated learning operation records to construct a multi-dimensional user-enterprise-transaction relationship network and render a risk heat map; the compliance audit unit generates an operation chain with a time stamp through blockchain evidence storage, analyzes the mapping relationship between the federated learning records and the privacy protocol terms, and outputs an audit report integrating the evaluation items of the GDPR General Data Protection Regulation, the PIPL Personal Information Protection Law, and the traceability matrix based on the blockchain operation chain; the independent verification module verifies the identity authentication information, the integrity of gradient transmission, and the compliance of differential privacy noise addition;
[0088] S6 specifically includes the following steps:
[0089] S6-1. The graph computing engine constructs a heterogeneous relationship graph including the enterprise shareholder chain, the guarantee relationship, and the actual controller path, calculates the node centrality score using the graph embedding algorithm based on random walk, and the risk propagation model identifies the hidden related-party credit risk by simulating the fund flow direction. The specific contents are as follows: clarify the graph walk parameters and the centrality calculation formula, and enhance the technical repeatability after supplementation:
[0090] S6-1-1. Construction of the heterogeneous relationship graph: Define three types of edge attributes: shareholder relationship (shareholding ratio ≥ 5%), guarantee relationship (joint liability type), and controller path (penetrating the shareholding level ≥ 3 layers); fuse the risk scores of S4 and the dynamic portrait features of S5 for node attributes, and interpolate and complete the missing attribute values through the Graph Convolutional Network (GCN);
[0091] S6-1-2. Optimization of the graph embedding algorithm: Use the Node2Vec algorithm to generate node embedding vectors, with the bias parameters set as p = 0.5 (breadth-first) and q = 2.0 (depth-first), and the walk length is 80 steps;
[0092] Improved formula for calculating the centrality score:
[0093] C(v) = α·PageRank(v) + β·Betweenness(v)
[0094] where α = 0.6, β = 0.4, and optimize the parameter combination through grid search;
[0095] S6-1-3. Risk propagation model: Based on Monte Carlo simulation of the fund flow path, the initial risk source is the high-risk enterprise node marked by S4, and the propagation probability formula:
[0096] P u→v = 1 - exp(-λ·w uv ·R u )
[0097] where λ is the attenuation factor (λ = 0.85), w uvis the edge weight, R u is the risk score of node u;
[0098] Hidden related party identification adopts the subgraph isomorphism detection algorithm (VF2 algorithm) to match the minimum spanning tree structure of known risk patterns;
[0099] S6-2. The blockchain evidence storage module organizes the data operation records into a hash chain structure in chronological order. Each block contains the operation type, the digital identity of the executor, and the fingerprint information of the operation object. The evidence storage data achieves distributed ledger consistency through the Byzantine fault-tolerant consensus algorithm;
[0100] S6-3. The compliance validator uses the non-interactive zero-knowledge proof protocol to verify the cross-domain data transmission path. The prover generates an evidence chain containing the data hash value and the routing nodes, and the verifier confirms that there is no man-in-the-middle tampering during the transmission process through elliptic curve bilinear pair operations.
[0101] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A risk control credit monitoring method based on cloud computing, characterized in that, It includes the following steps: S1. Real-time collect user transaction records, credit history and enterprise financial data through structured data units to generate structured data tags; integrate the NLP natural language processing engine through unstructured text units to parse the free text in social media, enterprise announcements and judicial documents, extract entity keywords and generate text feature tags; in a cross-institutional scenario, the federated learning client divides multi-source data by institutional domain and generates encrypted feature indexes to provide metadata for distributed training; S2. Based on the encrypted feature indexes generated in S1, use the homomorphic encryption algorithm to process sensitive data, support ciphertext transmission, storage and partial calculations; the federated learning framework aligns the feature dimensions of multi-source encrypted data to generate intermediate parameters for cross-domain joint training; the dynamic access control unit generates dynamic tokens based on the unique hash values generated from user roles and device hardware information, authorizing entities to access the desensitized data subset within the specified time window; S3. Based on the intermediate parameters and desensitized data in S2, the private cloud node stores the core sensitive data and performs redundant backups, and calls the elastic resources of the public cloud to perform non-sensitive feature extraction and model training; The federated learning coordinator splits the training task into local subtasks and global aggregation tasks; the task distribution controller allocates computing tasks based on the structured data tags, text feature tags and public cloud resource status generated in S1, and forcibly routes sensitive calculations to the private cloud node through the private cloud gateway policy; S4. Based on the local subtask training results in S3, the rule engine calls the risk control rule library to output risk events and confidence levels. The machine learning unit aggregates the local base models of multiple participants and constructs a global scoring model through weighted average or model integration methods; the integrated SHAP analysis unit generates a feature contribution report, and the local differential privacy mechanism adds noise to the model gradient to obscure sensitive features; the outputs of the rule engine and the model unit are fused through weighting to generate a risk score, and the weights are dynamically adjusted according to the model accuracy and recall rate calculated in real time; S5. Based on the risk score in S4, the stream data engine receives the user transaction records in S1 and the risk score in S4, and processes the user behavior sequence to generate a dynamic portrait; The adaptive update unit monitors the model decay metrics and triggers incremental learning to update the model parameters; during the asynchronous update process of the federated framework, the parameter server periodically aggregates the global model parameters, and the federated learning coordinator manages task splitting and resource allocation; S6. Aggregate the text features in S1, the risk metrics in S4, the dynamic portrait in S5 and the federated learning operation records to construct a multi-dimensional relationship network of users-enterprises-transactions and render a risk heat map; The compliance audit unit generates an operation chain with a timestamp through blockchain evidence storage, parses the mapping relationship between the federated learning records and the privacy protocol terms, and outputs an audit report integrating the evaluation items of the GDPR General Data Protection Regulation, the PIPL Personal Information Protection Law and the traceable matrix based on the blockchain operation chain; the independent verification module verifies the compliance of the identity authentication information, the integrity of the gradient transmission and the addition of differential privacy noise.
2. The risk control credit monitoring method based on cloud computing according to claim 1, characterized in that, The specific steps of S1 include the following: S1-1. The structured data unit collects the transaction flow data of financial institutions in real time through a distributed message queue, uses pattern matching technology based on regular expressions to clean the missing values and outliers in the data, and applies a normalization method to map numerical features to a standardized interval, generating structured data tags including account identifiers, transaction timestamps, and amount dimensions. The structured data tags are persistently stored in a columnar storage format. S1-2. The unstructured text parsing module loads a pre-trained language model to perform semantic analysis on judicial documents and enterprise announcements, identifies legal entities and financial indicators in the text through a bidirectional neural network architecture, stores the extraction results in a graph structure data format, and establishes an entity association mapping with the structured data. S1-3. The cross-institutional federated data catalog constructs a distributed index based on a cryptographic hash table, performs spatial alignment on multi-source heterogeneous data features, extracts data distribution features through a deep neural network to generate encrypted metadata, and uses a decentralized synchronization protocol to ensure consistency during metadata updates.
3. A risk control credit monitoring method based on cloud computing according to claim 1, characterized in that, The specific steps of S2 are as follows: S2-1. The homomorphic encryption module performs polynomial transformation operations on sensitive fields, completes feature cross-operation in the ciphertext space, generates a verifiable encrypted feature matrix, and optimizes the matrix dimension through linear algebra methods. S2-2. The feature hashing process uses a non-collision hash function to map multi-source data to a unified vector space, calculates the statistical distribution features within the hash bucket to generate intermediate parameters for joint training, and uses a compact binary format for parameter serialization to improve transmission efficiency. S2-3. The dynamic token generator fuses device hardware fingerprints and digital certificate information, generates time-sensitive access credentials through the HMAC algorithm, binds the access credentials to the device geofence, and implements a time window verification mechanism.
4. A risk control credit monitoring method based on cloud computing according to claim 1, characterized in that, The specific steps of S3 are as follows: S3-1. The private cloud storage cluster is deployed using a distributed file system architecture, implements redundant storage of data blocks based on the erasure code algorithm, integrates a hardware-level encryption module during the storage process to protect data security, and realizes key lifecycle management through a dedicated security device. S3-2. The hybrid cloud resource scheduler monitors the computing load metrics of the graphics processor, dynamically allocates feature processing tasks to different architecture acceleration chips according to preset policies, and implements load balancing and failover mechanisms during the task distribution process. S3-3. The trusted task routing controller, based on the data security level tags, enables a hardware isolation protection environment for highly sensitive tasks, establishes an encrypted communication channel that complies with the Transport Layer Security standard, and uses an intelligent path selection algorithm for network routing policies.
5. The risk control credit monitoring method based on cloud computing according to claim 1, characterized in that, The specific steps of S4 are as follows: S4-1. The rule engine loads an anti-fraud rule library containing account abnormal behavior patterns, uses the Rete algorithm of the Drools rule engine for multi-condition inference and matching. The rule conditions include combinations of risk event features such as sudden increase in transaction frequency, cross-regional device login, and blacklist of payees for large transfers. The risk events that match successfully trigger the process of generating alarm work orders. S4-2. The federal model aggregator receives the local model parameter updates from each participant through a secure multi-party computation protocol, and globally aggregates the hidden layer weight matrix of the credit scoring model using a weighted average algorithm. During the aggregation process, bilinear mapping is used to verify the parameter integrity and source authenticity. S4-3. The differential privacy module injects random noise that satisfies a probability density distribution into the forward model update before gradient transmission. The noise parameters are dynamically adjusted according to the data sensitivity level to ensure that the intermediate parameters during the training process cannot be reverse-engineered to obtain the original training samples.
6. The risk control credit monitoring method based on cloud computing according to claim 1, characterized in that The specific steps of S5 are as follows: S5-1. The user portrait generator processes the time-series transaction behavior data through a long short-term memory neural network. The network structure includes a bidirectional recurrent layer and an attention mechanism, and outputs a dynamic feature vector that characterizes the consumption periodicity and capital flow pattern. The dimension of the feature vector is automatically aligned with the input layer of the credit scoring model. S5-2. The model decay monitor continuously calculates the relative change rate of the area under the ROC curve of the validation set. When it is detected that the model prediction performance drops by more than a preset threshold, an incremental training task that includes samples from the most recent time period is triggered. The elastic weight consolidation algorithm is used for incremental training to prevent catastrophic forgetting. S5-3. The asynchronous update coordinator uses a historical gradient momentum buffering mechanism to optimize the parameter synchronization process. The coordinator maintains an independent version control queue for each participant and determines the parameter fusion ratio by comparing the cosine similarity between the global model and the local model.
7. A risk control credit monitoring method based on cloud computing according to claim 1, characterized in that The specific steps of S6 are as follows: S6-1. The graph computing engine constructs a heterogeneous relationship graph that includes the enterprise shareholder chain, guarantee relationships, and the actual controller path, and calculates the node centrality score using a graph embedding algorithm based on random walk. The risk propagation model identifies the hidden related-party credit risks by simulating the capital flow. S6-2. The blockchain evidence storage module organizes the data operation records into a hash chain structure in chronological order. Each block contains the operation type, the digital identity of the executor, and the fingerprint information of the operation object. The evidence storage data achieves distributed ledger consistency through the Byzantine fault-tolerant consensus algorithm. S6-3. The compliance validator uses a non-interactive zero-knowledge proof protocol to verify the cross-domain data transmission path. The prover generates an evidence chain that includes the data hash value and the routing nodes, and the verifier confirms that no man-in-the-middle tampering has occurred during the transmission process through elliptic curve bilinear pair operations.
Citation Information
Patent Citations
Federal learning method and application based on block chain and homomorphic encryption
CN114491616A
Credit granting data processing method and system of comprehensive credit system
CN118037440A
Supply chain financial credit risk assessment method and system based on federated learning and privacy protection, and storage medium
CN118864087A
Intelligent risk control system and method for cross-border e-commerce digitization
CN119398479A
Sample data processing method and device, equipment and medium
CN119417593A
Cited By
Multi-center collaborative acquisition and management method and system for student physique test data
CN120528947A
Project data quality checking method and system based on machine learning
CN120543124A
Tough city resource allocation method and system based on block chain and edge computing
CN120579803A
Resilient city resource allocation method and system based on blockchain and edge computing
CN120579803B
Resource metering method and device, electronic equipment and storage medium
CN120750683A