Credit full link checking control method and device, equipment and medium
By acquiring multi-source data and extracting multi-dimensional features, and performing matching and difference analysis based on these features, the problem of low accuracy in traditional data verification schemes has been solved, achieving efficient and accurate data verification across the entire credit chain.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI XINXIAOFEI DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional data verification methods typically use a single dimension for verification, resulting in low verification accuracy.
Acquire multi-source data (accounting data, business and financial data, and financial data), extract multi-dimensional features, perform matching based on multi-dimensional features, and generate a full-chain verification report on credit through difference analysis and intelligent repair.
By using intelligent identification and difference analysis of multi-dimensional features, the accuracy and efficiency of data verification are improved, manual intervention is reduced, and a closed-loop repair mechanism is constructed.
Smart Images

Figure CN121883146A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data quality monitoring technology, and in particular to a method, apparatus, equipment and medium for verification and control of the entire credit chain. Background Technology
[0002] In the daily operations and compliance management of financial institutions, the accurate verification of business data is a key link in ensuring the authenticity of financial reports, preventing operational risks, and meeting regulatory requirements.
[0003] Traditional data verification methods typically use a single dimension for verification, resulting in low verification accuracy. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for verification and control across the entire credit chain, aiming to solve the problem that traditional data verification schemes typically use a single dimension for verification, resulting in low verification accuracy.
[0005] In a first aspect, embodiments of this application provide a verification and control method for the entire credit chain, the verification and control method for the entire credit chain including:
[0006] Acquire multi-source data, wherein the multi-source data includes accounting subject data, business and financial data, and financial data;
[0007] Extract multi-dimensional features from the multi-source data;
[0008] Based on the multi-dimensional features, the target verification type corresponding to the multi-dimensional features is determined, wherein the target verification type includes one or more of accounting subject verification, business and finance verification, or financial and economic verification;
[0009] The multi-dimensional features are used as the features to be matched, and the features to be matched are matched based on the preset matching rules of the target verification type to obtain the matching results;
[0010] Determine whether the matching result meets the preset matching requirements;
[0011] If the matching result does not meet the preset matching requirements, then the multi-dimensional features are subjected to difference analysis and repair to obtain the repaired multi-dimensional features;
[0012] The repaired multi-dimensional features are used as new features to be matched, and the process is repeated in the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, and then the verification report of the credit full-link is generated.
[0013] A further technical solution is that extracting multi-dimensional features from the multi-source data includes:
[0014] The multi-source data is preprocessed to obtain preprocessed multi-source data.
[0015] Basic features are extracted from the preprocessed multi-source data, wherein the basic features include numerical features, temporal features, classification features, text features, and identifier features;
[0016] Based on the aforementioned basic features, derived features are constructed from the preprocessed multi-source data, wherein the derived features include at least one of statistical features, time window features, ratio features, trend features, and cross features;
[0017] The basic features and the derived features are transformed by feature encoding to obtain multi-dimensional features.
[0018] A further technical solution is that the basic features and the derived features are subjected to feature encoding transformation to obtain multi-dimensional features, including:
[0019] Identify the low-cardinality classification features, high-cardinality classification features, and ordered classification features among the basic features and the derived features;
[0020] One-hot encoding is performed on the low cardinality classification features to obtain low cardinality numerical features;
[0021] The high cardinality classification features are subjected to target encoding and embedding encoding to obtain high cardinality numerical features;
[0022] The low-cardinality and high-cardinality numerical features are subjected to feature scaling to obtain a vector of multi-dimensional features.
[0023] A further technical solution is that, after obtaining the multi-dimensional features, the method further includes:
[0024] The attention mechanism is used to calculate the weights of features in different dimensions;
[0025] Based on the weights of features of different dimensions, the multi-dimensional features are weighted and fused to obtain preliminary fused features;
[0026] The preliminary fusion features are processed by feature interaction modeling to obtain low-order fusion interaction features;
[0027] The low-order fusion interaction features are input into a preset neural network for high-order nonlinear processing to obtain deep fusion features.
[0028] A further technical solution is that, after obtaining the deep fusion features, the method further includes:
[0029] The deep fusion features are verified and checked, including data integrity check, logical consistency check, numerical range check, time series continuity check, and outlier check.
[0030] A further technical solution is that the multi-dimensional features are subjected to difference analysis and repair to obtain the repaired multi-dimensional features, including:
[0031] Perform a difference analysis on the multi-dimensional features to obtain the difference results;
[0032] The root cause analysis results are obtained by using a preset root cause identification algorithm to perform root cause analysis on the difference results. The preset root cause identification algorithm includes at least one of decision tree construction, random forest algorithm, expert rule matching, Granger causality test analysis algorithm, and K-means clustering analysis algorithm.
[0033] Based on the root cause analysis results and historical cases, the target remediation plan was determined;
[0034] The multi-dimensional features are repaired based on the target repair scheme to obtain the repaired multi-dimensional features.
[0035] A further technical solution is that the multi-dimensional features are repaired based on the target repair scheme to obtain the repaired multi-dimensional features, including:
[0036] The target repair scheme is decomposed into multiple atomic schemes;
[0037] A distributed lock is used to control the concurrent repair operations of multiple atomic schemes on the multi-dimensional features, thereby obtaining the repaired multi-dimensional features.
[0038] Secondly, embodiments of this application also provide a verification and control device for the entire credit chain, which includes a unit for performing the above-described method.
[0039] Thirdly, embodiments of this application also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0040] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.
[0041] This application provides a method, apparatus, device, and medium for end-to-end credit verification control. The method includes: acquiring multi-source data; extracting multi-dimensional features from the multi-source data; determining the target verification type corresponding to the multi-dimensional features based on the multi-dimensional features; using the multi-dimensional features as features to be matched, and matching the features to be matched based on a preset matching rule for the target verification type to obtain a matching result; determining whether the matching result meets preset matching requirements; if the matching result does not meet the preset matching requirements, performing difference analysis and repair on the multi-dimensional features to obtain repaired multi-dimensional features; using the repaired multi-dimensional features as new features to be matched, and returning to the step of matching the features to be matched based on the preset matching rule for the target verification type, until the matching result meets the preset matching requirements, and generating an end-to-end credit verification report.
[0042] This application embodiment comprehensively analyzes multi-source data, including accounting subject data, business and financial data, and financial data, which avoids missed detections and misjudgments caused by the limitations of a single data source. Furthermore, by extracting multi-dimensional features from multi-source data and intelligently identifying the verification type based on these features, the matching accuracy can be improved. In addition, when the matching results do not meet the preset requirements, not only are the differences identified, but further difference analysis and intelligent repair are also carried out, i.e., a closed-loop repair mechanism is constructed to reduce manual intervention, thereby improving the accuracy and efficiency of data verification. Attached Figure Description
[0043] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0046] Figure 1 A flowchart illustrating the first embodiment of a credit end-to-end verification and control method provided in this application;
[0047] Figure 2 The overall architecture diagram of the credit end-to-end verification platform provided for this application;
[0048] Figure 3The three-in-one self-healing system architecture diagram provided for this application;
[0049] Figure 4 The hardware deployment architecture diagram provided for this application;
[0050] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0052] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0053] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0054] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0055] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0056] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0057] To address the aforementioned issues, this application provides a full-chain verification and control method for credit, which can improve the accuracy and efficiency of data verification.
[0058] See Figure 1 , Figure 1 The flowchart of a first embodiment of a credit end-to-end verification and control method provided in this application is shown. The credit end-to-end verification and control method includes the following steps:
[0059] Step 110: Obtain multi-source data.
[0060] The multi-source data includes accounting subject data, business and financial data, and financial data.
[0061] For example, accounting data may include structured data such as accounting subject codes, balances, and details; business and financial data may include semi-structured data such as business transactions, financial vouchers, and accounting rules; and financial data may include time-series data such as cash flow, position information, and report data.
[0062] Step 120: Extract multi-dimensional features from the multi-source data.
[0063] Among them, multi-dimensional features can include time-dimensional features, monetary-dimensional features, and business-dimensional features.
[0064] For example, time-dimensional features can be timestamp information such as transaction time, accounting time, and settlement time; amount-dimensional features can be numerical features such as transaction amount, exchange rate, and fees; and business-dimensional features can be classification features such as product type, customer level, and risk rating.
[0065] Step 130: Based on the multi-dimensional features, determine the target verification type corresponding to the multi-dimensional features.
[0066] The target verification type includes one or more of the following: accounting subject verification, business and finance verification, or financial and economic verification.
[0067] Step 140: Using the multi-dimensional features as the features to be matched, match the features to be matched based on the preset matching rules of the target verification type to obtain the matching result.
[0068] Step 150: Determine whether the matching result meets the preset matching requirements.
[0069] Step 160: If the matching result does not meet the preset matching requirements, perform difference analysis and repair on the multi-dimensional features to obtain the repaired multi-dimensional features.
[0070] Step 170: Take the repaired multi-dimensional features as new features to be matched, and return to the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, then generate the verification report of the credit full-link.
[0071] This embodiment comprehensively analyzes multi-source data, including accounting subject data, business and financial data, and financial data, to avoid missed detections and misjudgments caused by the limitations of a single data source. Furthermore, by extracting multi-dimensional features from the multi-source data and intelligently identifying the verification type based on these features, it can improve matching accuracy. In addition, when the matching results do not meet the preset requirements, it not only identifies the differences but also conducts further difference analysis and intelligent repair, i.e., it constructs a closed-loop repair mechanism to reduce manual intervention, thereby improving the accuracy and efficiency of data verification.
[0072] In some possible implementations, step 120, namely extracting multi-dimensional features from the multi-source data, includes:
[0073] Step 121: Preprocess the multi-source data to obtain preprocessed multi-source data.
[0074] Preprocessing can include data cleaning, format standardization, and data association, as detailed below:
[0075] 1. Data cleaning and processing:
[0076] 1-1) Null value detection: Use the pandas.isnull() function to detect null values in the data and count the proportion of null values;
[0077] 1-2) Duplicate Data Identification: Identify duplicate records based on the business primary key (transaction serial number + timestamp);
[0078] 1-3) Outlier detection: Use the 3σ principle to detect outliers in the amount field, and set the threshold to mean ± 3 times the standard deviation;
[0079] 1-4) Data type validation: Validate that the amount field is numeric and the time field is datetime format;
[0080] 1-5) Business logic verification: Check whether the debit and credit amounts are balanced and whether the account codes comply with accounting standards.
[0081] 2. Format standardization:
[0082] 2-1) Standardize time format: Convert all time fields to ISO 8601 format, such as YYYY-MM-DDTHH:mm:ss.sssZ;
[0083] 2-2) Amount precision control: The precision of the amount field is uniformly set to two decimal places, and the Decimal type is used to avoid floating-point errors;
[0084] 2-3) Coding standardization: Subject codes are uniformly formatted as 8-digit numbers, and customer numbers are uniformly formatted as 16-digit strings;
[0085] 2-4) Currency standardization: Use 3-letter currency codes based on the ISO 4217 standard (e.g., CNY, USD);
[0086] 2-5) Uniform character encoding: All text fields shall use UTF-8 encoding format.
[0087] 3. Establishing data associations:
[0088] 3-1) Primary Key Mapping Table Construction: Establish a primary key mapping relationship table between different systems;
[0089] 3-2) Foreign key constraint verification: Verify the integrity of foreign key constraints between related tables;
[0090] 3-3) Time window association: Establish the association between transactions and accounting entries based on a ±5 minute time window;
[0091] 3-4) Fuzzy matching algorithm: Use the Levenshtein distance algorithm to perform fuzzy matching of customer names;
[0092] 3-5) Relevance score: Calculate a confidence score (between 0 and 1) for each relationship.
[0093] Step 122: Extract basic features from the preprocessed multi-source data, wherein the basic features include numerical features, temporal features, classification features, text features, and identifier features.
[0094] For example, the extraction of basic features can be referred to the following:
[0095] 1. Numerical Feature Extraction: Extract key numerical fields from structured data, such as transaction amount, account balance, loan interest rate, credit period, etc., for quantitative analysis of core indicators such as capital scale, risk exposure and business scale.
[0096] 2. Time Feature Decomposition: The original timestamp information is broken down into time components such as year, month, day, hour, minute, and weekday, which supports the modeling of the periodicity and time sequence of business behavior and improves the ability to identify time-sensitive scenarios (such as repayment date and batch processing window).
[0097] 3. Categorical Feature Encoding: Extract discrete categorical fields such as product type, customer level, and transaction channel (e.g., online banking, mobile banking, counter) and convert them into numerical forms that the model can process using methods such as one-hot encoding or label encoding to reflect the business differences between different categories.
[0098] 4. Text Feature Processing: Perform natural language processing on unstructured text fields such as transaction summaries and remarks, including Chinese word segmentation, stop word filtering, and keyword extraction, to uncover the business intent or abnormal information contained therein (such as sensitive words like "overdue", "compensation", and "reversal").
[0099] 5. Identifier Feature Generation: Based on business rules, construct Boolean-type derived features, such as "whether it is a holiday", "whether it is during working hours", "whether it is the end of the month / quarter", etc., to enhance the model's ability to perceive special time nodes or business rules and improve the contextual understanding level of the verification logic.
[0100] Step 123: Based on the basic features, construct derived features from the preprocessed multi-source data.
[0101] The derived features include at least one of statistical features, time window features, ratio features, trend features, and cross features.
[0102] For example, the construction of derived features can be seen below:
[0103] 1) Statistical feature calculation: Based on customers' historical transaction data, basic statistical measures are extracted, including the mean, variance, maximum and minimum of transaction amount, which are used to characterize customers' consumption level and behavioral stability.
[0104] 2) Time window characteristics: Transaction frequency and cumulative transaction amount are statistically analyzed according to different time granularities (past 7 days, 30 days, 90 days) to capture the changing trends of customer behavior activity in the short, medium and long term, and enhance the model's ability to perceive behavioral dynamics.
[0105] 3) Ratio feature construction: Calculate the ratio of the current transaction amount to the customer's historical average transaction amount to generate a normalized ratio feature, which is used to identify abnormal transaction behavior (such as a large sudden increase) and improve the sensitivity of fraud or risky transactions.
[0106] 4) Trend Feature Analysis: Perform linear regression fitting on the customer account balance sequence and extract the slope value of the trend as a trend feature to determine the customer's fund flow (such as continuous outflow may indicate the risk of fund transfer).
[0107] 5) Cross-feature generation: Combine discrete features, such as combining "product type" and "customer level" by Cartesian product to generate high-dimensional cross features, explore the behavioral differences of different customer groups on specific products, and enhance the non-linear expressive power of the model.
[0108] Step 124: Perform feature encoding transformation on the basic features and the derived features to obtain multi-dimensional features.
[0109] In some possible implementations, step 124, namely, performing feature encoding transformation on the basic features and the derived features to obtain multi-dimensional features, includes:
[0110] 1) Determine the low-cardinality classification features, high-cardinality classification features, and ordered classification features among the basic features and the derived features;
[0111] Low-cardinality classification features can include product type, transaction channel, etc.; high-cardinality classification features can include historical difference rate and customer ID, etc.; and ordered classification features can be customer level.
[0112] 2) Perform one-hot encoding on the low cardinality classification features to obtain low cardinality numerical features;
[0113] 3) Perform target encoding and embedding encoding on the high cardinality classification features to obtain high cardinality numerical features;
[0114] Among these methods, high cardinality classification features can be target-encoded based on historical difference rates, and customer IDs can be embedded and encoded using Word2Vec.
[0115] 4) Perform feature scaling on the low-cardinality and high-cardinality numerical features to obtain a vector of multi-dimensional features.
[0116] Feature scaling can be performed using StandardScaler to standardize numerical features.
[0117] In some embodiments, after obtaining multi-dimensional features, data fusion processing can be performed. The data fusion stage may include dimension alignment processing, multi-dimensional feature fusion, and consistency verification. For details, please refer to the following embodiments.
[0118] The dimension alignment process can include the following steps:
[0119] 1) Time alignment: Align data from different systems according to a unified time granularity (minute level);
[0120] 2) Business key matching: Data matching is performed based on business keys such as customer number, product number, and transaction serial number;
[0121] 3) Missing value imputation: Use methods such as forward imputation, backward imputation, and linear interpolation to handle missing values;
[0122] 4) Data resampling: Resampling data from different frequencies to a unified time frequency;
[0123] 5) Dimensional expansion: Extend low-dimensional data to high dimensions through copying or interpolation.
[0124] In some possible implementations, multidimensional feature fusion may include the following steps 125-128:
[0125] Step 125: Use the attention mechanism to calculate the weights of features in different dimensions.
[0126] Step 126: Based on the weights of features of different dimensions, perform weighted fusion processing on the multi-dimensional features to obtain preliminary fused features.
[0127] For example, attention weights can be used to weight and fuse the three-dimensional features of subjects, business and finance, and finance.
[0128] Step 127: Perform feature interaction modeling processing on the preliminary fusion features to obtain low-order fusion interaction features.
[0129] Step 128: Input the low-order fusion interaction features into a preset neural network for high-order nonlinear processing to obtain deep fusion features.
[0130] For steps 125-128, the attention mechanism in the Transformer architecture is introduced through attention weight calculation to dynamically evaluate the relative importance of each dimension of subject, business and finance, and finance in the current task, so as to achieve adaptive focusing on key information.
[0131] Subsequently, based on the calculated attention weights, weighted feature fusion is performed to integrate the feature representations of the three dimensions in a weighted manner, generating a preliminary fusion feature vector to ensure that the high-value dimension dominates the fusion result.
[0132] To further capture the nonlinear relationships between features, feature interaction modeling is adopted. The Factorization Machine (FM) model is used to explicitly model the second-order cross relationships between features, effectively identifying business-meaning combination patterns such as "high amount + non-working time".
[0133] Based on this, a multi-layer feedforward neural network is constructed through deep feature learning. The fused features are nonlinearly transformed and abstracted, and higher-order feature combinations are extracted layer by layer to enhance the model's expressive power and generalization performance.
[0134] In addition, residual connections can be introduced into the network structure to directly pass input information to the deep network, which can alleviate the gradient vanishing problem during the training of deep models and improve the stability and convergence efficiency of the fusion network.
[0135] In some possible implementations, after obtaining the deep fusion features, the method further includes:
[0136] Step 129: Perform verification checks on the deep fusion features, wherein the verification checks include data integrity checks, logical consistency checks, numerical range checks, time series continuity checks, and outlier checks.
[0137] Step 129 may specifically include the following:
[0138] 1) Data integrity check: Verify the integrity of the merged data to ensure no data loss;
[0139] 2) Logical consistency verification: Check the consistency of business logic such as debit and credit balance and account correspondence;
[0140] 3) Numerical range verification: Verify whether the fused numerical features are within a reasonable range;
[0141] 4) Time series continuity: Examine the continuity and monotonicity of time series data;
[0142] 5) Outlier detection: The Isolation Forest algorithm is used to detect outliers in the merged data.
[0143] In some possible implementations, step 160 involves performing difference analysis and repair on the multi-dimensional features to obtain repaired multi-dimensional features, including:
[0144] Step 161: Perform a difference analysis on the multi-dimensional features to obtain the difference results.
[0145] Step 161 may include precise differential localization and impact propagation analysis, as detailed below:
[0146] For example, precise difference localization can be achieved in the following ways:
[0147] 1) Data fingerprint calculation: Calculate the data fingerprint for each record using the MD5 hash algorithm;
[0148] 2) Difference comparison algorithm: The Myers difference algorithm is used to accurately locate inconsistent fields;
[0149] 3) Hierarchical positioning: narrowing the range of differences layer by layer from table level to row level to field level;
[0150] 4) Time window analysis: Analyze the time window and duration of the differences;
[0151] 5) Determine the spatial scope: Determine the scope of systems, modules, and data tables involved in the differences.
[0152] For example, influence propagation analysis can be performed in the following ways:
[0153] 1) Dependency Graph Traversal: Based on a pre-built data lineage and system dependency graph, the breadth-first search (BFS) algorithm is used to start from the source node of the difference and traverse its downstream dependent nodes layer by layer to identify all potentially affected data tables, metrics, services and system modules, forming a complete impact propagation path graph.
[0154] 2) Impact Quantification: For each affected downstream node, a weighted scoring model is designed based on its data importance weight, call frequency, and dependency depth to calculate the degree of impact from the difference, quantifying it into an impact score of 0-100. The higher the score, the more severely the data reliability of that node is compromised.
[0155] 3) Business Process Mapping: Associating and mapping the affected technical nodes with core business processes. For example, mapping "abnormal reconciliation results" to "end-of-day clearing process", and "customer balance calculation error" to "funds transfer process", thereby transforming technical issues into understandable business process risks.
[0156] 4) User Impact Assessment: Further analyze the user groups served by the affected business processes, assess the scope of impact (such as the number of customers involved and the number of transactions) and degree of impact (such as whether it leads to service interruption, financial loss or compliance risks) on end users, and form a user-side impact report.
[0157] 5) Risk Level Classification: Based on the comprehensive impact scope, business criticality, user impact, and remediation timeliness requirements, the differences are classified into five risk levels: P0 (Critical), P1 (Severe), P2 (Moderate), P3 (Minor), and P4 (Warning). This classification result serves as the core basis for subsequent root cause analysis, priority scheduling, and remediation resource allocation.
[0158] Step 162: Using a preset root cause identification algorithm, perform root cause analysis on the difference results to obtain root cause analysis results. The preset root cause identification algorithm includes at least one of decision tree construction, random forest algorithm, expert rule matching, Granger causality test analysis algorithm, and K-means clustering analysis algorithm.
[0159] Root cause analysis includes causal reasoning, time series analysis, and correlation analysis. Causal reasoning is based on the causal graph model to analyze the causal relationship of inconsistent data. Time series analysis mainly analyzes the time series of data changes and identifies abnormal change points. Correlation analysis mainly analyzes the correlation and influence path between different data items.
[0160] For example, the preset root cause identification algorithm can be found below:
[0161] 1) Decision Tree Construction: Based on historical difference datasets, a root cause identification decision tree model is trained using CART or C4.5 algorithms. This model generates interpretable "if-then" decision rules by recursively partitioning the feature space, intuitively demonstrating the reasoning path from difference features to root cause categories, making it easy for business personnel to understand and verify.
[0162] 2) Feature Importance Analysis: An ensemble learning approach is used to learn the differential data using the random forest algorithm. By calculating the average impurity decrease (or Gini importance) of each feature during the decision-making process, its contribution to root cause judgment is quantified. This result is used to identify key driving factors, providing a quantitative basis for root cause ranking and prioritization.
[0163] 3) Expert Rule Matching: The feature vector of the current discrepancy is matched against a pre-built expert rule base for pattern matching. The rule base is composed of common problem patterns summarized by domain experts (e.g., "borrowing imbalance occurs at the end of the month" → "closing script not executed"). The rule engine performs inference and outputs the root cause hypothesis for matching. This mechanism effectively integrates prior knowledge and improves the efficiency and reliability of identifying known problems.
[0164] 4) Temporal Causality Analysis: For differential data with time-series characteristics, the Granger Causality Test is used to analyze the temporal influence relationships between different variables. By examining whether the historical value of one variable has a significant predictive power for the current value of another variable, it is determined whether there is a statistically significant causal direction, which helps to identify abnormal propagation paths.
[0165] 5) Cluster Analysis: Unsupervised clustering algorithms such as K-means are applied to group historical difference samples according to patterns to discover potential outlier categories. The clustering results can be used to identify new or unknown problem patterns and provide a basis for subsequent root cause analysis, enabling "same approach to similar problems".
[0166] Step 163: Based on the root cause analysis results and historical cases, determine the target remediation plan.
[0167] Among these methods, a repair strategy can be generated by predefined repair strategy templates for common problems, or by recommending the optimal repair strategy based on historical repair experience, or by assessing the risks and scope of impact of the repair operation.
[0168] For example, the target remediation plan can be generated in the following manner:
[0169] (1) Template library retrieval: based on predefined template matching and repair schemes;
[0170] (2) Rule engine reasoning: Derive repair solutions based on business rules;
[0171] (3) Case-Based Reasoning (CBR): Generate repair solutions based on similar historical cases;
[0172] (4) Constraint Solving: Generate feasible repair solutions based on the constraint satisfaction problem (CSP).
[0173] In some embodiments, 3-5 candidate remediation solutions can be generated for each problem, such as generating a list of remediation suggestions containing multiple candidate solutions and their scores, applicable scenarios and risk warnings, for human decision-making or automated systems to select and execute.
[0174] In some embodiments, the generated repair solution may be comprehensively evaluated, wherein the comprehensive evaluation may include the following:
[0175] 1) Feasibility assessment: Assess the technical feasibility and resource availability of the proposed solution;
[0176] 2) Risk assessment matrix: Construct a two-dimensional assessment matrix of risk probability × impact degree;
[0177] 3) Cost-benefit analysis: Calculate the implementation costs and expected benefits of the proposed solution;
[0178] 4) Time complexity analysis: Evaluate the execution time and resource consumption of the proposed solution;
[0179] 5) Side effect analysis: Analyze the potential negative impacts and side effects of the proposed solution.
[0180] In some embodiments, the optimal solution can be selected as the target repair solution in the following manner.
[0181] 1) Multi-objective optimization: The NSGA-II algorithm is used for multi-objective optimization selection;
[0182] 2) Analytic Hierarchy Process (AHP): The AHP method is used to assign weights to the evaluation indicators;
[0183] 3) TOPSIS method: The TOPSIS method is used to rank and select alternatives;
[0184] 4) Monte Carlo simulation: Using the Monte Carlo method to simulate the uncertainties in the execution of the scheme;
[0185] 5) Expert review: Key solutions need to be reviewed by the expert review committee.
[0186] In some embodiments, to ensure the safety, reliability, and recoverability of the repair operation, a series of comprehensive pre-check steps are required before actually implementing the repair plan. These steps aim to ensure the smooth progress of the repair process and to enable a rapid rollback to a stable state in the event of anomalies. Specifically, these include the following five aspects:
[0187] 1) Data Backup: Use incremental backup technology to back up the relevant data tables and configuration files. This step ensures that even in the event of unexpected situations during the repair process, the system can be quickly restored based on the latest backup data. Incremental backups only copy the parts that have changed since the last backup, thereby improving backup efficiency and reducing storage requirements.
[0188] 2) Environment Check: Conduct a comprehensive check of system resources (such as CPU, memory, and disk space), network connectivity, and permission configurations. Confirm that all necessary resources and services are in normal working order to avoid repair failures due to environmental issues. For example, ensure there is sufficient disk space for temporary file storage, a stable network connection, and that the current user has the necessary permissions to perform the required operations.
[0189] 3) Dependency Verification: Verify that all services and components on which the repair operation depends are functioning correctly. This includes, but is not limited to, critical infrastructure such as database services, message queues, and external API interfaces. Through health checks or heartbeat mechanisms, ensure that all dependencies are available to prevent cascading failures caused by the unavailability of third-party services.
[0190] 4) Rollback Preparation: Prepare rollback scripts and related data in advance to ensure that a rollback operation can be performed immediately should problems occur during the repair process. The rollback plan should detail each step and its expected effect, and be thoroughly tested to ensure its effectiveness. In addition, a snapshot of the state before the repair should be saved to facilitate accurate revert to the original state.
[0191] 5) Monitoring Configuration: Configure real-time monitoring and alarm mechanisms for the repair process to promptly grasp the repair progress and system status changes. Set thresholds for key indicators (such as response time and error rate), and trigger alarms to notify relevant personnel when the set values are reached or exceeded. Simultaneously, log information is recorded for post-repair auditing and analysis to ensure the entire repair process is transparent and controllable.
[0192] Step 164: Repair the multi-dimensional features based on the target repair scheme to obtain the repaired multi-dimensional features.
[0193] In some possible implementations, step 164, namely, repairing the multi-dimensional features based on the target repair scheme to obtain the repaired multi-dimensional features, includes:
[0194] 1) Decompose the target repair scheme into multiple atomic schemes;
[0195] 2) Use a distributed lock to control the concurrent repair operations of multiple atomic schemes on the multi-dimensional features to obtain the repaired multi-dimensional features.
[0196] In other words, to ensure the controllability, consistency, and recoverability of the repair operation, this solution adopts a refined step-by-step execution control mechanism, decomposing complex repair tasks into orderly and manageable execution units, and achieving safe and reliable automated execution through multiple safeguards. The specific implementation method is as follows:
[0197] 1) Execution Plan Decomposition: The high-level repair plan is broken down into a series of atomic execution steps. Each step represents an indivisible minimum unit of operation (such as "update configuration item A", "execute SQL script 1", "restart service B"), ensuring that each operation has clear inputs, outputs and expected results, providing an execution basis for subsequent control mechanisms.
[0198] 2) Checkpoint Setup: Persistent checkpoints are set after critical steps are successfully executed to record the status and context information of the completed steps. This mechanism supports resuming interrupted tasks. If a repair task is interrupted due to an exception, the system can resume execution from the last successful checkpoint, avoiding repeated operations or starting from scratch, significantly improving execution efficiency and fault tolerance.
[0199] 3) Transaction management: Critical steps involving data modification are encapsulated in database transactions and executed in accordance with ACID principles.
[0200] 4) Concurrency control: A distributed lock mechanism (such as Redis or ZooKeeper) is used to coordinate repair tasks globally, ensuring that only one execution instance modifies shared resources (such as core data tables and configuration centers) at any given time.
[0201] 5) Progress Tracking: A real-time progress tracking system will be built to dynamically collect and display the execution status of repair tasks. A visual interface will present the current execution steps, completion percentage, time statistics, and key event logs, achieving end-to-end observability of the repair process. Furthermore, progress information can be used for external system integration, alarm triggering, and automated scheduling decisions.
[0202] In some embodiments, the repair results can be verified, such as by re-executing the verification algorithm to verify data consistency; or by verifying whether the repaired data meets business logic constraints; or by inviting business users to conduct acceptance tests to confirm the repair effect.
[0203] In some embodiments, a dynamic rule management approach can be adopted, such as defining standard rule templates that support parameterized configuration; or supporting rule version management and canary release; or setting the priority and dependencies of rule execution.
[0204] The rule matching process may include a rule parsing phase, a rule execution phase, and a rule learning phase, as detailed below:
[0205] 1. Rule parsing phase:
[0206] 1-1) Rule compilation and processing: For example, using regular expressions to decompose business rule text into token sequences, constructing an abstract syntax tree (AST) to represent the logical structure of the rules, verifying the existence and type matching of fields and functions referenced in the rules; converting the AST into executable Python or Java code, or optimizing the generated bytecode to reduce execution overhead.
[0207] 1-2) Dependency analysis: For example, constructing a directed acyclic graph (DAG) to represent the dependencies between rules; using the Kahn algorithm to perform topological sorting of rules and determine the execution order; using depth-first search to detect circular dependencies between rules; calculating dependency strength weights based on data flow between rules; and identifying rule groups that can be executed in parallel to improve execution efficiency.
[0208] 1-3) Execution path optimization: For example, pruning inapplicable rule branches in advance based on data characteristics; implementing short-circuit evaluation of logical expressions to reduce unnecessary calculations; caching intermediate calculation results to avoid duplicate calculations; creating indexes for rule condition fields to accelerate condition matching speed; generating the optimal rule execution plan to minimize execution time.
[0209] 2. Rule enforcement phase:
[0210] 2-1) Conditional matching algorithms: such as converting input data into feature vector representation; using vectorization operations to batch evaluate rule conditions; using edit distance for fuzzy matching of text conditions; using interval tree data structure to optimize numerical range queries; and using compiled regular expressions for text pattern matching.
[0211] 2-2) Parallel execution mechanisms: such as using thread pools to manage rules for task execution and control concurrency; splitting large amounts of data into different threads for parallel processing; using the MapReduce model to aggregate the execution results of each thread; implementing an exception isolation mechanism so that the failure of a single rule does not affect other rules; and dynamically adjusting the task allocation between threads to achieve load balancing.
[0212] 2-3) Result evaluation calculation: For example, calculate the confidence score (0-1) of the result based on the degree of rule matching; check the consistency between the results of multiple rules and identify conflicts; assign different weights according to the historical accuracy of the rules; use a weighted voting mechanism to merge the prediction results of multiple rules; use Bayesian methods to quantify the uncertainty of the prediction results.
[0213] 3. Rule learning phase:
[0214] 3-1) Collection of effect feedback: For example, collect user feedback on rule results in real time through API interface; process historical execution results and manually labeled data in batches on a regular basis; evaluate the quality and credibility of feedback data; remove noisy feedback and malicious feedback data; and calculate the accuracy, recall, F1 score and other indicators of the rules.
[0215] 3-2) Automatic parameter tuning: For example, using grid search to find the optimal combination of rule parameters; using Bayesian optimization algorithm to efficiently search the parameter space; using gradient descent to optimize differentiable rule parameters; using k-fold cross-validation to evaluate the effect of parameter tuning; and implementing an early stopping mechanism to prevent overfitting of parameter optimization.
[0216] 3-3) Rule evolution algorithm: For example, encode rules as chromosomes and define gene representation methods; design fitness functions to evaluate the overall performance of rules; use roulette wheel selection to select excellent rule individuals; implement rule crossover and mutation operations to generate new rules; retain the best rule individuals in each generation to prevent degeneration.
[0217] In some embodiments, a real-time data verification algorithm may be employed, specifically employing streaming data processing and incremental verification mechanisms, as described below:
[0218] 1. Streaming data processing:
[0219] 1-1) Flink stream processing configuration: For example, configure a Kafka data source, set consumer groups and partitioning strategies; or use Avro serialization format to ensure data type safety and backward compatibility; enable checkpointing mechanism and set the checkpoint interval to 30 seconds; use RocksDB state backend to support large state storage; configure restart strategy and fault recovery mechanism.
[0220] 1-2) Window aggregation processing: For example, use a 5-minute sliding window with a 1-minute step to aggregate data; or set a session window based on user activity with a timeout of 30 minutes; or configure a composite trigger based on time and data volume; or use a custom aggregation function to calculate the total amount and the number of transactions; or set the allowed late time to 2 minutes, and write the timeout data to the output stream on the side.
[0221] 1-3) Time semantic management: For example, generating watermarks based on data timestamps, tolerating 5 seconds of out-of-order delivery; or extracting business time from the message body as event time; or recording the processing time when data arrives at the system; or uniformly converting the time of different systems to UTC time or using the NTP protocol to ensure clock synchronization of each node.
[0222] 2. Incremental verification mechanism:
[0223] 2-1) Change log processing: such as CDC data capture, change event parsing, change data filtering, change serialization, and change deduplication, i.e., removing duplicate change events based on transaction ID and LSN.
[0224] 2-2) Merkle tree construction: For example, divide the data into blocks of a fixed size (1000 records); or use the SHA-256 algorithm to calculate the hash value for each data block; or build a binary Merkle tree structure from bottom to top; or calculate the root hash value of the entire dataset; or store the tree node information in the Redis cache.
[0225] 2-3) Fast difference location: such as root hash comparison, recursive decomposition, leaf node location, record-level comparison, etc.
[0226] 3. Parallel processing optimization:
[0227] 3-1) Data sharding strategies: such as hash sharding, range sharding, consistent hashing, sharding routing, sharding monitoring, etc.
[0228] 3-2) Parallel computing architecture: such as thread pool configuration, task queue, work stealing, memory management and exception isolation, etc.
[0229] 3-3) Result merging algorithms: such as phased merging, sorted merging, deduplication, priority sorting, batch output, etc.
[0230] In some embodiments, the flow of the intelligent difference analysis algorithm can be referred to as follows:
[0231] Among them, differential feature extraction includes statistical feature calculation, time series feature analysis, and correlation feature mining.
[0232] For example, the calculation of statistical characteristics may include the following:
[0233] 1-1) Basic statistics: Calculate the mean, median, standard deviation, skewness, and kurtosis of the difference amounts;
[0234] 1-2) Quantile analysis: Calculate the 25%, 50%, 75%, 95%, and 99% quantiles;
[0235] 1-3) Frequency statistics: Statistically count the frequency, relative frequency, and cumulative frequency of differences;
[0236] 1-4) Distribution Fitting: Use normal distribution and log-normal distribution to fit the distribution of difference amounts;
[0237] 1-5) Outlier identification: Use box plots to identify statistical outliers.
[0238] 2. Time series feature analysis:
[0239] 2-1) Periodic detection: Using FFT transform to detect periodic patterns of differences;
[0240] 2-2) Trend Analysis: Linear regression analysis was used to analyze the long-term trend of the difference in amount;
[0241] 2-3) Seasonal decomposition: Use STL decomposition to extract seasonal, trend, and random components;
[0242] 2-4) Autocorrelation analysis: Calculate the autocorrelation function and partial autocorrelation function of the differential time series;
[0243] 2-5) Change point detection: Use the CUSUM algorithm to detect change points in the difference pattern.
[0244] 3. Feature mining:
[0245] 3-1) Correlation analysis: Calculate the Pearson correlation coefficients between different types of differences;
[0246] 3-2) Mutual information calculation: Mutual information is used to quantify the nonlinear association between differential variables;
[0247] 3-3) Causal relationship inference: Granger causality test was used to analyze the causal relationship between differences;
[0248] 3-4) Association rule mining: Using the Apriori algorithm to mine association rules of differential combinations;
[0249] 3-5) Network Analysis: Construct differential correlation networks and analyze network topology characteristics.
[0250] In some embodiments, clustering analysis algorithms may employ K-means clustering, hierarchical clustering, DBSCAN clustering, Gaussian mixture models, etc.
[0251] In some embodiments, anomaly detection can use Isolation Forest, or One-Class SVM to identify differences that deviate from the normal pattern; or LOF algorithm to detect differences with local density anomalies; or deep autoencoder to detect anomalies with large reconstruction errors; or statistical methods such as Grubbs test and Dixon test to detect outliers.
[0252] In some embodiments, a multi-level caching architecture of L1 local cache + L2 distributed cache can be adopted; or caching methods such as intelligent preloading can be adopted to reduce the verification processing time from hours to minutes, while improving memory usage efficiency by 60% and system throughput by 500%.
[0253] Based on the above embodiments, the overall architecture of the credit end-to-end verification platform provided in this application is as follows: Figure 2 As shown, the overall architecture of the credit full-chain verification platform includes core modules such as business system access module, data acquisition and processing module, three-dimensional verification module, AI intelligent analysis module and service interface module.
[0254] The AI intelligent analysis module has an intelligent analysis layer and a self-healing repair layer, while the three-dimensional verification module has a verification engine layer and an operation management layer. In other words, the credit full-link verification platform adopts a layered microservice architecture design, which is divided into seven core layers from bottom to top: Layer 1: Data access layer, Layer 2: Data processing layer, Layer 3: Verification engine layer, Layer 4: Intelligent analysis layer, Layer 5: Self-healing repair layer, Layer 6: Service interface layer, and Layer 7: Operation management layer.
[0255] For details, please refer to the following description for each floor:
[0256] The first layer is the data access layer, which is mainly used for real-time data stream access, batch data synchronization, and API data interfaces.
[0257] Among them, real-time data stream access: supports real-time data streams from message queues such as Kafka and RocketMQ;
[0258] Batch data synchronization: Supports batch data synchronization methods such as direct database connection and file transfer;
[0259] API Data Interface: Provides RESTful API and GraphQL interface for data interaction.
[0260] The second layer is the data processing layer, which mainly includes data cleaning, data transformation, and data routing.
[0261] Data cleaning includes: removing duplicate data, handling outliers, and standardizing data formats.
[0262] Data conversion: Enables data format conversion and mapping between different systems;
[0263] Data routing: Routing data to the appropriate verification engine according to business rules;
[0264] The third layer: the verification engine layer, which may include the subject verification engine, the business and finance verification engine, and the financial verification engine.
[0265] Among them, the subject verification engine is responsible for verifying data related to accounting subjects;
[0266] Business-Finance Reconciliation Engine: Handles data reconciliation between business systems and financial systems;
[0267] Financial Reconciliation Engine: Enables data reconciliation between the financial system and the capital system.
[0268] The fourth layer: the intelligent analysis layer, is mainly used for the following operations:
[0269] 1) Rules Engine: Performs intelligent matching and validation based on business rules;
[0270] 2) Machine learning model: Using AI algorithms for anomaly detection and pattern recognition;
[0271] 3) Knowledge Graph: Construct a business knowledge graph to support complex reasoning.
[0272] Fifth layer: Self-healing and repair layer, mainly used for the following operations:
[0273] 1) Discrepancy analysis: In-depth analysis of the root causes of data inconsistencies;
[0274] 2) Strategy Generation: Generate remediation strategies based on historical experience and business rules;
[0275] 3) Automatic execution: Automatically executes repair operations and verifies the repair effect.
[0276] Layer 6: Service Interface Layer, mainly used for the following operations:
[0277] 1) Verification Service: Provides a standardized verification service interface;
[0278] 2) Monitoring services: Real-time monitoring of system operation status and business indicators;
[0279] 3) Configuration service: Supports dynamic configuration management and rule updates.
[0280] The seventh layer: the operations management layer, primarily responsible for the following operations:
[0281] 1) Visual monitoring: Provides real-time monitoring dashboards and business dashboards;
[0282] 2) Report Analysis: Generate various verification reports and analysis reports;
[0283] 3) System Management: Functions such as user permission management and system configuration.
[0284] In addition, the three-in-one self-healing system and hardware deployment provided in this application can be found in [reference needed]. Figure 3 and Figure 4 .
[0285] Based on the above embodiments, the credit end-to-end verification and control method provided in this application has the following technical effects:
[0286] 1) Efficiency Improvement Effects: 1-1) Improved Processing Speed: The verification and processing time has been reduced from 2 hours to 3 minutes, improving efficiency by 97.5%; 1-2) Improved Automation: 92.3% of problems are automatically repaired, significantly reducing manual intervention; 1-3) Improved Real-Time Performance: Supports millisecond-level real-time verification, enabling timely detection and handling of problems.
[0287] 2) Improved accuracy: 2-1) Verification accuracy: 99.6%, 14.4% higher than traditional methods; 2-2) Difference detection rate: 98.8%, significantly improving the ability to detect problems; 2-3) Reduced false alarm rate: The false alarm rate is controlled below 0.4%.
[0288] 3) Security Enhancement Effects: 3-1) Risk Prevention and Control Capabilities: Identify potential risks in advance, with a prevention rate of 95%; 3-2) Data Consistency Guarantee: Ensure 99.9% data consistency; 3-3) Business Continuity Guarantee: Quickly fix problems and avoid business interruption.
[0289] Thus, by constructing a three-in-one self-healing system integrating subject, business and finance, and financial management, this application effectively solves the problems of low accuracy, poor efficiency, and inability to automatically repair traditional verification methods.
[0290] Corresponding to the above-described end-to-end credit verification and control method, this application also provides an end-to-end credit verification and control device. This end-to-end credit verification and control device includes a unit for executing the end-to-end credit verification and control method of any of the above embodiments, and can be configured in a desktop computer, tablet computer, laptop computer, or other terminal.
[0291] The verification and control device for the entire credit chain may include the following units:
[0292] An acquisition unit is used to acquire multi-source data, wherein the multi-source data includes accounting subject data, business and financial data, and financial data;
[0293] Extraction unit, used to extract multi-dimensional features from the multi-source data;
[0294] The determining unit is configured to determine the target verification type corresponding to the multi-dimensional features based on the multi-dimensional features, wherein the target verification type includes one or more of accounting subject verification, business and finance verification, or financial and economic verification;
[0295] The matching unit is used to take the multi-dimensional features as features to be matched, and match the features to be matched based on the preset matching rules of the target verification type to obtain the matching result;
[0296] A judgment unit is used to determine whether the matching result meets the preset matching requirements;
[0297] An analysis and repair unit is used to perform difference analysis and repair on the multi-dimensional features if the matching result does not meet the preset matching requirements, so as to obtain the repaired multi-dimensional features.
[0298] The generation unit is used to take the repaired multi-dimensional features as new features to be matched and return to the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, and then generate the verification report of the credit full-link.
[0299] like Figure 5 As shown in the figure, this application provides a computer device including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0300] Memory 113 is used to store computer programs;
[0301] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the credit end-to-end verification control method provided in any of the foregoing method embodiments, including:
[0302] Acquire multi-source data, wherein the multi-source data includes accounting subject data, business and financial data, and financial data;
[0303] Extract multi-dimensional features from the multi-source data;
[0304] Based on the multi-dimensional features, the target verification type corresponding to the multi-dimensional features is determined, wherein the target verification type includes one or more of accounting subject verification, business and finance verification, or financial and economic verification;
[0305] The multi-dimensional features are used as the features to be matched, and the features to be matched are matched based on the preset matching rules of the target verification type to obtain the matching results;
[0306] Determine whether the matching result meets the preset matching requirements;
[0307] If the matching result does not meet the preset matching requirements, then the multi-dimensional features are subjected to difference analysis and repair to obtain the repaired multi-dimensional features;
[0308] The repaired multi-dimensional features are used as new features to be matched, and the process is repeated in the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, and then the verification report of the credit full-link is generated.
[0309] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0310] Therefore, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, wherein when executed by a processor, the computer program implements the steps of the credit end-to-end verification and control method provided in any of the foregoing method embodiments, including:
[0311] Acquire multi-source data, wherein the multi-source data includes accounting subject data, business and financial data, and financial data;
[0312] Extract multi-dimensional features from the multi-source data;
[0313] Based on the multi-dimensional features, the target verification type corresponding to the multi-dimensional features is determined, wherein the target verification type includes one or more of accounting subject verification, business and finance verification, or financial and economic verification;
[0314] The multi-dimensional features are used as the features to be matched, and the features to be matched are matched based on the preset matching rules of the target verification type to obtain the matching results;
[0315] Determine whether the matching result meets the preset matching requirements;
[0316] If the matching result does not meet the preset matching requirements, then the multi-dimensional features are subjected to difference analysis and repair to obtain the repaired multi-dimensional features;
[0317] The repaired multi-dimensional features are used as new features to be matched, and the process is repeated in the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, and then the verification report of the credit full-link is generated.
[0318] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.
[0319] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0320] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0321] The steps in the methods of this application embodiment can be adjusted, merged, or deleted according to actual needs. The units in the apparatus of this application embodiment can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0322] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0323] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0324] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Since these modifications and variations fall within the scope of the claims and their equivalents, this application also intends to include these modifications and variations.
[0325] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for verification and control across the entire credit chain, characterized in that, The method includes: Acquire multi-source data, wherein the multi-source data includes accounting subject data, business and financial data, and financial data; Extract multi-dimensional features from the multi-source data; Based on the multi-dimensional features, the target verification type corresponding to the multi-dimensional features is determined, wherein the target verification type includes one or more of accounting subject verification, business and finance verification, or financial and economic verification; The multi-dimensional features are used as the features to be matched, and the features to be matched are matched based on the preset matching rules of the target verification type to obtain the matching results; Determine whether the matching result meets the preset matching requirements; If the matching result does not meet the preset matching requirements, then the multi-dimensional features are subjected to difference analysis and repair to obtain the repaired multi-dimensional features; The repaired multi-dimensional features are used as new features to be matched, and the process is repeated in the step of matching the features to be matched based on the preset matching rules of the target verification type until the matching result meets the preset matching requirements, and then the verification report of the credit full-link is generated.
2. The credit end-to-end verification and control method according to claim 1, characterized in that, The extraction of multi-dimensional features from the multi-source data includes: The multi-source data is preprocessed to obtain preprocessed multi-source data. Basic features are extracted from the preprocessed multi-source data, wherein the basic features include numerical features, temporal features, classification features, text features, and identifier features; Based on the aforementioned basic features, derived features are constructed from the preprocessed multi-source data, wherein the derived features include at least one of statistical features, time window features, ratio features, trend features, and cross features; The basic features and the derived features are transformed by feature encoding to obtain multi-dimensional features.
3. The verification and control method for the entire credit chain according to claim 2, characterized in that, The step of performing feature encoding transformation on the basic features and the derived features to obtain multi-dimensional features includes: Identify the low-cardinality classification features, high-cardinality classification features, and ordered classification features among the basic features and the derived features; One-hot encoding is performed on the low cardinality classification features to obtain low cardinality numerical features; The high cardinality classification features are subjected to target encoding and embedding encoding to obtain high cardinality numerical features; The low-cardinality and high-cardinality numerical features are subjected to feature scaling to obtain a vector of multi-dimensional features.
4. The verification and control method for the entire credit chain according to claim 1, characterized in that, After obtaining the multi-dimensional features, the method further includes: The attention mechanism is used to calculate the weights of features in different dimensions; Based on the weights of features of different dimensions, the multi-dimensional features are weighted and fused to obtain preliminary fused features; The preliminary fusion features are processed by feature interaction modeling to obtain low-order fusion interaction features; The low-order fusion interaction features are input into a preset neural network for high-order nonlinear processing to obtain deep fusion features.
5. The credit end-to-end verification and control method according to claim 4, characterized in that, After obtaining the deep fusion features, the method further includes: The deep fusion features are verified and checked, including data integrity check, logical consistency check, numerical range check, time series continuity check, and outlier check.
6. The verification and control method for the entire credit chain according to claim 1, characterized in that, The step of performing difference analysis and repair on the multi-dimensional features to obtain the repaired multi-dimensional features includes: Perform a difference analysis on the multi-dimensional features to obtain the difference results; The root cause analysis results are obtained by using a preset root cause identification algorithm to perform root cause analysis on the difference results. The preset root cause identification algorithm includes at least one of decision tree construction, random forest algorithm, expert rule matching, Granger causality test analysis algorithm, and K-means clustering analysis algorithm. Based on the root cause analysis results and historical cases, the target remediation plan was determined; The multi-dimensional features are repaired based on the target repair scheme to obtain the repaired multi-dimensional features.
7. The credit end-to-end verification and control method according to claim 6, characterized in that, The process of repairing the multi-dimensional features based on the target repair scheme to obtain the repaired multi-dimensional features includes: The target repair scheme is decomposed into multiple atomic schemes; A distributed lock is used to control the concurrent repair operations of multiple atomic schemes on the multi-dimensional features, thereby obtaining the repaired multi-dimensional features.
8. A verification and control device for the entire credit chain, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1-7.