Financial compliance risk assessment method, system and storage medium based on big data

By building a multi-level risk assessment index system and a cascaded integrated network, a comprehensive and accurate assessment and timely warning of financial compliance risks are achieved, and the problems of low efficiency and poor real-time performance in traditional methods are solved, and the level of automation of risk management is improved.

CN120182006BActive Publication Date: 2025-08-15FOSHAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510654978.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-15
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The traditional financial compliance risk assessment method is inefficient, difficult to cope with massive financial transaction data, lack consistency and objectivity, unable to monitor risk changes in real time, unable to fully grasp the systematic characteristics of risks, and weak risk quantification and early warning mechanisms.

Method used

By collecting and standardizing multi-source data, a structured data set is formed, a multi-level financial compliance risk assessment index system is built, and a multi-dimensional real-time analysis is performed using a cascaded fusion network training model, potential risk points are identified and quantitatively evaluated, and a hierarchical warning threshold is set to generate warning signals.

Benefits of technology

It has achieved comprehensive and accurate assessment and timely warning of financial compliance risks, improved the automation level of risk management and scientific decision-making capabilities, and solved the problems of strong artificial dependence, high subjectivity and poor real-time in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182006B_ABST
    Figure CN120182006B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology and discloses a financial compliance risk assessment method, system, and storage medium based on big data. The method includes: collecting and standardizing multi-source data to form a structured dataset; constructing a three-level assessment system consisting of risk areas, subcategories, and indicators; training a layered fusion network to obtain an assessment model; generating risk reports through real-time analysis; identifying risk points and conducting quantitative assessments; and setting early warning thresholds. When these thresholds are exceeded, early warning signals are generated and sent to the risk management department. This application enables efficient identification, accurate assessment, and timely early warning of compliance risks, enhancing the automation level and scientific decision-making capabilities of financial institutions' compliance risk management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a financial compliance risk assessment method, system and storage medium based on big data. Background Art

[0002] With the rapid development of financial markets and the continuous emergence of financial innovations, financial institutions face increasingly complex compliance challenges. Financial regulators are imposing increasingly stringent compliance requirements on financial institutions, requiring them to strictly adhere to laws and regulations in their business operations and mitigate various compliance risks. Traditional financial compliance risk assessment methods rely primarily on manual review and limited data analysis. Financial institutions typically assess compliance risks through regular compliance inspections, internal audits, and manual sampling checks. As business scale and complexity increase, the volume of financial transaction data is exploding, and the data dimensions and types are becoming increasingly diverse. Regulatory requirements are also constantly evolving and improving. Anti-money laundering, counter-terrorist financing, and customer identity verification compliance requirements introduced by regulators around the world are becoming increasingly stringent, requiring financial institutions to address more comprehensive and sophisticated compliance challenges.

[0003] However, traditional financial compliance risk assessment methods have numerous limitations. First, manual review is inefficient and unable to cope with massive amounts of financial transaction data and complex business scenarios, resulting in delayed risk assessments and an inability to timely identify potential risks. Second, traditional methods rely heavily on expert experience and subjective judgment, and different assessors may use different judgment criteria, resulting in inconsistent and unobjective assessment results. Furthermore, traditional methods often rely on static, periodic checks, lacking the ability to monitor dynamic risk dynamics in real time and struggling to address rapidly evolving risk profiles. Furthermore, traditional methods often focus on assessing a single risk dimension, ignoring the correlations and transmission effects between risks and failing to fully grasp the systemic nature of risk. Finally, traditional methods are relatively weak in risk quantification and early warning mechanisms, making it difficult to accurately measure risk and provide timely warnings, resulting in passive and delayed risk management. Summary of the Invention

[0004] This application provides a financial compliance risk assessment method, system and storage medium based on big data, which is used to achieve efficient identification, accurate assessment and timely warning of compliance risks, and improve the automation level and scientific decision-making ability of financial institutions' compliance risk management.

[0005] In a first aspect, the present application provides a financial compliance risk assessment method based on big data, the financial compliance risk assessment method based on big data comprising: collecting and standardizing financial transaction data, customer information data, regulatory requirements data, and market data to obtain a structured data set;

[0006] A multi-level financial compliance risk assessment indicator system is constructed based on the structured data set, and the indicator system is a three-level structure including risk areas, risk subcategories and risk indicators; the structured data set and the multi-level financial compliance risk assessment indicator system are used to train a stacked fusion network including a decision tree, a neural network and a support vector machine to obtain a financial compliance risk assessment model; the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report; potential compliance risk points are identified according to the compliance risk assessment report, and weights are assigned to the compliance risk points for quantitative assessment to obtain a quantitative assessment result of the risk points; a graded warning threshold is set based on the quantitative assessment result of the risk points, and when the risk quantitative indicator exceeds the preset threshold, a warning signal including a risk point description, risk level and handling suggestion is generated, and the warning signal is sent to the corresponding risk management department.

[0007] In a second aspect, the present application provides a financial compliance risk assessment system based on big data, the financial compliance risk assessment system based on big data comprising:

[0008] The acquisition module is used to collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain structured data sets;

[0009] A construction module for constructing a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories, and risk indicators;

[0010] A training module, configured to use the structured data set and the multi-level financial compliance risk assessment indicator system to train a stacked fusion network comprising a decision tree, a neural network, and a support vector machine to obtain a financial compliance risk assessment model;

[0011] An analysis module, configured to input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report;

[0012] A quantification module is used to identify potential compliance risk points based on the compliance risk assessment report, and to assign weights to the compliance risk points for quantitative assessment to obtain a quantitative assessment result of the risk points;

[0013] A generation module is used to set a graded warning threshold based on the quantitative assessment results of the risk points. When the risk quantitative index exceeds the preset threshold, a warning signal containing a description of the risk point, risk level, and handling suggestions is generated, and the warning signal is sent to the corresponding risk management department.

[0014] In a third aspect, a financial compliance risk assessment device based on big data is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the financial compliance risk assessment device based on big data executes the above-mentioned financial compliance risk assessment method based on big data.

[0015] In a fourth aspect, a computer-readable storage medium is provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned financial compliance risk assessment method based on big data.

[0016] In the technical solution provided in this application, through the organic combination of technical features such as the collection and standardization of multi-source data, the construction of a multi-level risk assessment indicator system, the training and application of a stacked fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points and graded warning, a comprehensive assessment, accurate identification and timely warning of financial compliance risks are achieved, and significant technical effects are achieved. First, through the comprehensive collection and standardization of financial transaction data, customer information data, regulatory requirements data and market data, a unified structured data set is established, which solves the problems of single data source, non-uniform format and uneven quality in traditional methods, and provides a comprehensive and accurate data foundation for subsequent risk analysis; secondly, the construction of a multi-level financial compliance risk assessment indicator system realizes the hierarchical management of risk areas, risk subcategories and risk indicators, making risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak correlation in traditional methods; thirdly, the use of structured data sets and multi-level indicator systems to train a stacked fusion network including decision trees, neural networks and support vector machines fully utilizes the advantages of different machine learning algorithms. The decision tree model is good at processing risk judgments with clear rules, and the neural network model is good at capturing complex nonlinear The support vector machine model performs well in high-dimensional space. Through the fusion architecture, the advantages of each algorithm are complemented, which significantly improves the accuracy and stability of risk identification. At the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis to generate a comprehensive compliance risk assessment report, which realizes dynamic and continuous monitoring of risks and solves the problem of lagging risk assessment in traditional methods. In addition, potential compliance risk points are identified and quantitatively assessed based on the compliance risk assessment report. Through scientific weight allocation and calculation methods, qualitative risks are converted into quantitative indicators, making risk assessment more objective and accurate. Finally, based on the quantitative assessment results of risk points, hierarchical warning thresholds are set, and warning signals containing risk point descriptions, risk levels, and handling suggestions are generated, which realizes early warning and accurate push of risks, and provides a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considers the contribution of different algorithm characteristics to the solution, integrates the advantages of multiple algorithms through a layered fusion network, and realizes the intelligent, precise and automated compliance risk assessment. It effectively solves the problems of strong manual dependence, high subjectivity and poor real-time performance in traditional methods, and provides financial institutions with a comprehensive, efficient and accurate compliance risk management solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a schematic diagram of an embodiment of a financial compliance risk assessment method based on big data in an embodiment of the present application;

[0019] Figure 2 This is a schematic diagram of an embodiment of a financial compliance risk assessment system based on big data in an embodiment of the present application;

[0020] Figure 3 It is a schematic block diagram of the structure of a financial compliance risk assessment device based on big data in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The embodiments of the present application provide a financial compliance risk assessment method, system and storage medium based on big data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a financial compliance risk assessment method based on big data includes:

[0023] Step S101: Collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain a structured data set;

[0024] Step S102: Construct a multi-level financial compliance risk assessment indicator system based on the structured data set. The indicator system is a three-level structure including risk areas, risk subcategories, and risk indicators.

[0025] Step S103: Using the structured data set and the multi-level financial compliance risk assessment indicator system, a stacked fusion network including a decision tree, a neural network, and a support vector machine is trained to obtain a financial compliance risk assessment model;

[0026] Step S104: Input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report;

[0027] Step S105: Identify potential compliance risk points based on the compliance risk assessment report, and perform quantitative assessment on the weights assigned to the compliance risk points to obtain a quantitative assessment result of the risk points;

[0028] Step S106: Set a graded warning threshold based on the quantitative assessment results of risk points. When the risk quantitative index exceeds the preset threshold, generate a warning signal containing risk point description, risk level, and treatment suggestions, and send the warning signal to the corresponding risk management department.

[0029] It is understandable that the execution entity of this application can be a financial compliance risk assessment system based on big data, or a terminal or server, which is not limited here. The embodiment of this application is described by taking the server as the execution entity as an example.

[0030] Specifically, financial transaction data is collected from financial institutions' internal systems, including transaction records, fund flows, and account information. These data capture the basic attributes and behavioral characteristics of transactions. Customer information data is collected from customer management systems, including basic customer information, credit history, and transaction preferences, reflecting customer identity backgrounds and behavioral patterns. Regulatory data includes legal and regulatory texts, compliance guidelines, and penalty cases, providing a basis for compliance assessments. Market data includes market indicators such as interest rates, exchange rates, and stock prices, reflecting market changes. The collected raw data undergoes data cleansing to remove duplicate records, correct formatted data, and address missing values. The data is then formatted and unified. Numerical data is normalized to standardize the data range, and categorical data is encoded and converted to facilitate model processing. Finally, the data is annotated with risk-related tags to form a structured dataset. A multi-level financial compliance risk assessment indicator system is constructed based on this structured dataset. Financial compliance risks are categorized into risk areas: anti-money laundering compliance risk, customer identity verification risk, internal control risk, market manipulation risk, conflict of interest risk, tax compliance risk, and cross-border business compliance risk, forming a risk area classification framework. Each risk area is segmented, risk subcategories are established, and relationships between areas and subcategories are constructed. For example, anti-money laundering compliance risk is segmented into subcategories such as customer due diligence risk, abnormal transaction monitoring risk, and suspicious transaction reporting risk. Based on the secondary risk classification structure, specific risk indicators are designed for each risk subcategory, such as transaction frequency, transaction amount, customer credit score, and regulatory compliance numerical indicators, forming a three-level indicator system. The relative importance of each indicator is determined through the Analytic Hierarchy Process (AHP), and weights are assigned to form a weighted indicator system. Based on historical data analysis, risk threshold ranges are established, dividing the range into normal value ranges, concern value ranges, and warning value ranges. A dynamic adjustment mechanism is established to update indicator definitions and threshold settings based on regulatory policy changes and risk events.

[0031] A stacked fusion network was trained using a structured dataset and a multi-level financial compliance risk assessment indicator system. Training features were extracted from the structured dataset, and the training, validation, and test sets were divided in a 7:2:1 ratio. Using risk indicator values as input features and the corresponding risk levels as target labels, a decision tree model, a five-layer feedforward neural network model, and a radial basis kernel support vector machine model were trained, forming three basic models. Decision tree models excel at classification problems and identifying obvious risk signatures; neural network models are capable of capturing complex nonlinear relationships and hidden risk patterns; and support vector machines excel at processing high-dimensional data and establishing clear decision boundaries. The three basic models were applied to the validation set, and the F1 scores of each model were calculated for different risk domains to form a performance matrix. Based on this performance matrix, a stacked fusion architecture was constructed. The best-performing model for each risk domain was selected as the primary model, with the remaining models serving as auxiliary models. The model outputs were combined through weighted voting to form the initial fusion network. Grid search parameter optimization was performed on the validation set, adjusting the weight coefficients and decision thresholds of each basic model to obtain a parameter-optimized fusion network. The optimized fusion network is applied to the test set, and the average precision and recall of each risk area are calculated to form a financial compliance risk assessment model.

[0032] The structured dataset is input into the financial compliance risk assessment model for multi-dimensional real-time analysis. The structured dataset is grouped according to customer, transaction, business, and institution dimensions and converted into a feature matrix based on the indicator definitions of the multi-level financial compliance risk assessment indicator system, generating multi-dimensional feature data. This multi-dimensional feature data undergoes batch processing and real-time incremental processing, with the results of both types of processing combined to form a complete feature vector. The complete feature vector is then input into the decision tree model, five-layer feedforward neural network model, and radial basis kernel support vector machine model within the financial compliance risk assessment model, generating three sets of preliminary risk assessment results. Based on the weights assigned to the primary and auxiliary models corresponding to each risk area in the stacked fusion network, the three sets of preliminary risk assessment results are weighted and fused to generate risk scores and risk probability values for each risk indicator, forming a comprehensive risk assessment result. Following the indicator hierarchy within the multi-level financial compliance risk assessment indicator system, the risk scores are aggregated and calculated step by step from the risk indicator level to the risk subcategory and risk area levels, forming a hierarchical risk scoring table. The hierarchical risk scoring table is divided into risk levels according to pre-set score ranges. A risk distribution heat map, key risk indicator trend chart, and risk correlation network diagram are generated, which are then combined to form a compliance risk assessment report. Compliance risk points are identified and quantitatively assessed based on the compliance risk assessment report. The compliance risk assessment report is parsed to extract risk indicators with scores exceeding pre-set thresholds, including their name, risk score, subcategory, and risk area. This is then used to form a high-risk indicator set. Based on this high-risk indicator set, risk signature patterns are analyzed to identify abnormal transaction behaviors and compliance risks within the business process, creating a preliminary risk point list. Data is traced for each risk point in the preliminary risk point list, extracting associated original transaction records, customer information, and relevant business data to form a chain of evidence for the risk point and create a complete risk point profile. Based on the three core factors of risk severity, impact scope, and frequency of occurrence in the risk point profile, a weight coefficient is assigned to each risk point. The weight is determined through a multi-factor comprehensive calculation method to form a risk point weight table. Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, a composite calculation is performed to obtain a quantitative assessment value for each risk point, taking into account both direct and associated risks. According to the quantitative value of risk points, risk points are divided into four risk levels: critical, high, medium and low, and risk points are sorted according to risk levels to form quantitative assessment results of risk points.

[0033] A risk early warning mechanism is implemented based on the quantitative assessment results of risk points. Statistical analysis is performed on the quantitative assessment results of risk points to determine the boundaries of the quantitative score ranges for different risk levels, which serve as benchmark thresholds for risk warnings. This baseline threshold table is then created. Based on the baseline threshold table and historical risk management records, differentiated warning thresholds are set for different risk types and levels. A hierarchical warning threshold matrix is established, forming a complete warning threshold system. The comprehensive warning threshold system monitors the quantitative assessment results of risk points. When the quantitative risk indicator exceeds the corresponding warning threshold, the warning signal generation process is triggered, generating a raw warning trigger mark. Based on the raw warning trigger mark and the risk point data from the quantitative assessment results, a warning report is generated containing the risk point description, risk type, location, impact scope, specific numerical indicators, and development trends. This risk description report is then matched and analyzed with the management cases in the historical risk database to generate treatment recommendations for different risk types and levels, which are then integrated into a warning signal package. Based on the risk level and business area information in the warning signal package, the warning receiving department and personnel are identified. The warning signal package is then sent to the relevant risk management department via email, text message, or system notification, completing the warning distribution process.

[0034] In the embodiment of the present application, through the organic combination of technical features such as the collection and standardization of multi-source data, the construction of a multi-level risk assessment indicator system, the training and application of a stacked fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points and graded warning, a comprehensive assessment, accurate identification and timely warning of financial compliance risks are achieved, and significant technical effects are achieved. First, through the comprehensive collection and standardization of financial transaction data, customer information data, regulatory requirements data and market data, a unified structured data set is established, which solves the problems of single data source, non-uniform format and uneven quality in traditional methods, and provides a comprehensive and accurate data foundation for subsequent risk analysis; secondly, the construction of a multi-level financial compliance risk assessment indicator system realizes the hierarchical management of risk areas, risk subcategories and risk indicators, making risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak correlation in traditional methods; thirdly, the use of structured data sets and multi-level indicator systems to train a stacked fusion network including decision trees, neural networks and support vector machines fully utilizes the advantages of different machine learning algorithms. The decision tree model is good at processing risk judgments with clear rules, and the neural network model is good at capturing complex nonlinear The support vector machine model performs well in high-dimensional space. Through the fusion architecture, the advantages of each algorithm are complemented, which significantly improves the accuracy and stability of risk identification. At the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis to generate a comprehensive compliance risk assessment report, which realizes dynamic and continuous monitoring of risks and solves the problem of lagging risk assessment in traditional methods. In addition, potential compliance risk points are identified and quantitatively assessed based on the compliance risk assessment report. Through scientific weight allocation and calculation methods, qualitative risks are converted into quantitative indicators, making risk assessment more objective and accurate. Finally, based on the quantitative assessment results of risk points, hierarchical warning thresholds are set, and warning signals containing risk point descriptions, risk levels, and handling suggestions are generated, which realizes early warning and accurate push of risks, and provides a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considers the contribution of different algorithm characteristics to the solution, integrates the advantages of multiple algorithms through a layered fusion network, and realizes the intelligent, precise and automated compliance risk assessment. It effectively solves the problems of strong manual dependence, high subjectivity and poor real-time performance in traditional methods, and provides financial institutions with a comprehensive, efficient and accurate compliance risk management solution.

[0035] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0036] Extract financial transaction data such as transaction records, fund flows, and account information from the internal systems of financial institutions, and obtain customer information such as basic customer information, credit records, and transaction preferences from the customer management system to obtain the original data set;

[0037] Perform data cleaning on the original data set to obtain a cleaned data set, and convert the format of the cleaned data set to obtain a data set with a unified format;

[0038] Normalize the numerical data in the unified format dataset and perform encoding conversion on the categorical data in the unified format dataset to obtain a standardized dataset;

[0039] Annotate the standardized data set, add risk-related labels, and obtain a labeled data set;

[0040] Integrate and store the labeled data set into the data warehouse, establish data index and access control mechanism, and form a structured data set.

[0041] Specifically, financial transaction data is extracted from the financial institution's internal systems. By establishing data interfaces to connect to the institution's core business systems, transaction processing systems, and account management systems, transaction record data is extracted using API calls, database queries, or file transfers. This includes fields such as transaction serial number, transaction time, transaction amount, transaction type, and counterparty information. Fund flow data is extracted, including fields such as fund source, fund destination, fund amount, and transfer time. Account information data is extracted, including fields such as account opening time, account status, account balance changes, and related account information. Customer information data, including basic customer information, credit history, and transaction preferences, is also obtained from the customer management system. This data is extracted through a database connector and temporarily stored in a data buffer, forming the raw data set.

[0042] The original dataset requires data cleaning. First, duplicate data is identified and processed. Duplicate records are identified by comparing key fields, and the latest or most complete records are retained. Erroneous data, including formatting errors, value errors, and logical errors, are then identified and corrected. Missing data is then processed. For missing non-key fields, mean filling, mode filling, or predictive model filling are used. Records missing key fields are marked for manual review or directly eliminated. After data cleaning is completed, the cleaned dataset is formatted, converting data from different sources and structures into a unified format. This includes standardizing field naming, data types, and encoding methods, resulting in a unified dataset.

[0043] Normalize the numerical data in the uniformly formatted dataset to convert numerical data of different dimensions and ranges to a unified scale. Minimum-maximum normalization is primarily used to linearly transform the raw data to the interval [0, 1]. For fields susceptible to outliers, Z-score normalization is used to transform the data to a distribution with a mean of 0 and a standard deviation of 1. For data exhibiting skewed distributions, such as transaction amounts and transaction frequencies, a logarithmic transformation is performed before normalization to bring the data distribution closer to a normal distribution. Categorical data in the uniformly formatted dataset is also encoded, converting non-numeric categorical data into a numerical form that the model can process. For ordered categorical variables, ordinal encoding is used; for unordered categorical variables with a small number of values, one-hot encoding is used; for unordered categorical variables with a large number of values, embedded encoding or hash encoding is used.

[0044] Data annotation and risk-related labeling are performed on standardized datasets. This primarily involves identifying data records with similar characteristics based on the patterns of historical violations; adding compliance risk labels to data records based on regulatory requirements and compliance requirements; and identifying potential risk points based on expert experience and industry best practices. The annotation process utilizes a semi-automated approach, with preliminary labeling performed by a rules engine, followed by review and confirmation by compliance experts. For complex or ambiguous cases, decisions are made through consultation among multiple experts. Labeling includes information such as risk type, risk level, and risk rationale. This creates a labeled dataset, linking data records to risk information.

[0045] Consolidate and store labeled datasets in a data warehouse, establish a unified data storage architecture, and integrate data of different types and sources into the data warehouse to form a consistent data view. Establish a data indexing mechanism to create indexes for key fields to improve data query efficiency; establish a data access control mechanism to set different levels of data access rights to ensure data security and compliance; establish a data update mechanism to implement regular and real-time data updates to maintain data timeliness. Through these steps, a structured dataset is formed.

[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0047] Combining the data characteristics of the structured dataset, financial compliance risks are divided into anti-money laundering compliance risk, customer identity identification risk, internal control risk, market manipulation risk, conflict of interest risk, tax compliance risk, and cross-border business compliance risk, thus obtaining a risk area classification framework.

[0048] Subdivide each risk area in the risk area classification framework, establish risk subcategories under each risk area, and build the relationship between risk areas and risk subcategories to obtain a secondary risk classification structure;

[0049] Based on the two-level risk classification structure, numerical indicators of transaction frequency, transaction amount, customer credit score, and regulatory compliance are designed for each risk subcategory, forming a three-level indicator system;

[0050] Assign weights to each level of the three-level indicator system, determine the relative importance of each indicator in the overall evaluation through the analytic hierarchy process, and obtain a weighted indicator system;

[0051] Based on the weighted indicator system and historical data analysis, a risk threshold interval is established for each risk indicator, which is divided into normal value range, concern value range and warning value range to obtain an indicator system with thresholds;

[0052] Establish a dynamic adjustment mechanism for the indicator system with thresholds, update indicator definitions and threshold settings based on changes in regulatory policies and new risk events, and form a multi-level financial compliance risk assessment indicator system.

[0053] Specifically, risk areas are classified based on the data characteristics within structured datasets. By analyzing the characteristic patterns of financial transaction data, customer information data, regulatory requirements data, and market data within structured datasets, the main types of financial compliance risks are identified and categorized into seven core areas: anti-money laundering compliance risk, customer identification risk, internal control risk, market manipulation risk, conflict of interest risk, tax compliance risk, and cross-border business compliance risk. Anti-money laundering compliance risk refers to the compliance risk faced by financial institutions when they fail to effectively identify, monitor, and report suspicious transactions or fail to fulfill their customer due diligence obligations; customer identification risk refers to the risk of failing to accurately identify and verify customer identity information; internal control risk involves the risk of flaws in internal management mechanisms and operational procedures; market manipulation risk refers to the risk of influencing financial market prices through improper trading practices; conflict of interest risk refers to the risk arising from the misalignment between the interests of financial institutions or practitioners and those of their clients; tax compliance risk refers to the risk of failing to comply with tax laws and regulations; and cross-border business compliance risk refers to the risk of failing to meet the regulatory requirements of different countries or regions in cross-border financial activities. This classification is based on cluster analysis of risk characteristics within structured datasets and a review of regulatory provisions, forming a risk area classification framework.

[0054] Each risk area in the risk area classification framework is subdivided, risk subcategories are established under each risk area, and the relationship between risk areas and risk subcategories is constructed. Taking anti-money laundering compliance risk as an example, it is subdivided into subcategories such as customer due diligence risk, abnormal transaction monitoring risk, suspicious transaction reporting risk, and anti-terrorist financing risk; customer identity identification risk is subdivided into subcategories such as identity information authenticity risk, identity information integrity risk, and identity information update risk; internal control risk is subdivided into subcategories such as job responsibility risk, approval process risk, and internal supervision risk. Each risk area is usually subdivided into 3-5 risk subcategories. By defining the characteristic attributes and boundary conditions of the risk subcategories, the differences between the subcategories are clarified, and a mapping relationship is established between risk areas and risk subcategories to form a secondary risk classification structure. This step is completed through characteristic analysis of risk events in structured data sets, combined with expert experience and regulatory classification standards.

[0055] Based on the secondary risk classification structure, specific risk indicators are designed for each risk subcategory. The transaction frequency indicator is measured by counting the number of transactions per unit time and is used to identify abnormal transaction behavior. The transaction amount indicator assesses risk by analyzing the size and changing trends of the transaction amount. Large transactions or sudden changes in amount usually carry higher risks. The customer credit score indicator integrates factors such as the customer's credit history, repayment ability, and default record to provide a basis for customer risk assessment. The regulatory compliance indicator measures the extent to which transactions or behaviors comply with regulatory requirements and is scored by comparing them with regulatory regulations. For each risk subcategory, an appropriate indicator combination is selected based on its characteristics, and the indicator calculation method, data source, and update frequency are clarified to form a complete three-level indicator system. This step is achieved through statistical analysis and feature extraction of relevant data in the structured dataset.

[0056] Assign weights to each level of the three-level indicator system, and use the analytic hierarchy process (AHP) to determine the relative importance of each indicator in the overall assessment. The AHP is a multi-criteria decision-making method that decomposes complex problems into a hierarchical structure and determines the relative importance of each factor through pairwise comparison. Specific steps include: establishing a hierarchical model to divide risk assessment indicators into a target layer, a criterion layer, and an indicator layer; constructing a judgment matrix to form a judgment matrix by performing pairwise comparisons between each pair of elements in each layer; calculating weight vectors by using the eigenvalue method to calculate the eigenvectors of the judgment matrix as the relative weights of the elements at that level to the previous level; and performing a consistency test to verify the consistency of the judgment matrix and ensure the rationality of the weight assignment. Using the AHP, determine the weights of the risk area layer, risk subcategory layer, and risk indicator layer to form a complete weighted indicator system.

[0057] Based on a weighted indicator system and historical data analysis, risk threshold ranges are established for each risk indicator. First, historical risk event data is collected and the performance characteristics of each risk indicator in these events are analyzed. Statistical analysis methods such as quantile analysis and cluster analysis are used to determine the distribution patterns of indicator values. Based on the data distribution and risk level, indicator values are divided into three ranges: normal value range, concern value range, and warning value range. The normal value range indicates that the risk indicator value is at a safe level and does not require special attention. The concern value range indicates that the risk indicator value is abnormal and requires close monitoring. The warning value range indicates that the risk indicator value is significantly abnormal, indicating a high risk and requiring immediate action. By setting these threshold ranges, an indicator system with thresholds is formed, providing judgment criteria for risk monitoring and early warning.

[0058] Finally, a dynamic adjustment mechanism for the indicator system with thresholds was established to enable it to adapt to changes in the financial environment and regulatory requirements. This dynamic adjustment mechanism includes a regular evaluation mechanism to assess the effectiveness of the indicator system and analyze its contribution to risk identification; an event-triggered mechanism to trigger indicator updates when new regulatory policies or major risk events emerge; and a feedback optimization mechanism to continuously optimize indicator definitions and threshold settings based on actual usage feedback. Through these mechanisms, the indicator system is updated in a timely manner to ensure its continued effectiveness. The resulting multi-level financial compliance risk assessment indicator system provides a framework and standard for identifying, evaluating, and providing early warning of financial compliance risks.

[0059] For example, a bank applied this approach to construct an AML compliance risk assessment indicator system. First, based on transaction data and regulatory requirements, AML risk was subdivided into three subcategories: customer due diligence risk, transaction monitoring risk, and reporting compliance risk. Specific indicators were then designed: for customer due diligence risk, indicators included the proportion of high-risk customers and the completeness of identity information; for transaction monitoring risk, indicators included the frequency of large transactions and the amount of cross-border transactions; and for reporting compliance risk, indicators included the timeliness of suspicious transaction reporting and a report quality score. Weights were determined using the Analytic Hierarchy Process (AHP). For example, within the overall AML risk, customer due diligence risk was weighted 0.4, transaction monitoring risk 0.4, and reporting compliance risk 0.2. Thresholds were set for each indicator. For example, for the large transaction frequency indicator, fewer than five transactions per month were considered normal, 5-10 transactions were considered concerning, and more than 10 transactions were considered warning. When regulators issued new AML regulations, indicator definitions and threshold settings were immediately updated. This multi-tiered indicator system enabled the bank to comprehensively assess and monitor AML compliance risks, effectively identify potential risks, and implement timely countermeasures.

[0060] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0061] Based on the correspondence between risk areas, risk subcategories, and risk indicators in the multi-level financial compliance risk assessment indicator system, training features were extracted from the structured data set. The training set, validation set, and test set were divided into a ratio of 7:2:1 to obtain the model training data set.

[0062] Using the risk indicator values in the model training dataset as input features and the risk levels corresponding to the risk indicators as target labels, we trained a decision tree model, a five-layer feedforward neural network model, and a radial basis kernel support vector machine model, respectively, to obtain three basic models.

[0063] The three basic models were applied to the validation set, and the F1 score of each model was calculated for each risk area in the multi-level financial compliance risk assessment indicator system, forming a performance matrix that contains the mapping relationship between risk areas and model performance;

[0064] A layered fusion architecture is constructed based on the performance matrix. For each risk domain, the model with the best performance is selected as the primary model, and the remaining models are used as auxiliary models. The output results of the three basic models are combined through weighted voting to obtain the initial fusion network.

[0065] Perform grid search parameter optimization on the initial fusion network on the validation set, adjust the weight coefficients and decision thresholds of each basic model, and obtain a fusion network with optimized parameters;

[0066] The parameter-optimized fusion network is applied to the test set, and the average precision and recall of each risk area are calculated to form a financial compliance risk assessment model. The financial compliance risk assessment model is used to receive structured data sets and output risk scores and risk levels.

[0067] Specifically, training features are extracted from structured datasets based on the correspondence between risk areas, risk subcategories, and risk indicators in the multi-level financial compliance risk assessment indicator system. This process essentially involves filtering out risk-assessment-related data fields from the structured dataset based on the indicator system definition and converting these fields into feature vectors that the model can process. During data extraction, the data source and calculation method for each risk indicator are determined, and then the corresponding data is retrieved from the structured dataset. For example, for the transaction frequency indicator, the number of transactions per unit time is counted from the transaction data table; for the transaction amount indicator, statistics such as the mean, variance, and maximum value of the transaction amount are calculated; and for the customer credit score indicator, the credit score is extracted from customer information data or calculated using a credit scoring model. The extracted feature data is divided into training, validation, and test sets in a 7:2:1 ratio. The training set accounts for 70% of the data used for model training, the validation set accounts for 20% of the data used for model selection and parameter tuning, and the test set accounts for 10% of the data used for final model performance evaluation. Stratified sampling is used in the data partitioning process to ensure that the risk level distribution in each subset is similar to that of the original dataset, thereby avoiding sampling bias.

[0068] Using the risk indicator values in the model training dataset as input features and the corresponding risk levels as target labels, three different machine learning models were trained. The decision tree model is a tree-based classification model that partitions the data space into distinct regions through recursive binary partitioning, with each region corresponding to a prediction result. During training, information gain or the Gini coefficient is selected as the partitioning criterion to find the optimal features and threshold for node splitting until a stopping condition is met. The decision tree model has the advantages of strong interpretability and the ability to reflect the importance ranking of features, making it suitable for problems with clear regularities in financial risk assessment. The five-layer feedforward neural network model is a neural network consisting of an input layer, three hidden layers, and an output layer, with neurons in each layer fully connected to the previous layer. The training process uses the backpropagation algorithm, using gradient descent to minimize the loss function between the predicted values and the true labels, while continuously adjusting the network weights and biases. The advantage of neural network models is that they can capture complex nonlinear relationships between features, making them suitable for processing hidden patterns in financial risk assessment. The radial basis kernel support vector machine model uses a radial basis function as the kernel function. It maps data into a high-dimensional space to find the decision boundary with the largest margin. During training, support vectors and model parameters are found by solving a quadratic programming problem, and a classification model is constructed. The advantages of the support vector machine model are its excellent performance in high-dimensional spaces and its strong resistance to overfitting, making it suitable for processing high-dimensional feature data used in financial risk assessment.

[0069] The three basic models were applied to the validation set, and the F1 score of each model was calculated for each risk area in the multi-level financial compliance risk assessment indicator system. The F1 score is the harmonic mean of precision and recall, calculated as F1 = 2 × (precision × recall) / (precision + recall), where precision is the proportion of positive predictions that are actually positive, and recall is the proportion of positive predictions that are actually positive. The F1 score, which comprehensively considers both precision and recall, is an important metric for evaluating classification model performance. It is particularly suitable for evaluating financial risk prediction models, as both false negatives (low recall) and false positives (low precision) have negative consequences in risk assessment. For each risk area (such as anti-money laundering compliance risk and customer identity verification risk), the F1 scores of the three basic models were calculated, forming a performance matrix that maps risk areas to model performance. The rows of the performance matrix represent different risk areas, the columns represent different models, and the matrix elements are the F1 scores of the corresponding models for the corresponding risk areas.

[0070] A cascading fusion architecture is constructed based on the performance matrix. For each risk domain, the best-performing model is selected as the primary model, with the remaining models serving as auxiliary models. Cascading fusion is an ensemble learning method that combines the predictions of multiple base models to generate more accurate and stable predictions. For each risk domain, the model with the highest F1 score, based on the performance matrix, is selected as the primary model and assigned a higher weight. The remaining two models serve as auxiliary models and are assigned lower weights. The outputs of the three base models are combined using a weighted voting method to form an initial fusion network. Specifically, for each prediction example, the three base models each output a prediction result and its confidence score. These are then weighted and summed according to preset weights to produce the final prediction result. The initial fusion network uses a linear weighting scheme: prediction result = w1 × model 1 result + w2 × model 2 result + w3 × model 3 result, where w1, w2, and w3 are corresponding weight coefficients, with initial values set based on the model's F1 score for the corresponding risk domain.

[0071] A grid search parameter optimization was performed on the validation set for the initial ensemble network, adjusting the weight coefficients and decision thresholds of each base model. Grid search is a hyperparameter optimization method that systematically tries different combinations in the parameter space to find the optimal parameter settings. The weight coefficients were searched in the interval [0, 1], and the decision threshold was searched in the interval [0.3, 0.7]. The step size was determined based on computing resources and time requirements. For each parameter combination, the ensemble network's performance metrics (such as the F1 score) were calculated on the validation set, and the optimal parameter combination was selected as the final parameter combination. By adjusting the weight coefficients, the contribution of different models to the final prediction was altered. By adjusting the decision threshold, the trade-off between precision and recall was balanced, resulting in the ensemble network achieving optimal performance in a specific risk domain. The parameter optimization process was iterated until the optimal parameter combination was found, forming a parameter-optimized ensemble network.

[0072] The parameter-optimized fusion network was applied to the test set, and the average precision and recall for each risk area were calculated to evaluate the model's final performance. Average precision, the average of the precisions at different thresholds, provides a more comprehensive assessment of model performance; recall reflects the model's ability to detect all real-world risk cases. Performance evaluation on the test set validated the fusion network's generalization capabilities, ensuring the model's ability to effectively identify financial compliance risks in real-world applications. The resulting financial compliance risk assessment model includes specialized prediction modules for different risk areas. It accepts structured datasets as input and outputs risk scores and risk levels, providing decision support for compliance risk management at financial institutions.

[0073] Taking a financial institution's compliance risk assessment as an example, the institution extracted 20,000 records containing various risk indicators from historical data, each associated with a post-validated risk level label. Data features included 50 indicators, including customer transaction frequency, transaction amount, account activity, and customer credit score, with the target label being the risk level (low, medium, and high). The data was partitioned in a 7:2:1 ratio, resulting in 14,000 training data items, 4,000 validation data items, and 2,000 test data items. When training the decision tree model, the Gini coefficient was used as the partitioning criterion, with a maximum depth of 5 and a minimum number of sample splits of 20. When training the five-layer feedforward neural network, the number of neurons in the three hidden layers was set to 128, 64, and 32, respectively. The ReLU activation function was used, the Adam optimizer was employed, and the learning rate was 0.001. When training the radial basis kernel support vector machine, the kernel parameter γ was set to 0.1, and the regularization parameter C was set to 10. The performance of the three models was evaluated on a validation set. In the anti-money laundering risk domain, the neural network achieved an F1 score of 0.83, the decision tree achieved 0.76, and the support vector machine achieved 0.79. In the customer identity risk domain, the decision tree performed best, with an F1 score of 0.85. Based on these results, a cascaded fusion network was constructed. For the anti-money laundering risk domain, the neural network weight was set to 0.6, while the other two models each had an F1 score of 0.2. For the customer identity risk domain, the decision tree weight was set to 0.6, with the others adjusted accordingly. By optimizing the weights and thresholds through grid search, the resulting fusion model achieved an average precision of 0.87 and a recall of 0.84 on the test set, significantly outperforming each of the individual models and demonstrating the effectiveness of the cascaded fusion architecture in financial compliance risk assessment.

[0074] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0075] The structured data set is grouped according to the customer dimension, transaction dimension, business dimension, and institution dimension. Based on the indicator definition of the multi-level financial compliance risk assessment indicator system, the features corresponding to the risk indicators are extracted to obtain a multi-dimensional feature matrix.

[0076] Perform data standardization and missing value processing on the multi-dimensional feature matrix to make it meet the input requirements of the financial compliance risk assessment model and obtain pre-processed feature data;

[0077] The pre-processed feature data were input into the decision tree model, five-layer feedforward neural network model, and radial basis kernel support vector machine model in the financial compliance risk assessment model to obtain the risk assessment values of the three models;

[0078] Based on the model performance mapping relationship of each risk area in the performance matrix, the risk assessment values of the three groups of models are weighted and fused. The risk score of each risk indicator is generated using the parameter-optimized fusion network to obtain the comprehensive risk assessment result.

[0079] Based on the comprehensive risk assessment results, and in accordance with the hierarchical relationship of the multi-level financial compliance risk assessment indicator system, the scores of the risk indicator level are weighted and aggregated to obtain the risk scores of the risk subcategory level and the risk area level in turn, forming a hierarchical risk score table;

[0080] Compare the hierarchical risk scoring table with the risk level classification standards to generate a risk level distribution chart and a risk trend analysis chart, which are combined into a compliance risk assessment report.

[0081] Specifically, structured datasets are grouped and features extracted according to different dimensions. Structured datasets are uniformly formatted after preliminary data collection, cleaning, and standardization. They contain financial transaction records, customer information, regulatory data, and more. The data grouping process categorizes the data into customer, transaction, business, and institution dimensions. Customer-dimensional grouping focuses on individual customer characteristics, including basic information, behavioral patterns, and credit status; transaction-dimensional grouping focuses on transaction-level characteristics, including transaction frequency, amount, and time distribution; business-dimensional grouping focuses on characteristics of different business lines, such as retail, corporate, and investment; and institution-dimensional grouping focuses on institutional compliance, such as branch compliance performance and departmental compliance. Based on these groupings, corresponding features are extracted from the data in each dimension, according to the indicator definitions of the multi-level financial compliance risk assessment indicator system. For example, from the customer dimension, features such as customer risk level and transaction activity are extracted; from the transaction dimension, features such as outliers in transaction amount and changes in transaction frequency are extracted; from the business dimension, features such as business compliance and business risk concentration are extracted; and from the institution dimension, features such as institutional risk exposure and compliance management effectiveness are extracted. In this way, a multi-dimensional feature matrix is formed, in which the rows represent the evaluation objects (such as customers, transactions, business lines, etc.) and the columns represent different risk indicator characteristics.

[0082] Perform data standardization and missing value processing on the multi-dimensional feature matrix to ensure that the data meets model input requirements. Data standardization is the process of converting features of different dimensions and distributions to a unified scale. It mainly includes two methods: min-max standardization and Z-score standardization. Min-max standardization maps feature values to the interval [0, 1] and is suitable for features with clear upper and lower bounds. Z-score standardization transforms features into a distribution with a mean of 0 and a standard deviation of 1 and is suitable for features that approximate a normal distribution. Select appropriate standardization methods for different feature types, such as using a logarithmic transformation followed by Z-score standardization on transaction amounts and using min-max standardization on risk scoring indicators. Missing value processing mainly uses methods such as mean filling, median filling, mode filling, or model prediction filling. For randomly missing numerical features, the mean or median is used; for missing categorical features, the mode is used; and for missing values that are strongly correlated with other features, the model prediction based on other features is used to fill them. Through standardization and missing value processing, complete, unified, and standardized preprocessed feature data is obtained, ensuring that the data quality and format meet the requirements of subsequent model analysis. The preprocessed feature data is fed into the three basic models within the financial compliance risk assessment model to generate risk assessments. The decision tree model constructs a tree-like structure for risk classification, dividing the feature space into distinct regions. Each leaf node corresponds to a risk prediction result. Feature data is filtered layer by layer through the internal nodes of the decision tree, ultimately reaching the leaf nodes to generate risk assessments, including the risk classification result and confidence level. The five-layer feedforward neural network model consists of an input layer, three hidden layers, and an output layer, using activation functions and weighted connections to transmit and transform information. Feature data is first normalized before input into the network. Nonlinear transformations are then performed in each hidden layer, ultimately generating a risk assessment value at the output layer. The radial basis kernel support vector machine model uses a kernel function to map features into a high-dimensional space and construct an optimal separating hyperplane for classification. After kernel transformation, the distance between the feature data and the support vector is calculated to generate a risk assessment value. These three models independently process the feature data, generating three different sets of risk assessment results, including risk classification results and corresponding probability values or confidence scores.

[0083] Based on the model performance mapping relationships across risk areas in the performance matrix, a weighted fusion calculation is performed on the risk assessment values of the three sets of models. This performance matrix, generated during the model training phase, contains each model's performance metrics (such as the F1 score) for each risk area. During the fusion calculation, the weights of each model in each risk area are determined based on the performance matrix, with models with better performance receiving higher weights. For each risk indicator, the assessment values corresponding to the three models are extracted and summed according to pre-set weights to produce the fusion assessment result. For example, in the anti-money laundering risk area, if the neural network model performs best, it will be given a higher weight in the fusion calculation; in the customer identification risk area, if the decision tree model performs best, its weight will be increased accordingly. Through a parameter-optimized fusion network, the outputs of the three basic models are intelligently combined, integrating the strengths of each model to generate a risk score for each risk indicator, forming a comprehensive risk assessment result.

[0084] Based on the comprehensive risk assessment results, a weighted aggregation calculation is performed on the risk indicator scores at each level, following the hierarchical relationship of the multi-level financial compliance risk assessment indicator system. The multi-level financial compliance risk assessment indicator system consists of three levels: risk area, risk subcategory, and risk indicator. Each upper-level node is composed of multiple lower-level nodes, each with a corresponding weight. The weighted aggregation calculation process is a bottom-up aggregation process. First, the risk score of the lowest-level risk indicator is obtained. Then, based on the weights set in the indicator system, the risk score of each risk subcategory is calculated. The specific calculation method is to multiply the score of each risk indicator under the risk subcategory by the corresponding weight and sum them to obtain the risk subcategory score. Similarly, the score of each risk subcategory under the risk area is multiplied by the corresponding weight and summed to obtain the risk area score. Through this layer-by-layer aggregation method, the scores of each risk area and the overall compliance risk are ultimately obtained, forming a hierarchical risk score table, which clearly shows the scoring and logical relationship from specific risk indicators to overall risk.

[0085] Compare the hierarchical risk scoring table with the risk classification standards to generate a risk level distribution chart and a risk trend analysis chart, which are then combined into a compliance risk assessment report. The risk classification standards are developed based on historical data analysis and regulatory requirements. Risks are generally divided into several levels, such as low risk, medium risk, and high risk, with each level corresponding to a specific scoring range. Compare the scores in the hierarchical risk scoring table with the classification standards to determine the risk level of each risk indicator, risk subcategory, and risk area. Based on the risk level information, a risk level distribution chart is generated to intuitively display the risk level distribution of different risk areas and different risk subcategories; a risk trend analysis chart is generated to show the trend of risk scores over time, facilitating the identification of risk change patterns. Finally, the risk scoring table, risk level distribution chart, risk trend analysis chart, and related explanatory notes are combined into a complete compliance risk assessment report, providing financial institutions with comprehensive risk assessment information and decision support.

[0086] For example, data is extracted from structured datasets for grouping and feature extraction. The customer dimension extracts basic information such as age, occupation, and income, as well as behavioral characteristics such as account activity and trading preferences. The transaction dimension extracts features such as transaction frequency, amount, and geographical distribution. The business dimension extracts features such as transaction volume and compliance monitoring indicators for each business line. The institution dimension extracts features such as branch risk control measures and compliance staffing. After combining these features into a multidimensional feature matrix, numerical data such as transaction amount are Z-score normalized, categorical data such as transaction type are one-hot encoded, and missing transaction IP addresses are filled with the most recent IP address to form preprocessed feature data. This feature data is then fed into three models: a decision tree model classifies customers based on features such as transaction frequency and transaction amount; a neural network model uses deep learning to capture the complex relationships in customer transaction patterns; and a support vector machine model handles nonlinear relationships between high-dimensional features. The performance matrix shows that the neural network performs best in transaction fraud detection (F1=0.86). For this purpose, the neural network weight was set to 0.6, while the other models were each set to 0.2. The decision tree performs best in customer identity risk assessment, and the weights were adjusted accordingly. Each customer's score for each risk indicator is calculated through weighted fusion and then aggregated according to the indicator hierarchy. For example, a customer scores 0.75 for the abnormal transaction frequency indicator and 0.3 for the abnormal transaction location indicator. After weighted aggregation, the customer's score for the transaction monitoring risk subcategory is 0.6. The weighted aggregation of the customer's scores across these three risk subcategories yields a total anti-money laundering risk score of 0.55. Based on the risk classification criteria (a score above 0.7 is high risk, 0.4-0.7 is medium risk, and below 0.4 is low risk), the customer is rated as medium risk. The generated risk assessment report includes the customer's risk score table, risk level distribution, and historical trend chart.

[0087] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0088] Parse the compliance risk assessment report and extract risk indicators whose risk scores exceed the preset threshold, including the risk indicator name, risk score, risk subcategory, and risk area, to obtain a set of high-risk indicators;

[0089] Analyze risk characteristic patterns based on a set of high-risk indicators, identify abnormal transaction behaviors and compliance risks in business processes, and obtain a preliminary list of risk points;

[0090] Conduct data tracing for each risk point in the preliminary risk point list, extract associated original transaction records, customer information, and related business data, form a risk point evidence chain, and obtain a risk point file;

[0091] According to the three core factors of risk severity, impact scope, and occurrence frequency in the risk point file, a weight coefficient is assigned to each risk point to obtain a risk point weight table;

[0092] Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, perform a composite calculation to obtain a quantitative assessment value for each risk point, taking into account both direct and associated risks to arrive at a quantitative value for the risk point;

[0093] According to the quantitative value of risk points, risk points are divided into four risk levels: critical, high, medium and low, and risk points are sorted according to risk levels to form quantitative assessment results of risk points.

[0094] Specifically, the compliance risk assessment report is parsed. The compliance risk assessment report is a comprehensive document generated by the fusion model, containing risk indicator scores, risk distribution, and trend analysis. The parsing process utilizes structured text analysis to extract information from the report regarding risk indicators whose scores exceed preset thresholds. These thresholds are established based on historical risk data analysis and regulatory requirements, and typically vary for different risk types. The parsing program compares the score of each risk indicator in the report against the corresponding threshold. When an indicator score exceeds a threshold, the parsing program extracts complete information for that indicator, including the indicator name (e.g., "Frequency of Large Transactions" or "Cross-border Capital Flows"), the actual risk score, the risk subcategory (e.g., "Transaction Monitoring Risk"), and the risk area (e.g., "Anti-Money Laundering Compliance Risk"). These high-risk indicators are aggregated into a high-risk indicator set, which serves as the basis for subsequent risk point identification.

[0095] Risk signature patterns are analyzed based on a collection of high-risk indicators to identify unusual transaction behaviors and compliance risks within business processes. Risk signature pattern analysis utilizes pattern recognition and anomaly detection techniques to match high-risk indicators against a predefined risk pattern library to identify potential unusual behavior patterns. This risk pattern library, a knowledge base built from historical risk cases, regulatory requirements, and industry best practices, contains characteristic descriptions of various known risks. The matching process not only considers individual indicator violations but also analyzes the combined relationships and temporal variations among multiple indicators to uncover complex risk patterns. For example, when both the "Frequency of Large-Value Transactions by Customers in a Short Period" and the "High Counterparty Risk Rating" indicators exceed the limits simultaneously, this may indicate a money laundering risk. The combination of "Frequent Cross-Border Fund Flows" and "Incomplete Customer Identity Information" may indicate the risk of illegal cross-border transactions. Through this signature pattern analysis, risk indicator violations are linked to specific business scenarios and transaction behaviors, resulting in a preliminary risk point list. Each risk point includes a risk type description, a combination of risk trigger indicators, and a preliminary risk assessment. Data is traced for each risk point in the preliminary risk point list to form a chain of evidence for the risk point. Data tracing is the process of tracing back to the original data source to collect original evidence supporting risk assessments. The tracing process begins by identifying data sources relevant to the risk point, primarily including transaction databases, customer information systems, and business processing systems. Query conditions are then designed to precisely extract data records directly related to the risk point. For example, for the "large-amount, frequent transactions" risk point, all recent transactions of that customer are traced, including detailed information such as transaction time, amount, counterparty, and transaction type. For the "customer identity anomaly" risk point, the customer's identity verification history, identification document information, and address change records are traced. Tracing is not limited to directly relevant data but also extends to related data, such as counterparty information, related account activity, and similar case studies. By integrating and analyzing multi-source data, a complete chain of evidence for risk is established, forming a detailed risk point profile and providing data support and evidence for each risk point.

[0096] Based on the information in the risk point profile, risk severity is assessed and weighted based on three core factors. Risk severity refers to the potential loss or negative impact caused by a risk event. This is determined by assessing factors such as potential financial losses, regulatory penalties, and reputational impact, and is categorized into four levels: extremely high, high, medium, and low. Impact refers to the business scope and customer base that a risk event could affect. This is determined by evaluating factors such as the number of business lines, number of customers, and capital size, and is also categorized into four levels. Frequency refers to the number of times a risk behavior occurs within a given period. This is assessed by statistically analyzing historical occurrences or predicting future likelihood, and is also categorized into four levels. These three core factors form a three-dimensional risk assessment framework. Each dimension is independently scored and then combined to form an overall risk point weight. This weighting process takes into account the varying characteristics of different risk types. For example, fraud risk generally prioritizes severity, while operational risk prioritizes frequency. Through this multi-dimensional assessment, each risk point is assigned an objective and reasonable weight, forming a risk point weighting table.

[0097] Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, a composite calculation is performed to obtain a quantitative assessment value for each risk point. Composite calculation is a scoring method that comprehensively considers multiple factors, taking into account not only the scores of the indicators directly associated with the risk point, but also indirectly associated risk factors. Direct risk refers to the risk level of the risk point itself, calculated by multiplying the score of the indicator directly associated with the risk point by a weight coefficient. Associated risk refers to the additional risk brought about by other risk points or risk factors associated with the risk point, determined through risk association map analysis. The composite calculation process first calculates the direct risk value, which is the weighted sum of the scores of each risk indicator associated with the risk point and the corresponding indicator weights. It then analyzes the correlations between risk points, identifies the risk transmission paths and impact levels, and calculates the associated risk value. Finally, the direct and associated risk values are combined according to a preset ratio to obtain the final quantitative assessment value for the risk point. This composite calculation method comprehensively considers the multidimensional attributes and associated impacts of risk, making the risk assessment results more comprehensive and accurate.

[0098] Based on the quantitative value of the risk point, the risk points are divided into four risk levels and ranked. The risk level division is a risk classification based on the size of the risk quantitative value, combined with the risk tolerance of the institution and regulatory requirements. Critical risk is the risk point with the highest quantitative value, indicating that the risk has become apparent or is about to occur, and immediate intervention measures are required; high risk means that the risk is in a high-incidence state and needs to be handled as a priority; medium risk means that the risk is near the warning line and requires regular monitoring; low risk means that the risk is within an acceptable range and only requires routine management. The risk level division not only considers the absolute size of the quantitative value, but also the relative position of the risk distribution. The quartile or natural breakpoint method is usually used to determine the level boundaries. After determining the risk level, the risk points are ranked according to the risk level and the size of the quantitative value to form a priority processing order and generate the final quantitative assessment results of the risk points. This result clearly shows the risk level, priority and mutual relationship of each risk point, providing a decision-making basis for subsequent risk response and management.

[0099] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0100] Conduct statistical analysis on the quantitative assessment results of risk points to determine the boundaries of the quantitative score intervals for different risk levels, which serve as the benchmark thresholds for risk warnings and generate a benchmark threshold table;

[0101] Based on the benchmark threshold table and historical risk disposal records, differentiated warning thresholds are set for different risk types and risk levels, and a hierarchical warning threshold matrix is established to obtain a warning threshold system;

[0102] Use the early warning threshold system to monitor the quantitative assessment results of risk points. When the risk quantitative index exceeds the corresponding early warning threshold, the early warning signal generation process is triggered to obtain the original early warning trigger mark;

[0103] Based on the original warning trigger mark and the risk point data in the risk point quantitative assessment results, a warning content is generated, including the risk point description, risk type, occurrence location, impact range, specific numerical indicators and development trends, to obtain a risk description report;

[0104] Match and analyze the risk description report with the disposal cases in the historical risk database to generate treatment suggestions for different risk types and risk levels, and integrate them into an early warning signal package;

[0105] Based on the risk level and business area information in the early warning signal package, the warning receiving department and personnel are determined, and the early warning signal package is sent to the corresponding risk management department via email, text message, or system notification to complete the early warning distribution.

[0106] Specifically, statistical analysis is performed on the quantitative risk assessment results to determine appropriate warning thresholds. This statistical analysis utilizes data analysis techniques to analyze the distribution characteristics of the quantitative risk values and identify natural data points as risk level boundaries. The specific process involves collecting the quantitative values of all risk points to form a dataset; analyzing the data distribution characteristics, including calculating statistical quantities such as the mean, standard deviation, and quartiles; and using quantile or cluster analysis methods to determine the data's natural data points. The quantile method determines boundaries based on the position of the data quantiles, such as the 25th, 50th, and 75th percentiles as the dividing lines between low, medium, and high risk. Cluster analysis uses clustering algorithms such as K-means to cluster the quantitative risk values into several categories, with the boundaries between these categories representing the level boundaries. This statistical analysis determines the quantitative score ranges corresponding to the four risk levels of critical, high, medium, and low, creating a baseline threshold table. The baseline threshold table maps risk levels to quantitative score ranges and serves as the basis for triggering warnings.

[0107] Based on the baseline threshold table and historical risk management records, differentiated early warning thresholds are set for different risk types and risk levels. The baseline threshold table provides a generalized early warning standard, but different risk types have varying importance and urgency in different business scenarios, necessitating the establishment of differentiated early warning thresholds. This differentiated setting process begins by analyzing historical risk management records and collecting historical case studies for different risk types, including information such as the quantitative value at the time of risk occurrence, the management process, and the results of the management. The sensitivity and tolerance levels of different risk types are calculated, with stricter early warning thresholds set for risk types with high sensitivity and low tolerance. Furthermore, considering business importance and regulatory requirements, more sensitive thresholds are set for key businesses and areas of high regulatory attention. Through this differentiated adjustment, a hierarchical early warning threshold matrix is established. This matrix is a multidimensional structure, with the horizontal axis representing different risk types (such as anti-money laundering risk and fraud risk) and the vertical axis representing different risk levels (critical, high, medium, and low). The matrix elements represent specific early warning threshold values. Together, the hierarchical early warning threshold matrix and the baseline threshold table form a complete early warning threshold system, providing precise criteria for risk monitoring. The early warning threshold system monitors the quantitative assessment results of risk points in real time. When a risk quantification indicator exceeds the corresponding early warning threshold, the early warning signal generation process is triggered. The monitoring process involves continuous data comparison, comparing the real-time calculated risk quantification indicator with the thresholds in the early warning threshold system. Specific implementation involves setting monitoring cycles and determining the monitoring frequency for different risk types. High-risk types, such as anti-money laundering compliance risks, may require daily or even real-time monitoring, while low-risk types, such as general operational risks, may require weekly or monthly monitoring. Regular batch monitoring is performed to check the quantitative values of all risk points according to a preset cycle. A real-time trigger mechanism is established to immediately assess new risk points or those with significant changes in risk values. When a risk quantification indicator exceeds the threshold set in the early warning threshold system, an initial early warning trigger mark is generated. The mark includes information such as the trigger time, trigger risk point ID, trigger threshold, actual value, and degree of exceedance, providing basic data for subsequent early warning content generation.

[0108] Detailed warning content is generated based on the original warning trigger marks and the risk point data from the risk point quantitative assessment results. Warning content generation is the process of converting the original warning marks into a human-understandable risk description report. The specific generation steps include extracting basic information about the risk point, obtaining basic information such as the risk point name, business to which it belongs, and involved customers from the risk point archive based on the risk point ID; compiling a risk point description, using clear and concise language to describe the nature, characteristics, and manifestations of the risk; organizing risk type information to clarify the category and subcategory to which the risk belongs; clarifying the location of the risk, including spatial location information such as business lines, departments, and regions; assessing the scope of risk impact, and analyzing the business scope, customer groups, and capital scale that the risk may affect; providing specific numerical indicators, including quantitative information such as the indicator value that triggers the warning and historical change trends; analyzing risk development trends, and predicting future risk changes through time series analysis. These contents are integrated into a structured risk description report with a unified report format, comprehensive content, and accurate description, providing a basis for subsequent disposal recommendations.

[0109] The risk description report is matched and analyzed with the disposal cases in the historical risk database to generate treatment recommendations. The historical risk database is a repository of risk disposal experience accumulated by financial institutions, containing information such as disposal methods and effect evaluations of various risk events in history. The matching and analysis process includes feature matching, which extracts the key features of the current risk point and matches them with the risk events in the historical case database to identify similar cases; case screening, which selects successful cases with good disposal results from the matched cases; experience extraction, which extracts experience such as disposal methods, key steps, and precautions from successful cases; and recommendation generation, which forms targeted treatment recommendations based on the extracted experience and the specific circumstances of the current risk. The treatment recommendations include emergency response recommendations, risk mitigation methods, responsible departments, and time limits. The risk description report and treatment recommendations together form a complete early warning signal package, which is a structured document containing all the information that needs to be conveyed to risk management personnel.

[0110] Based on the risk level and business domain information in the early warning signal package, the recipients of the warning are determined and the warning is distributed. Warning distribution is the process of accurately and promptly delivering the warning signal package to the relevant responsible individuals. The specific distribution process includes recipient identification: determining the departments and personnel receiving the warning information based on the risk level, risk type, and business affiliation, and forming a list of recipients; distribution channel selection: selecting appropriate notification channels based on the urgency of the risk and the characteristics of the recipient, such as instant messaging methods such as SMS and phone calls for high-urgency risks, and email or in-system messages for routine risks; distribution content customization: appropriately filtering and formatting the warning signal package content based on the recipient's role and responsibilities to ensure that the recipient receives information relevant to their responsibilities; distribution execution: sending the customized warning content to the recipient via the selected channel; and distribution confirmation: tracking the delivery and reading status of the warning information to ensure effective information transmission. This precise distribution ensures that risk warning information reaches the responsible departments and personnel in a timely manner, promoting timely risk response and resolution.

[0111] Taking a bank's credit business as an example, the bank analyzed the distribution of historical risk point quantitative values and discovered natural breakpoints in the data at 0.3, 0.5, and 0.7. Based on this, the bank categorized risk levels into four levels, creating a baseline threshold table: 0-0.3 is low risk, 0.3-0.5 is medium risk, 0.5-0.7 is high risk, and 0.7 and above is critical risk. Analysis of historical handling records revealed that fraud risks pose a greater threat and require longer handling cycles, requiring earlier intervention. Operational risks, on the other hand, are simpler to handle and have a higher tolerance. Based on this, a hierarchical early warning threshold matrix was established. For example, for fraud risk, the critical threshold was adjusted to 0.65, and the high threshold to 0.45; the threshold for operational risk remained unchanged. During routine monitoring, the system detected a risk value of 0.68 for a corporate loan customer, exceeding the critical fraud risk threshold of 0.65, and immediately generated a preliminary early warning flag. Based on this tag, the system extracts data such as basic customer information, loan status, and risk profile, generating a risk profile report. This report includes a description of the risk of "the enterprise provided false financial statements to obtain a loan," the risk type of "credit fraud," the location of the incident at "Corporate Banking Department Branch X," the scope of impact of "involving three loans totaling 5 million yuan," specific indicators of "financial data inconsistent with tax filings and recent large-scale abnormal fund transfers," and a trend analysis of "worsening risk." This report is then compared with historical fraud cases to extract actionable recommendations such as "immediate account freeze, on-site inspection, and initiation of legal proceedings," creating a comprehensive early warning signal package. Based on risk level and business affiliation, the alert recipients are identified as the Director of the Corporate Banking Department, the branch risk control manager, and the head office risk management department. Emergency contacts are notified via text message, and the complete early warning signal package is also sent via email, ensuring that relevant personnel can take swift action to address the risk.

[0112] The above describes the financial compliance risk assessment method based on big data in the embodiment of the present application. The following describes the financial compliance risk assessment system based on big data in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a financial compliance risk assessment system based on big data includes:

[0113] The collection module 201 is used to collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain a structured data set;

[0114] A construction module 202 is configured to construct a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories, and risk indicators;

[0115] A training module 203 is configured to use the structured data set and the multi-level financial compliance risk assessment indicator system to train a stacked fusion network including a decision tree, a neural network, and a support vector machine to obtain a financial compliance risk assessment model;

[0116] An analysis module 204 is configured to input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report;

[0117] Quantification module 205, configured to identify potential compliance risk points based on the compliance risk assessment report, and assign weights to the compliance risk points for quantitative assessment to obtain risk point quantitative assessment results;

[0118] The generation module 206 is used to set a graded warning threshold based on the quantitative assessment results of the risk points. When the risk quantitative index exceeds the preset threshold, a warning signal including a risk point description, risk level, and treatment suggestions is generated, and the warning signal is sent to the corresponding risk management department.

[0119] Through the collaborative efforts of the various components mentioned above, and through the organic combination of technical features such as the collection and standardization of multi-source data, the construction of a multi-level risk assessment indicator system, the training and application of a stacked fusion network, multi-dimensional real-time analysis, quantitative assessment of risk points, and graded warnings, a comprehensive assessment, accurate identification, and timely warning of financial compliance risks have been achieved, achieving significant technical results. First, through the comprehensive collection and standardization of financial transaction data, customer information data, regulatory requirements data, and market data, a unified structured data set has been established, solving the problems of single data sources, inconsistent formats, and uneven quality in traditional methods, providing a comprehensive and accurate data foundation for subsequent risk analysis; second, the construction of a multi-level financial compliance risk assessment indicator system has achieved hierarchical management of risk areas, risk subcategories, and risk indicators, making risk assessment more systematic and structured, and effectively avoiding the defects of scattered indicators and weak correlation in traditional methods; third, using structured data sets and a multi-level indicator system to train a stacked fusion network consisting of decision trees, neural networks, and support vector machines, fully leveraging the advantages of different machine learning algorithms. The decision tree model is good at processing risk judgments with clear rules, and the neural network model is good at capturing complex nonlinear data. The support vector machine model performs well in high-dimensional space. Through the fusion architecture, the advantages of each algorithm are complemented, which significantly improves the accuracy and stability of risk identification. At the same time, the structured data set is input into the financial compliance risk assessment model for multi-dimensional real-time analysis to generate a comprehensive compliance risk assessment report, which realizes dynamic and continuous monitoring of risks and solves the problem of lagging risk assessment in traditional methods. In addition, potential compliance risk points are identified and quantitatively assessed based on the compliance risk assessment report. Through scientific weight allocation and calculation methods, qualitative risks are converted into quantitative indicators, making risk assessment more objective and accurate. Finally, based on the quantitative assessment results of risk points, hierarchical warning thresholds are set, and warning signals containing risk point descriptions, risk levels, and handling suggestions are generated, which realizes early warning and accurate push of risks, and provides a time window and action guide for risk disposal. The technical solution of the present invention applies artificial intelligence algorithms and models in the specific field of financial compliance risk management, fully considers the contribution of different algorithm characteristics to the solution, integrates the advantages of multiple algorithms through a layered fusion network, and realizes the intelligent, precise and automated compliance risk assessment. It effectively solves the problems of strong manual dependence, high subjectivity and poor real-time performance in traditional methods, and provides financial institutions with a comprehensive, efficient and accurate compliance risk management solution.

[0120] above Figure 2 The financial compliance risk assessment system based on big data in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The financial compliance risk assessment device based on big data in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0121] Figure 3 This is a schematic diagram of the structure of a big data-based financial compliance risk assessment device provided by an embodiment of the present invention. This big data-based financial compliance risk assessment device 300 may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing applications 333 or data 332. The memory 320 and storage medium 330 may be either ephemeral or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instructions operating on the big data-based financial compliance risk assessment device 300. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instructions stored in the storage medium 330 on the big data-based financial compliance risk assessment device 300 to implement the steps of the above-described big data-based financial compliance risk assessment method.

[0122] The financial compliance risk assessment device 300 based on big data may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the big data-based financial compliance risk assessment device shown does not constitute a limitation on the big data-based financial compliance risk assessment device provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0123] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the big data-based financial compliance risk assessment method.

[0124] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a big data-based financial compliance risk assessment device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0126] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A financial compliance risk assessment method based on big data, characterized in that: The method comprises: Financial transaction data, customer information data, regulatory requirements data, and market data are collected and standardized to obtain a structured data set, including: extracting financial transaction data such as transaction records, fund flows, and account information from the internal systems of financial institutions, and obtaining customer information data such as basic customer information, credit records, and transaction preferences from the customer management system to obtain an original data set, wherein the transaction records include transaction serial number, transaction time, transaction amount, transaction type, and counterparty information; performing data cleansing on the original data set to obtain a cleaned data set, and performing format conversion on the cleaned data set to obtain a unified format data set; normalizing the numerical data in the unified format data set, and performing encoding conversion on the categorical data in the unified format data set to obtain a standardized data set; data labeling on the standardized data set, adding risk-related tags to obtain a labeled data set; integrating and storing the labeled data set in a data warehouse, establishing a data index and access control mechanism, and forming a structured data set; Constructing a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories, and risk indicators; The structured data set and the multi-level financial compliance risk assessment indicator system are used to train a stacked fusion network including a decision tree, a neural network and a support vector machine to obtain a financial compliance risk assessment model, including: extracting training features from the structured data set according to the correspondence between risk areas, risk subcategories and risk indicators in the multi-level financial compliance risk assessment indicator system, and dividing the training set, validation set and test set into a ratio of 7:2:1 to obtain a model training data set; using the risk indicator value in the model training data set as an input feature and the risk level corresponding to the risk indicator as a target label, respectively training a decision tree model, a five-layer feedforward neural network model and a radial basis kernel support vector machine model to obtain three basic models; applying the three basic models to the validation set, and For each risk area in the regulatory risk assessment indicator system, the F1 score of each model is calculated to form a performance matrix containing the mapping relationship between risk areas and model performance; based on the performance matrix, a stacked fusion architecture is constructed, and for each risk area, the model with the best performance is selected as the main model, and the remaining models are used as auxiliary models. The output results of the three basic models are combined by weighted voting to obtain an initial fusion network; grid search parameter optimization is performed on the initial fusion network on the validation set, and the weight coefficient and decision threshold of each basic model are adjusted to obtain a parameter-optimized fusion network; the parameter-optimized fusion network is applied to the test set, and the average precision and recall rate of each risk area are calculated to form the financial compliance risk assessment model, which is used to receive the structured data set and output risk scores and risk levels; Inputting the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report, including: grouping the structured data set according to customer dimension, transaction dimension, business dimension and institution dimension, and extracting features corresponding to risk indicators according to the indicator definition of the multi-level financial compliance risk assessment indicator system to obtain a multi-dimensional feature matrix; performing data standardization and missing value processing on the multi-dimensional feature matrix to make it meet the input requirements of the financial compliance risk assessment model to obtain pre-processed feature data; inputting the pre-processed feature data into the decision tree model, five-layer feedforward neural network model and radial basis kernel support vector machine in the financial compliance risk assessment model respectively model, obtaining risk assessment values of the three groups of models; performing weighted fusion calculation on the risk assessment values of the three groups of models according to the model performance mapping relationship of each risk field in the performance matrix, generating risk scores of each risk indicator using the parameter-optimized fusion network, and obtaining a comprehensive risk assessment result; performing weighted aggregation calculation on the scores of the risk indicator level according to the hierarchical relationship of the multi-level financial compliance risk assessment indicator system, and sequentially obtaining risk scores of the risk subclass level and the risk field level to form a hierarchical risk score table; comparing the hierarchical risk score table with the risk level classification standard to generate a risk level distribution diagram and a risk trend analysis diagram, which are combined into the compliance risk assessment report; Identify potential compliance risk points based on the compliance risk assessment report, assign weights to the compliance risk points, and perform quantitative assessment to obtain quantitative risk point assessment results; A graded warning threshold is set based on the quantitative assessment results of the risk points. When the risk quantification index exceeds the preset threshold, a warning signal containing a description of the risk point, risk level, and handling suggestions is generated, and the warning signal is sent to the corresponding risk management department.

2. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The multi-level financial compliance risk assessment indicator system is constructed based on the structured data set. The indicator system is a three-level structure including risk areas, risk subcategories and risk indicators, including: Combining the data features in the structured dataset, financial compliance risks are divided into anti-money laundering compliance risk, customer identity identification risk, internal control risk, market manipulation risk, conflict of interest risk, tax compliance risk, and cross-border business compliance risk, thus obtaining a risk area classification framework. Subdivide each risk area in the risk area classification framework, establish risk subcategories under each risk area, and build a correlation between risk areas and risk subcategories to obtain a secondary risk classification structure; Based on the two-level risk classification structure, numerical indicators of transaction frequency, transaction amount, customer credit score, and regulatory compliance are designed for each risk subcategory to form a three-level indicator system; Assigning weights to the indicators at each level of the three-level indicator system, determining the relative importance of each indicator in the overall evaluation through the analytic hierarchy process, and obtaining a weighted indicator system; Based on the weighted indicator system and historical data analysis, a risk threshold interval is established for each risk indicator, and the normal value range, the concern value range and the warning value range are divided to obtain an indicator system with thresholds; Establish a dynamic adjustment mechanism for the indicator system with thresholds, update indicator definitions and threshold settings based on changes in regulatory policies and new risk events, and form the multi-level financial compliance risk assessment indicator system.

3. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The identifying of potential compliance risk points based on the compliance risk assessment report and the quantitative assessment of the compliance risk points by assigning weights to obtain quantitative assessment results of risk points include: Parsing the compliance risk assessment report, extracting risk indicators whose risk scores exceed a preset threshold, including the risk indicator name, risk score, risk subcategory, and risk area, to obtain a set of high-risk indicators; Analyze risk characteristic patterns based on the set of high-risk indicators, identify abnormal transaction behaviors and compliance risks in business processes, and obtain a preliminary list of risk points; Conduct data tracing for each risk point in the preliminary risk point list, extract associated original transaction records, customer information, and relevant business data, form a risk point evidence chain, and obtain a risk point file; According to the three core factors of risk severity, impact scope, and occurrence frequency in the risk point file, a weight coefficient is assigned to each risk point to obtain a risk point weight table; Using the risk point weight table and the risk indicator scores in the compliance risk assessment report, a composite calculation is performed to obtain a quantitative assessment value for each risk point, taking into account direct risks and associated risks to obtain a quantitative value for the risk point; According to the quantitative value of the risk point, the risk point is divided into four risk levels: critical, high, medium, and low, and the risk points are sorted according to the risk level to form the quantitative assessment result of the risk point.

4. The financial compliance risk assessment method based on big data according to claim 1 is characterized in that: The step of setting a graded warning threshold based on the quantitative risk assessment results is as follows: when the risk quantitative index exceeds the preset threshold, generating a warning signal including a risk point description, risk level, and treatment suggestions, and sending the warning signal to the corresponding risk management department, including: Performing statistical analysis on the quantitative assessment results of the risk points to determine the boundaries of the quantitative score intervals of different risk levels as benchmark thresholds for risk warning, and obtaining a benchmark threshold table; According to the benchmark threshold table and historical risk disposal records, differentiated warning thresholds are set for different risk types and risk levels, and a hierarchical warning threshold matrix is established to obtain a warning threshold system; The risk point quantitative assessment results are monitored using the early warning threshold system. When the risk quantitative index exceeds the corresponding early warning threshold, the early warning signal generation process is triggered to obtain the original early warning trigger mark; Based on the original warning trigger mark and the risk point data in the risk point quantitative assessment result, a warning content including risk point description, risk type, occurrence location, impact range, specific numerical indicators and development trend is generated to obtain a risk description report; Match and analyze the risk description report with the disposal cases in the historical risk database to generate treatment suggestions for different risk types and risk levels, and integrate them into an early warning signal package; According to the risk level and business area information in the early warning signal package, the early warning receiving department and personnel are determined, and the early warning signal package is sent to the corresponding risk management department via email, text message, or system notification to complete the early warning distribution.

5. A financial compliance risk assessment system based on big data, characterized in that: For implementing the financial compliance risk assessment method based on big data according to any one of claims 1 to 4, the financial compliance risk assessment system based on big data comprises: The acquisition module is used to collect and standardize financial transaction data, customer information data, regulatory requirements data, and market data to obtain a structured data set, including: extracting financial transaction data such as transaction records, fund flows, and account information from the internal system of the financial institution, and obtaining customer information data such as basic customer information, credit records, and transaction preferences from the customer management system to obtain an original data set, wherein the transaction record includes the transaction serial number, transaction time, transaction amount, transaction type, and counterparty information; performing data cleaning on the original data set to obtain a cleaned data set, and performing format conversion on the cleaned data set to obtain a unified format data set; normalizing the numerical data in the unified format data set, and performing code conversion on the categorical data in the unified format data set to obtain a standardized data set; data labeling on the standardized data set, adding risk-related tags to obtain a labeled data set; integrating and storing the labeled data set in a data warehouse, establishing a data index and access control mechanism, and forming a structured data set; A construction module for constructing a multi-level financial compliance risk assessment indicator system based on the structured data set, wherein the indicator system is a three-level structure including risk areas, risk subcategories, and risk indicators; A training module is used to use the structured data set and the multi-level financial compliance risk assessment indicator system to train a stacked fusion network including a decision tree, a neural network and a support vector machine to obtain a financial compliance risk assessment model, including: extracting training features from the structured data set according to the correspondence between risk areas, risk subcategories and risk indicators in the multi-level financial compliance risk assessment indicator system, and dividing the training set, validation set and test set into a ratio of 7:2:1 to obtain a model training data set; using the risk indicator value in the model training data set as the input feature and the risk level corresponding to the risk indicator as the target label, respectively training a decision tree model, a five-layer feedforward neural network model and a radial basis kernel support vector machine model to obtain three basic models; applying the three basic models to the validation set, and Calculate the F1 score of each model for each risk area in the financial compliance risk assessment indicator system to form a performance matrix containing the mapping relationship between risk areas and model performance; construct a stacked fusion architecture based on the performance matrix, select the model with the best performance for each risk area as the main model, and the remaining models as auxiliary models, and combine the output results of the three basic models through weighted voting to obtain an initial fusion network; perform grid search parameter optimization on the initial fusion network on the validation set, adjust the weight coefficient and decision threshold of each basic model, and obtain a parameter-optimized fusion network; apply the parameter-optimized fusion network to the test set, calculate the average precision and recall rate of each risk area, and form the financial compliance risk assessment model, which is used to receive the structured data set and output risk scores and risk levels; The analysis module is used to input the structured data set into the financial compliance risk assessment model for multi-dimensional real-time analysis to obtain a compliance risk assessment report, including: grouping the structured data set according to customer dimension, transaction dimension, business dimension and institution dimension, and extracting features corresponding to risk indicators according to the indicator definition of the multi-level financial compliance risk assessment indicator system to obtain a multi-dimensional feature matrix; performing data standardization and missing value processing on the multi-dimensional feature matrix to make it meet the input requirements of the financial compliance risk assessment model to obtain pre-processed feature data; and inputting the pre-processed feature data into the decision tree model, the five-layer feedforward neural network model and the radial basis kernel support model in the financial compliance risk assessment model respectively. A vector machine model is used to obtain risk assessment values for the three groups of models; based on the model performance mapping relationship of each risk area in the performance matrix, a weighted fusion calculation is performed on the risk assessment values of the three groups of models, and a risk score for each risk indicator is generated using the parameter-optimized fusion network to obtain a comprehensive risk assessment result; based on the comprehensive risk assessment result, a weighted aggregation calculation is performed on the scores of the risk indicator level in accordance with the hierarchical relationship of the multi-level financial compliance risk assessment indicator system to obtain risk scores for the risk subclass level and the risk area level in turn, forming a hierarchical risk score table; the hierarchical risk score table is compared with the risk level classification standard to generate a risk level distribution diagram and a risk trend analysis diagram, which are combined into the compliance risk assessment report; A quantification module is used to identify potential compliance risk points based on the compliance risk assessment report, and to assign weights to the compliance risk points for quantitative assessment to obtain a quantitative assessment result of the risk points; A generation module is used to set a graded warning threshold based on the quantitative assessment results of the risk points. When the risk quantitative index exceeds the preset threshold, a warning signal containing a description of the risk point, risk level, and handling suggestions is generated, and the warning signal is sent to the corresponding risk management department.

6. A financial compliance risk assessment device based on big data, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor implements the big data-based financial compliance risk assessment method described in any one of claims 1 to 4 when executing the computer program.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to execute the financial compliance risk assessment method based on big data according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Electrical operation risk assessment and early warning system based on big data analysis

    CN119539496A